SearcharxivSearch

arXiv subjects

Luyang Jiang

Publications and source records attributed to Luyang Jiang.

4 recordsLinked to original sources

SIMSplat: Language-Aligned 4D Gaussian Splatting for Driving Scenario Generation

Driving scene manipulation using real-world sensor data has emerged as a promising alternative to traditional driving simulators. Despite advances in language control and neural scene representations, existing methods treat grounding, editing, and simulation as loosely connected stages, relying on heuristic object localization, manual guidance, and single-agent validation, thereby constraining semantic expressiveness and hindering scalable, reactive scenario generation. We introduce SIMSplat, a driving scene editor built on scene-graph-based 4D Gaussian Splatting augmented with language-aligned features. By embedding appearance, motion, and location semantics directly into Gaussian scene-graph nodes, SIMSplat makes reconstructed scenes queryable through free-form natural language, bridging language understanding to object-level editing and multi-agent simulation within a single framework. Building on this language-grounded scene graph, SIMSplat supports diverse edits including fine-grained pedestrian manipulation, while a multi-agent path refinement module propagates changes across all agents to ensure reactive, physically plausible simulations. The pipeline further integrates with Vision-Language Models for automated scenario mining. Experiments show that SIMSplat more than doubles baseline grounding accuracy, achieves the highest task completion rate, and produces the lowest failure rates across diverse driving scenarios.

cs.RO

VoQA: Visual-only Question Answering

Visual understanding requires interpreting both natural scenes and the textual information that appears within them, motivating tasks such as Visual Question Answering (VQA). However, current VQA benchmarks overlook scenarios with visually embedded questions, whereas advanced agents should be able to see the question without separate text input as humans. We introduce Visual-only Question Answering (VoQA), where both the scene and the question appear within a single image, requiring models to perceive and reason purely through vision. This setting supports more realistic visual understanding and interaction in scenarios where questions or instructions are embedded directly in the visual scene. Evaluations under pure visual-only zero-shot, prompt-guided and OCR-assisted settings show that current models exhibit a clear performance drop compared to traditional VQA. To address this, we investigate question-alignment fine-tuning strategies designed to guide models toward interpreting the visual question prior to reasoning. Leveraging VoQA dataset together with these strategies yields robust vision-only reasoning while preserving cross-task generalization to traditional VQA, reflecting the complementary visual and textual reasoning capabilities fostered through VoQA training. The code and data are publicly available.

cs.CV

The Clever Hans Mirage: A Comprehensive Survey on Spurious Correlations in Machine Learning

Back in the early 20th century, a horse named Hans appeared to perform arithmetic and other intellectual tasks during exhibitions in Germany, while it actually relied solely on involuntary cues in the body language from the human trainer. Modern machine learning models are no different. These models are known to be sensitive to spurious correlations between non-essential features of the inputs (e.g., background, texture, and secondary objects) and the corresponding labels. Such features and their correlations with the labels are known as "spurious" because they tend to change with shifts in real-world data distributions, which can negatively impact the model's generalization and robustness. In this paper, we provide a comprehensive survey of this emerging issue, along with a fine-grained taxonomy of existing state-of-the-art methods for addressing spurious correlations in machine learning models. Additionally, we summarize existing datasets, benchmarks, and metrics to facilitate future research. The paper concludes with a discussion of the broader impacts, the recent advancements, and future challenges in the era of generative AI, aiming to provide valuable insights for researchers in the related domains of the machine learning community.

cs.LG

Direct Joule-Heated Non-Equilibrium Synthesis Enables High Performing Thermoelectrics

High-throughput synthesis of bulk inorganic materials is crucial for accelerating functional materials discovery but is hindered by slow, energy-intensive solid-state methods. We introduce Direct Joule-Heated Synthesis (DJS), a rapid, single-step and scalable solid-state synthesis technique achieving a $10^5$-fold speedup and 20,000x energy efficiency improvement over conventional synthesis. DJS enables the synthesis of dense, bulk chalcogenides ($\mathrm{Bi_{0.5}Sb_{1.5}Te_3}$, $\mathrm{AgSbTe_2}$), achieving a zT of 2.3 at 573 K in optimally Cd/Se co-doped $\mathrm{AgSbTe_2}$, one of the highest for polycrystalline materials at this temperature. DJS enables optimal co-doping and rapid, non-equilibrium solidification, producing lamellar microstructures, interfacial regions, and cation-ordered nanodomains that scatter all-scale phonons, achieving ultralow lattice thermal conductivity (~0.2 $W m^{-1} K^{-1}$ at 573 K). DJS establishes a new benchmark for scalable and fast synthesis, accelerating functional material discovery.

cond-mat.mtrl-sci