SearcharxivSearch

arXiv subjects

Jiaxin Luo

Publications and source records attributed to Jiaxin Luo.

4 recordsLinked to original sources

World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning

Vision-language models (VLMs) have shown strong performance on static visual understanding, yet they still struggle with dynamic spatial reasoning that requires imagining how scenes evolve under egocentric motion. Recent efforts address this limitation either by scaling spatial supervision with synthetic data or by coupling VLMs with world models at inference time. However, the former often lacks explicit modeling of motion-conditioned state transitions, while the latter incurs substantial computational overhead. In this work, we propose World2VLM, a training framework that distills spatial imagination from a generative world model into a vision-language model. Given an initial observation and a parameterized camera trajectory, we use a view-consistent world model to synthesize geometrically aligned future views and derive structured supervision for both forward (action-to-outcome) and inverse (outcome-to-action) spatial reasoning. We post-train the VLM with a two-stage recipe on a compact dataset generated by this pipeline and evaluate it on multiple spatial reasoning benchmarks. World2VLM delivers consistent improvements over the base model across diverse benchmarks, including SAT-Real, SAT-Synthesized, VSI-Bench, and MindCube. It also outperforms the test-time world-model-coupled methods while eliminating the need for expensive inference-time generation. Our results suggest that world models can serve not only as inference-time tools, but also as effective training-time teachers, enabling VLMs to internalize spatial imagination in a scalable and efficient manner.

cs.CV

Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models

Large Vision-Language Models (LVLMs) achieve strong performance on many multimodal tasks, but object hallucinations severely undermine their reliability. Most existing studies focus on the text modality, attributing hallucinations to overly strong language priors and insufficient visual grounding. In contrast, we observe that abnormal attention patterns within the visual modality can also give rise to hallucinated objects. Building on this observation, we propose Segmentation-based Attention Entropy (SAE), which leverages semantic segmentation to quantify visual attention uncertainty in a semantic space. Based on SAE, we further design a reliability score for hallucination detection and an SAE-guided attention adjustment method that modifies visual attention at inference time to mitigate hallucinations. We evaluate our approach on public benchmarks and in real embodied multimodal scenarios with quadruped robots. Experiments show that SAE reduces hallucinations without additional training, improving LVLM reliability.

cs.CV

Symmetry-driven giant magneto-optical Kerr effects in altermagnet hematite

Altermagnets have attracted tremendous interest for revealing intriguing physics and promising spintronics applications. In contrast to conventional antiferromagnets, altermagnets break both PT and Tt symmetries, and simultaneously exhibit spin-split band structures with a vanishing net magnetization. To quantify insulating altermagnets without conduction electron, we propose to use magneto-optical Kerr effect (MOKE) to identify the altermagnetic fingerprints. In particular, we demonstrate not only the giant MOKE responses, but also their connection with the orientations of Neel vectors at room temperature in altermagnet hematite alpha-Fe_2O_3. Specifically, under the Neel vector along the [1-100] axis, we find a giant polar Kerr rotation angle 93.4 mdeg in the (11-20) plane, which is allowed by the magnetic space group C2'/c'. Under the Neel vector along the [11-20] axis, we find a longitudinal Kerr angle 9.6 mdeg in the (0001) plane, which is allowed by the magnetic space group C2/c. Further, we show that such pronounced MOKE effects directly enable an optical imaging of altermagnetic domains, together with their reversible domain wall (DW) motion. Our studies not only suggest MOKE can be used to identify altermagnet candidates, but also signify the feasibility of exploring altermagnetic optical and DW spintronics, which could largely expand the current research paradigm of altermagnetism.

cond-mat.mtrl-sci

Direct visualization of electric current induced dipoles of atomic impurities

Learning the electron scattering around atomic impurities is a fundamental step to fully understand the basic electronic transport properties of realistic conducting materials. Although many efforts have been made in this field for several decades, atomic scale transport around single point-like impurities has yet been achieved. Here, we report the direct visualization of the electric current induced dipoles around single atomic impurities in epitaxial bilayer graphene by multi-probe low temperature scanning tunneling potentiometry as the local current density is raised up to around 25 A/m, which is considerably higher than that in previous studies. We find the directions of these dipoles which are parallel or anti-parallel to local current are determined by the charge polarity of the impurities, revealing the direct evidence for the existence of the carrier density modulation effect proposed by Landauer in 1976. Furthermore, by $in$ $situ$ tuning local current direction with contact probes, these dipoles are redirected correspondingly. Our work paves the way to explore the electronic quantum transport phenomena at single atomic impurity level and the potential future electronics toward or beyond the end of Moore's Law.

cond-mat.mes-hall