Searcharxiv⌕ Search

arXiv subjects

Xiaoyu Zhu

Publications and source records attributed to Xiaoyu Zhu.

At least 19 recordsLinked to original sources

Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents

Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning that provides little useful guidance for action generation while incurring substantial inference cost. To address this challenge, we introduce Selection-based Structured Reasoning (SSR), a framework that reformulates reasoning as selection instead of open-ended generation. SSR represents recurring high-level reasoning as pre-specified, reusable natural-language candidates. At each turn, the model selects from these reasoning candidates based on their likelihoods given the current context, without requiring an auxiliary task head. Using pre-specified reasoning traces enables parallel scoring, where teacher-forced prefilling computes token likelihoods concurrently within and across candidates using a shared context KV cache. We evaluate SSR on seven multimodal search benchmarks using 2B and 4B models. Across multiple reinforcement learning objectives and supervised fine-tuning, SSR delivers significant efficiency gains without sacrificing task performance. SSR achieves an average success rate competitive with leading search agents of the same scale, while reducing per-turn reasoning latency by over 90% and total per-question model inference latency by 28-54%. Project page: https://zfy0314.github.io/ssr-webpage/.

cs.LG↗

Coexistence and Interconversion of Multiple-Order Majorana Modes in Topological Superconductor Films with Varying Thickness

We theoretically investigate the thickness-dependent evolution of Majorana modes in $C_{2h}$-symmetric topological superconductor films (such as the recently discovered 2M-WS$_2$) proximity coupled with magnetic insulators. For sufficiently thick films, two Majorana bound states coexist as end modes at the surface and interface along a vortex line, with the interfacial mode evolving into a chiral Majorana edge mode upon increasing the proximity-induced exchange field. The intrinsic $C_{2h}$ crystalline symmetry selects two chiral Majorana modes circulating along the hinges on two of the four side surfaces of the film. When the penetration depth of the exchange field is sufficiently shallow, the two circulating modes are localized near the interface, but with qualitatively different subsequent evolutions. One of them further collapses to form two Majorana corner modes, while the other merges with the chiral Majorana mode circulating around the interface. Importantly, the corner modes are well decoupled from the interfacial chiral mode, thereby enabling an unprecedented coexistence of first-, second-, and third-order Majorana modes within a single material platform. We further show that such coexistence persists even in the ultrathin-film limit, where the electric-field-controlled two-dimensional $Z_2$ topology offers an extra advantage to readily interconvert the multiple-order Majorana modes. These findings highlight the pivotal role of the proper crystalline symmetry in enabling emergence, manipulation, and potential braiding of Majorana modes for demonstrating non-Abelian statistics and fault-tolerant quantum computation.

cond-mat.supr-con↗

Low-energy Muon-Nucleon scattering experiment: LUNE (White Paper)

The HIAF will provide high-intensity, high-quality muon beams with momenta from 0.5 to 7.5 GeV/c. This energy range is uniquely suited for precision muon scattering, bridging the gap between low-energy electron facilities and future high-energy lepton-ion colliders. In particular, HIAF will enable precision measurements with both positive and negative muon beams over a broad kinematic range, complementing existing electron-scattering facilities such as JLab, EicC and EIC. Based on HIAF muon source, the LUNE Collaboration has been established to address several fundamental questions in nuclear and particle physics, including the proton charge radius puzzle, nucleon electromagnetic structure, and the dynamics of quantum electrodynamics and hadronic interactions. The program proceeds in two phases, from elastic scattering to nucleon structure and beyond-Standard-Model searches. The experiment is expected to determine the proton charge radius with a precision of approximately 1.0\% using elastic muon-proton scattering. It will also perform systematic measurements of the proton electromagnetic form factors with both $μ^+$ and $μ^-$ beams, enabling precise studies of two-photon exchange effects and stringent tests of quantum electrodynamics. Beyond elastic scattering, LUNE will investigate TMD, gravitational form factors, and nuclear charge radii, providing new insights into the 3D structure of nucleons and nuclei. The experiment will further address important topics including Coulomb-distortion corrections, nuclear medium effects, and possible signatures of physics beyond the Standard Model. This white paper presents the scientific motivation, detector concept, expected performance, and long-term strategy of LUNE.

hep-ex↗

Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning

Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reasoning images, Visual CoT provides an intuitive mechanism for visual foresight but introduces substantial inference overhead, which is particularly problematic for proactive video reasoning. We ask whether models can learn to think visually during training while reasoning directly at inference. We introduce Internalized Visual Thinking (IVT), a post-training framework that jointly optimizes textual prediction and next-embedding prediction over unlabeled videos. Given a partially observed video, IVT predicts latent representations of future frames together with the target textual answer, encouraging the model to capture motion, object transitions, interactions, and latent intent. At inference, IVT generates the answer directly without synthesizing or re-encoding future frames. We conduct controlled studies across target representations, decoder designs, prediction horizons, data mixtures, training curricula, and predictive objectives. IVT improves over direct-answer fine-tuning on all six evaluation settings while retaining the same inference pathway. Compared with explicit Visual CoT, IVT achieves comparable or better performance and reduces average end-to-end latency by more than 5x. Together, our findings suggest that explicit pixel-space generation at inference time, as used in visual chain-of-thought, may not be necessary for effective proactive video reasoning. Predictive world modeling can be internalized during training to produce multimodal reasoners that are both more accurate and substantially more efficient.

cs.CV↗

On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMs

Reinforcement learning (RL) finetuning has become a key technique for enhancing large language models (LLMs) on reasoning-intensive tasks, motivating its extension to vision-language models (VLMs). While RL-tuned VLMs improve on visual reasoning benchmarks, they remain vulnerable to weak visual grounding, hallucinations, and over-reliance on textual cues. We show that simple, controlled textual perturbations, including misleading captions or incorrect chain-of-thought (CoT) traces, cause substantial drops in robustness and confidence, and that these effects are more pronounced when CoT consistency is taken into account across open-source multimodal reasoning models. In contrast, closed models exhibit similar failure modes but maintain markedly greater robustness and reasoning consistency, suggesting that the gap reflects a shortcoming in current open-source RL finetuning rather than an inherent limitation of the task. To better understand these vulnerabilities, we further analyze RL finetuning dynamics and uncover an accuracy-faithfulness trade-off: finetuning raises benchmark accuracy, but can simultaneously erode the reliability of the accompanying CoT and its robustness to contextual shifts. Although adversarial augmentation improves robustness, it does not by itself prevent faithfulness drift. Incorporating a faithfulness-aware reward can restore alignment between answers and reasoning, but when paired with augmentation, training risks collapsing onto shortcut strategies and robustness remains elusive. Together, these findings highlight the limitations of accuracy-only evaluations and motivate training and assessment protocols that jointly emphasize correctness, robustness, and the faithfulness of visually grounded reasoning.

cs.LG↗

An ILUES-based adaptive Gaussian process method for multimodal Bayesian inverse problems

Inverse problems are prevalent in both scientific research and engineering applications. In the context of Bayesian inverse problems, sampling from the posterior distribution can be particularly challenging when the forward models are computationally expensive. This challenge is further compounded when the posterior distribution is multimodal. To address this issue, we propose a Gaussian process (GP)-based method to indirectly build surrogates for the forward model. Specifically, the unnormalized posterior density is expressed as a product of an auxiliary density and an exponential GP surrogate. Iteratively, the auxiliary density converges to the posterior distribution, starting from an arbitrary initial density. However, the efficiency of GP regression is highly influenced by the quality of the training data. Therefore, we utilize the iterative local updating ensemble smoother (ILUES) to generate high-quality samples that are concentrated in regions with high posterior probability. Subsequently, based on the surrogate model and mode information extracted using a clustering method, Markov chain Monte Carlo (MCMC) with a Gaussian mixed (GM) proposal is used to draw samples from the auxiliary density. Through numerical examples, we demonstrate that the proposed method can accurately and efficiently represent the posterior with a limited number of forward simulations.

stat.CO↗

Quantum phase transition driven by competing intralayer and interlayer hopping in bilayer nickelates

Bilayer nickelates exhibit high-temperature superconductivity under proper hydrostatic pressure or epitaxial strain, signifying the emergence of quantum phase transitions whose physical mechanisms remain unclear. Using a minimal bilayer Hubbard model incorporating only the Ni-$d_{3z^2-r^2}$ orbitals, we demonstrate that a phase transition naturally arises from tuning the ratio of intralayer to interlayer hopping amplitudes. The transition point separates regimes with a rich interplay between superconducting and density-wave orders. In the regime of weaker intralayer hopping, the ground state is characterized by quasi-long-range spin-density-wave order. As the intralayer hopping increases, the system undergoes a transition marked by the opening of a finite spin gap and the disappearance of spin-density-wave order. Meanwhile, superconductivity is dramatically enhanced, accompanied by the emergence of quasi-long-range charge-density-wave order, indicating that the system enters Luther-Emery phase. This quantum phase transition, driven by the competition between intralayer and interlayer hopping, provides a plausible microscopic explanation for the experimentally observed correlation between the superconducting transition temperature and ratio of out-of-plane to in-plane lattice constants. Our findings reveal a possible link between the suppression of spin-density-wave order and the prominence of superconducting order, which may assist future efforts to optimize experimental conditions for further enhancing superconductivity in bilayer nickelates.

cond-mat.str-el↗

Irreducible tensor product modules over the Takiff Lie algebra for $\mathfrak{sl}_{2}$

In this paper, we construct a class of non-weight modules over the Takiff $\mathfrak{sl}_{2}$ by taking the tensor products of the irreducible free $U(\overline{\mathfrak{h}})$-modules of rank 1, where $\overline{\mathfrak{h}}$ is a natural Cartan subalgebra of the Takiff $\mathfrak{sl}_{2}$, with the irreducible highest weight modules. We characterize the irreducibility of these tensor product modules and determine the necessary and sufficient conditions for isomorphisms between them. We further prove that these non-weight modules are distinct from the known non-weight modules. Finally, we reformulate some tensor product modules over the Takiff $\mathfrak{sl}_{2}$ as induced modules derived from modules over certain subalgebras, and determine the necessary and sufficient conditions for the reducibility of these induced modules.

math.RT↗

Deep Learning-based Lightweight RGB Object Tracking for Augmented Reality Devices

Augmented Reality (AR) applications often require robust real-time tracking of objects in the user's environment to correctly overlay virtual content. Recent advances in computer vision have produced highly accurate deep learning-based object trackers, but these models are typically too heavy in computation and memory for wearable AR devices. In this paper, we present a lightweight RGB object tracking algorithm designed specifically for resource-constrained AR platforms. The proposed tracker employs a compact Siamese neural network architecture and incorporates optimization techniques such as model pruning, quantization, and knowledge distillation to drastically reduce model size and inference cost while maintaining high tracking accuracy. We train the tracker offline on large video datasets using deep convolutional neural networks and then deploy it on-device for real-time tracking. Experimental results on standard tracking benchmarks show that our approach achieves comparable accuracy to state-of-the-art trackers, yet runs in real-time on a mobile AR headset at around 30 FPS -- more than an order of magnitude faster than prior high-performance trackers on the same hardware. This work enables practical, robust object tracking for AR use-cases, opening the door to more interactive and dynamic AR experiences on lightweight devices.

cs.HC↗

Combinatorics of monoidal actions in Lie-algebraic context

This paper is, essentially, a survey related to the problem of understanding the combinatorics of the action of the monoidal category of finite dimensional modules over a simple finite dimensional Lie algebra on various categories of Lie algebra modules. A special attention is payed to the Lie algebras $\mathfrak{sl}_2$ and $\mathfrak{sl}_3$. A few new general results are collected at the end.

math.RT↗

OmniGeo: Towards a Multimodal Large Language Models for Geospatial Artificial Intelligence

The rapid advancement of multimodal large language models (LLMs) has opened new frontiers in artificial intelligence, enabling the integration of diverse large-scale data types such as text, images, and spatial information. In this paper, we explore the potential of multimodal LLMs (MLLM) for geospatial artificial intelligence (GeoAI), a field that leverages spatial data to address challenges in domains including Geospatial Semantics, Health Geography, Urban Geography, Urban Perception, and Remote Sensing. We propose a MLLM (OmniGeo) tailored to geospatial applications, capable of processing and analyzing heterogeneous data sources, including satellite imagery, geospatial metadata, and textual descriptions. By combining the strengths of natural language understanding and spatial reasoning, our model enhances the ability of instruction following and the accuracy of GeoAI systems. Results demonstrate that our model outperforms task-specific models and existing LLMs on diverse geospatial tasks, effectively addressing the multimodality nature while achieving competitive results on the zero-shot geospatial tasks. Our code will be released after publication.

cs.AI↗

Combinatorics of infinite rank module categories over finite dimensional $\mathfrak{sl}_3$-modules in Lie-algebraic context

We determine the combinatorics of transitive module categories over the monoidal category of finite dimensional $\mathfrak{sl}_3$-modules which arise when acting by the latter monoidal category on arbitrary simple $\mathfrak{sl}_3$-modules. This gives us a family of eight graphs which can be viewed as $\mathfrak{sl}_3$-generalizations of the classical infinite Dynkin diagrams.

math.RT↗

Direct demonstration of bulk-boundary correspondence in higher-order topological superconductors with chiral symmetry

A higher-order topological superconductor can experience topological phase transitions driven by variations in a bulk parameter without closing the bulk gap. This presents a challenge in establishing a direct bulk-boundary correspondence, as conventional bulk invariants change only upon the closure of the bulk gap. Our study of two-dimensional higher-order phases in the DIII and BDI symmetry classes, both characterized by chiral symmetry, demonstrates that zero-energy crossings facilitate a direct connection between the bulk Hamiltonian and Majorana zero modes at corners. These crossings, emerging as boundary conditions vary, can be identified from the bulk Hamiltonian. For both classes, we introduce a pair of topological invariants derived from these zero-energy crossings to characterize the higher-order topology. Phases in which at least one invariant assumes a nonzero value are anticipated to host Majorana corner modes. Moreover, these invariants may change with the closure of either bulk or edge gaps, thereby providing a clear and direct demonstration of bulk-boundary correspondence in higher-order phases. Our findings offer a promising framework for systematically exploring higher-order topology through boundary condition modulation.

cond-mat.mes-hall↗

STMT: A Spatial-Temporal Mesh Transformer for MoCap-Based Action Recognition

We study the problem of human action recognition using motion capture (MoCap) sequences. Unlike existing techniques that take multiple manual steps to derive standardized skeleton representations as model input, we propose a novel Spatial-Temporal Mesh Transformer (STMT) to directly model the mesh sequences. The model uses a hierarchical transformer with intra-frame off-set attention and inter-frame self-attention. The attention mechanism allows the model to freely attend between any two vertex patches to learn non-local relationships in the spatial-temporal domain. Masked vertex modeling and future frame prediction are used as two self-supervised tasks to fully activate the bi-directional and auto-regressive attention in our hierarchical transformer. The proposed method achieves state-of-the-art performance compared to skeleton-based and point-cloud-based models on common MoCap benchmarks. Code is available at https://github.com/zgzxy001/STMT.

cs.CV↗

Open-Vocabulary 3D Semantic Segmentation with Text-to-Image Diffusion Models

In this paper, we investigate the use of diffusion models which are pre-trained on large-scale image-caption pairs for open-vocabulary 3D semantic understanding. We propose a novel method, namely Diff2Scene, which leverages frozen representations from text-image generative models, along with salient-aware and geometric-aware masks, for open-vocabulary 3D semantic segmentation and visual grounding tasks. Diff2Scene gets rid of any labeled 3D data and effectively identifies objects, appearances, materials, locations and their compositions in 3D scenes. We show that it outperforms competitive baselines and achieves significant improvements over state-of-the-art methods. In particular, Diff2Scene improves the state-of-the-art method on ScanNet200 by 12%.

cs.CV↗

Infinite rank module categories over finite dimensional $\mathfrak{sl}_2$-modules in Lie-algebraic context

We study locally finitary realizations of simple transitive module categories of infinite rank over the monoidal category $\mathscr{C}$ of finite dimensional modules for the complex Lie algebra $\mathfrak{sl}_2$. Combinatorics of such realizations is governed by six infinite Coxeter diagrams. We show that five of these are realizable in our setup, while one (type $B_\infty$) is not. We also describe the $\mathscr{C}$-module subcategories of $\mathfrak{sl}_2$-mod generated by simple modules as well as the $\mathscr{C}$-module categories coming from the natural action of $\mathscr{C}$ on the categories of finite dimensional modules over Lie subalgebras of $\mathfrak{sl}_2$.

math.RT↗

Adversarially Masked Video Consistency for Unsupervised Domain Adaptation

We study the problem of unsupervised domain adaptation for egocentric videos. We propose a transformer-based model to learn class-discriminative and domain-invariant feature representations. It consists of two novel designs. The first module is called Generative Adversarial Domain Alignment Network with the aim of learning domain-invariant representations. It simultaneously learns a mask generator and a domain-invariant encoder in an adversarial way. The domain-invariant encoder is trained to minimize the distance between the source and target domain. The masking generator, conversely, aims at producing challenging masks by maximizing the domain distance. The second is a Masked Consistency Learning module to learn class-discriminative representations. It enforces the prediction consistency between the masked target videos and their full forms. To better evaluate the effectiveness of domain adaptation methods, we construct a more challenging benchmark for egocentric videos, U-Ego4D. Our method achieves state-of-the-art performance on the Epic-Kitchen and the proposed U-Ego4D benchmark.

cs.CV↗

Induced modules and central character quotients for Takiff $\mathfrak{sl}_{2}$

We construct a large new family of simple modules over Takiff $\mathfrak{sl}_{2}$. We prove that the quotient of the universal enveloping algebra of the Takiff Lie algebra for $\mathfrak{sl}_{2}$ by the ideal generated by a non-trivial central character is a simple algebra. In the case of the trivial central character, we show that the corresponding ideal is primitive by explicitly constructing a simple module whose annihilator coincides with that ideal. Together with the annihilators of simple $\mathfrak{sl}_{2}$-modules, we expect that the above ideals exhaust all primitive ideal.

math.RT↗