Searcharxiv⌕ Search

arXiv subjects

Xuan Luo

Publications and source records attributed to Xuan Luo.

At least 37 records · Page 2Linked to original sources

Discovery of a hybridization-wave electronic order in a van der Waals Kondo lattice

Kondo lattice systems, in which localized magnetic moments coherently hybridize with itinerant electrons, exhibit a rich landscape of emergent quantum phenomena. Within this framework, the hybridization strength itself has been theoretically proposed as a spatially modulated order parameter, giving rise to a so-called hybridization wave. However, direct experimental evidence of this quantum state has remained an outstanding challenge. Here, we report the direct observation of a hybridization wave in the layered transition metal dichalcogenide 6R-TaS2, a naturally occurring heterostructure composed of alternating 1T- and 1H-TaS2 layers. Using scanning tunneling microscopy and spectroscopy (STM/STS), we identify the hybridization gap in 1T layer, demonstrating the establishment of a coherent Kondo lattice. Notably, we discover that the hybridization gap present a uniaxial unit-cell doubling modulation, which breaks the both translational and rotational symmetries of the underlying Star-of-David superlattice. Such unit-cell doubling is not caused by structural topography, and therefore, constitutes the real-space visualization of the hybridization-wave order. Furthermore, the hybridization wave correlates with an energy-dependent nematic order that shares the same periodicity and orientation, revealing intertwined electronic instabilities. Our findings not only validate a long-standing prediction but also establish layer-engineered van der Waals materials as a versatile platform for exploring and controlling hybridization-driven quantum phases.

cond-mat.str-el↗

Speech-Omni-Lite: Portable Speech Interfaces for Vision-Language Models

While large-scale omni-models have demonstrated impressive capabilities across various modalities, their strong performance heavily relies on massive multimodal data and incurs substantial computational costs. This work introduces Speech-Omni-Lite, a cost-efficient framework for extending pre-trained Visual-Language (VL) backbones with speech understanding and generation capabilities, while fully preserving the backbones' vision-language performance. Specifically, the VL backbone is equipped with two lightweight, trainable plug-and-play modules, a speech projector and a speech token generator, while keeping the VL backbone fully frozen. To mitigate the scarcity of spoken QA corpora, a low-cost data construction strategy is proposed to generate Question-Text Answer-Text-Speech (QTATS) data from existing ASR speech-text pairs, facilitating effective speech generation training. Experimental results show that, even with only thousands of hours of speech training data, Speech-Omni-Lite achieves excellent spoken QA performance, which is comparable to omni-models trained on millions of hours of speech data. Furthermore, the learned speech modules exhibit strong transferability across VL backbones.

eess.AS↗

A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulness

This study reveals a critical safety blind spot in modern LLMs: learning-style queries, which closely resemble ordinary educational questions, can reliably elicit harmful responses. The learning-style queries are constructed by a novel reframing paradigm: HILL (Hiding Intention by Learning from LLMs). The deterministic, model-agnostic reframing framework is composed of 4 conceptual components: 1) key concept, 2) exploratory transformation, 3) detail-oriented inquiry, and optionally 4) hypotheticality. Further, new metrics are introduced to thoroughly evaluate the efficiency and harmfulness of jailbreak methods. Experiments on the AdvBench dataset across a wide range of models demonstrate HILL's strong generalizability. It achieves top attack success rates on the majority of models and across malicious categories while maintaining high efficiency with concise prompts. On the other hand, results of various defense methods show the robustness of HILL, with most defenses having mediocre effects or even increasing the attack success rates. In addition, the assessment of defenses on the constructed safe prompts reveals inherent limitations of LLMs' safety mechanisms and flaws in the defense methods. This work exposes significant vulnerabilities of safety measures against learning-style elicitation, highlighting a critical challenge of fulfilling both helpfulness and safety alignments.

cs.CR↗

Evaluating Proactive Risk Awareness of Large Language Models

As large language models (LLMs) are increasingly embedded in everyday decision-making, their safety responsibilities extend beyond reacting to explicit harmful intent toward anticipating unintended but consequential risks. In this work, we introduce a proactive risk awareness evaluation framework that measures whether LLMs can anticipate potential harms and provide warnings before damage occurs. We construct the Butterfly dataset to instantiate this framework in the environmental and ecological domain. It contains 1,094 queries that simulate ordinary solution-seeking activities whose responses may induce latent ecological impact. Through experiments across five widely used LLMs, we analyze the effects of response length, languages, and modality. Experimental results reveal consistent, significant declines in proactive awareness under length-restricted responses, cross-lingual similarities, and persistent blind spots in (multimodal) species protection. These findings highlight a critical gap between current safety alignment and the requirements of real-world ecological responsibility, underscoring the need for proactive safeguards in LLM deployment.

cs.CL↗

AEQ-Bench: Measuring Empathy of Omni-Modal Large Models

While the automatic evaluation of omni-modal large models (OLMs) is essential, assessing empathy remains a significant challenge due to its inherent affectivity. To investigate this challenge, we introduce AEQ-Bench (Audio Empathy Quotient Benchmark), a novel benchmark to systematically assess two core empathetic capabilities of OLMs: (i) generating empathetic responses by comprehending affective cues from multi-modal inputs (audio + text), and (ii) judging the empathy of audio responses without relying on text transcription. Compared to existing benchmarks, AEQ-Bench incorporates two novel settings that vary in context specificity and speech tone. Comprehensive assessment across linguistic and paralinguistic metrics reveals that (1) OLMs trained with audio output capabilities generally outperformed models with text-only outputs, and (2) while OLMs align with human judgments for coarse-grained quality assessment, they remain unreliable for evaluating fine-grained paralinguistic expressiveness.

cs.CL↗

Learning When Not to Attend Globally

When reading books, humans focus primarily on the current page, flipping back to recap prior context only when necessary. Similarly, we demonstrate that Large Language Models (LLMs) can learn to dynamically determine when to attend to global context. We propose All-or-Here Attention (AHA), which utilizes a binary router per attention head to dynamically toggle between full attention and local sliding window attention for each token. Our results indicate that with a window size of 256 tokens, up to 93\% of the original full attention operations can be replaced by sliding window attention without performance loss. Furthermore, by evaluating AHA across various window sizes, we identify a long-tail distribution in context dependency, where the necessity for full attention decays rapidly as the local window expands. By decoupling local processing from global access, AHA reveals that full attention is largely redundant, and that efficient inference requires only on-demand access to the global context.

cs.CL↗

Learning Efficient Fuse-and-Refine for Feed-Forward 3D Gaussian Splatting

Recent advances in feed-forward 3D Gaussian Splatting have led to rapid improvements in efficient scene reconstruction from sparse views. However, most existing approaches construct Gaussian primitives directly aligned with the pixels in one or more of the input images. This leads to redundancies in the representation when input views overlap and constrains the position of the primitives to lie along the input rays without full flexibility in 3D space. Moreover, these pixel-aligned approaches do not naturally generalize to dynamic scenes, where effectively leveraging temporal information requires resolving both redundant and newly appearing content across frames. To address these limitations, we introduce a novel Fuse-and-Refine module that enhances existing feed-forward models by merging and refining the primitives in a canonical 3D space. At the core of our method is an efficient hybrid Splat-Voxel representation: from an initial set of pixel-aligned Gaussian primitives, we aggregate local features into a coarse-to-fine voxel hierarchy, and then use a sparse voxel transformer to process these voxel features and generate refined Gaussian primitives. By fusing and refining an arbitrary number of inputs into a consistent set of primitives, our representation effectively reduces redundancy and naturally adapts to temporal frames, enabling history-aware online reconstruction of dynamic scenes. Our approach achieves state-of-the-art performance in both static and streaming scene reconstructions while running at interactive rates (15 fps with 350ms delay) on a single H100 GPU.

cs.CV↗

Anisotropy of linear magnetoresistance in Kagome metal ZrV$_6$Sn$_6$

The Kagome lattice has attracted extensive attention due to the diverse magnetic properties and non-trivial electronic states generated by its unique atomic arrangement, which provides an excellent system for exploring macroscopic quantum behavior. Here, we report the anomalous transport properties in 166-type Kagome metal ZrV$_6$Sn$_6$ single crystals. The quadratic and linear magnetoresistance (LMR) can be observed depending on the directions of the field and the current. Integrating Hall resistivity and quantum oscillation measurements, we found that the LMR could match well with the Abrikosov model. However, this model encounters difficulties in explaining the anisotropy of the magnetoresistance. To solve the issue, we extrapolate the Abrikosov model to the case of two-dimensional linear dispersion. It was found that when the field is parallel to the linear dependence momentum, the quantized energy is $ε_n^{\pm}$ = $\pm v\sqrt{p^2+2eHn/c}$, resulting in LMR. By contrast, when it is parallel to the non-linear dependence momentum, the energy is $ε_n^{\pm}$ = $\pm v\sqrt{2eHn/c}$, without yielding LMR. Through the combination of experiment and theory, the modified Abrikosov model could interpret the macroscopic quantum transport in ZrV$_6$Sn$_6$ crystal. The present research provides a new perspective for understanding the LMR behavior.

cond-mat.str-el↗

Investigating charmed hybrid baryons via QCD sum rules

We investigate charmed hybrid baryons using the QCD sum rule method within the framework of heavy quark effective theory. We construct twenty-eight interpolating currents for charmed hybrid baryons, seven of which are employed in QCD sum rule analyses of nineteen states with quark-gluon configurations $qqcg$, $qscg$, and $sscg$ ($q = u/d$). The masses of the lowest-lying charmed hybrid baryons in the $SU(3)$ flavor $\mathbf{6}_F$ representation are calculated to be $M_{Σ_{cg}(1/2^+)} = 3.36^{+0.27}_{-0.26}~\rm{GeV}$, $M_{Ξ^\prime_{cg}(1/2^+)} = 3.59\pm 0.20~\rm{GeV}$, and $M_{Ω_{cg}(1/2^+)} = 3.82\pm 0.21~\rm{GeV}$. We propose that future experiments search for these states via their $P$-wave decay channels $ND^{(*)}$, $ΛD^{(*)}$, and $ΞD^{(*)}$, respectively. Such investigations would provide valuable insight into the role of gluonic excitations in hadron structure.

hep-ph↗

A short review on QCD sum rule studies of P-wave single heavy baryons

Over the past few decades, the study of singly heavy baryons has entered a golden era, with numerous excited states observed by experimental collaborations. Various theoretical approaches have been developed to investigate their properties, with the QCD sum rule method being one of the most widely applied. This paper provides a review of these QCD sum rule studies. Over the last ten years, we have systematically studied $P$-wave singly heavy baryons using QCD sum rules and light-cone sum rules within the framework of heavy quark effective theory. These $P$-wave singly heavy baryons can explain many excited heavy baryons, including the $Λ_c(2595)^+$, $Λ_c(2625)^+$, $Ξ_c(2790)^{0/+}$, $Ξ_c(2815)^{0/+}$, $Σ_c(2800)^0$, $Ξ_c(2882)^0$, $Ξ_c(2923)^0$, $Ξ_c(2939)^0$, $Ξ_c(2965)^0$, $Ω_c(3000)^0$, $Ω_c(3066)^0$, $Ω_c(3090)^0$, $Ω_c(3050)^0$, $Ω_c(3119)^0$, $Λ_b(5912)^0$, $Λ_b(5920)^0$, $Ξ_b(6087)^0$, $Ξ_b(6095)^0/Ξ_b(6100)^-$, $Σ_b(6097)^\pm$, $Ξ_b(6227)^-$, $Ω_b(6316)^-$, $Ω_b(6330)^-$, $Ω_b(6340)^-$, and $Ω_b(6350)^-$, etc. Furthermore, we predict additional $P$-wave singly heavy baryons, including two $Λ_b$ states, two $Ξ_b$ states, three $Σ_b$ states, three $Ξ_b^\prime$ states, two $Ω_b$ states, two $Λ_c$ states, two $Ξ_c$ states, three $Σ_c$ states, and one $Ω_c$ state, all with relatively narrow decay widths, making them viable candidates for experimental observation. The study of singly heavy baryons is closely related to two meaningful questions:"What is the shortest possible lifetime of an observable particle?" and "How can one generally describe approximate (flavor) symmetries?".

hep-ph↗

Direct Multi-Token Decoding

Decoder-only transformers have become the standard architecture for large language models (LLMs) due to their strong performance. Recent studies suggest that, in pre-trained LLMs, early, middle, and late layers may serve distinct roles: Early layers focus on understanding the input context, middle layers handle task-specific processing, and late layers convert abstract representations into output tokens. We hypothesize that once representations have been processed by the early and middle layers, the resulting hidden states may encapsulate sufficient information to support the generation of multiple tokens using only the late layers, eliminating the need to repeatedly traverse the early and middle layers. We refer to this inference paradigm as Direct Multi-Token Decoding (DMTD). Unlike speculative decoding, our method introduces no additional parameters, auxiliary routines, or post-generation verification. Despite being trained on a limited dataset, a fine-tuned DMTD Qwen3-4B model has already demonstrated promising results, achieving up to a 2x speedup with only minor performance loss. Moreover, as shown in our scaling analysis, its performance is expected to further improve with larger training datasets.

cs.CL↗

QCD sum rule study of topped mesons within heavy quark effective theory

Motivated by the recent CMS observation of a near-threshold enhancement in top quark pair production, we investigate a novel class of hadronic systems containing a single top quark: the topped mesons ($t\bar{q}$, with $\bar q = \bar u, \bar d, \bar s$). In contrast to the extensively studied toponium ($t\bar{t}$) system, analyzed primarily within perturbative QCD, topped mesons offer a complementary nonperturbative probe of QCD dynamics in the heavy quark limit. These states are expected to exhibit longer lifetimes and narrower decay widths than toponium, as only a single top quark undergoes weak decay. We employ QCD sum rules within the framework of heavy quark effective theory to study the structure and mass spectrum of ground-state topped mesons. Our analysis predicts masses near 173.1 GeV, approximately 0.5-0.6 GeV above the top quark pole mass. Compared with singly topped baryons ($tqq$, with $q = u, d, s$) studied concurrently in [arXiv:2507.05895], topped mesons have a simpler quark composition and more favorable decay channels (a topped meson is anticipated to decay weakly into a $Υ$ meson and a charmed meson), enhancing their potential for both theoretical analysis and experimental discovery.

hep-ph↗

Adaptive Layer-skipping in Pre-trained LLMs

Various layer-skipping methods have been proposed to accelerate token generation in large language models (LLMs). However, limited attention has been paid to a fundamental question: How do computational demands vary across the generation of different tokens? In this work, we introduce FlexiDepth, a method that dynamically adjusts the number of Transformer layers used in text generation. By incorporating a plug-in router and adapter, FlexiDepth enables adaptive computation in LLMs without modifying their original parameters. Applied to Llama-3-8B, it skips 8 out of 32 layers while maintaining full benchmark performance. Our experiments reveal that computational demands in LLMs significantly vary based on token type. Specifically, generating repetitive tokens or fixed phrases requires fewer layers, whereas producing tokens involving computation or high uncertainty requires more layers. Despite the computational savings, FlexiDepth does not yet achieve wall-clock speedup due to varied skipping patterns and I/O overhead. To inspire future work and advance research on practical speedup, we open-sourced FlexiDepth and a dataset documenting its layer allocation patterns.

cs.CL↗

LVT: Large-Scale Scene Reconstruction via Local View Transformers

Large transformer models are proving to be a powerful tool for 3D vision and novel view synthesis. However, the standard Transformer's well-known quadratic complexity makes it difficult to scale these methods to large scenes. To address this challenge, we propose the Local View Transformer (LVT), a large-scale scene reconstruction and novel view synthesis architecture that circumvents the need for the quadratic attention operation. Motivated by the insight that spatially nearby views provide more useful signal about the local scene composition than distant views, our model processes all information in a local neighborhood around each view. To attend to tokens in nearby views, we leverage a novel positional encoding that conditions on the relative geometric transformation between the query and nearby views. We decode the output of our model into a 3D Gaussian Splat scene representation that includes both color and opacity view-dependence. Taken together, the Local View Transformer enables reconstruction of arbitrarily large, high-resolution scenes in a single forward pass. See our project page for results and interactive demos https://toobaimt.github.io/lvt/.

cs.CV↗

FastCuRL: Curriculum Reinforcement Learning with Stage-wise Context Scaling for Efficient Training R1-like Reasoning Models

Improving training efficiency continues to be one of the primary challenges in large-scale Reinforcement Learning (RL). In this paper, we investigate how context length and the complexity of training data influence the RL scaling training process of R1-distilled reasoning models, e.g., DeepSeek-R1-Distill-Qwen-1.5B. Our experimental results reveal that: (1) simply controlling the context length and curating the training data based on the input prompt length can effectively improve the training efficiency of RL scaling, achieving better performance with more concise CoT; (2) properly scaling the context length helps mitigate entropy collapse; and (3) carefully choosing the context length facilitates achieving efficient LLM training and reasoning. Inspired by these insights, we propose FastCuRL, a curriculum RL framework with stage-wise context scaling to achieve efficient LLM training and reasoning. Extensive experimental results demonstrate that FastCuRL-1.5B-V3 significantly outperforms state-of-the-art reasoning models on five competition-level benchmarks and achieves 49.6% accuracy on AIME 2024. Furthermore, FastCuRL-1.5B-Preview surpasses DeepScaleR-1.5B-Preview on five benchmarks while only using a single node with 8 GPUs and a total of 50% of training steps.

cs.CL↗

Scaling Transformer-Based Novel View Synthesis Models with Token Disentanglement and Synthetic Data

Large transformer-based models have made significant progress in generalizable novel view synthesis (NVS) from sparse input views, generating novel viewpoints without the need for test-time optimization. However, these models are constrained by the limited diversity of publicly available scene datasets, making most real-world (in-the-wild) scenes out-of-distribution. To overcome this, we incorporate synthetic training data generated from diffusion models, which improves generalization across unseen domains. While synthetic data offers scalability, we identify artifacts introduced during data generation as a key bottleneck affecting reconstruction quality. To address this, we propose a token disentanglement process within the transformer architecture, enhancing feature separation and ensuring more effective learning. This refinement not only improves reconstruction quality over standard transformers but also enables scalable training with synthetic data. As a result, our method outperforms existing models on both in-dataset and cross-dataset evaluations, achieving state-of-the-art results across multiple benchmarks while significantly reducing computational costs. Project page: https://scaling3dnvs.github.io/

cs.GR↗

Single spin asymmetry $A _ { U L } ^ { \sin ( 3 ϕ_ { h } - ϕ_{ R } ) }$ in dihadron production in SIDIS

In the field of particle physics, the phenomenon of dihadron production in semi-inclusive deep inelastic scattering (SIDIS) process has always been a significant focus. This paper focuses on the single longitudinal spin asymmetry $A_{UL }^{\sin(3ϕ_{h}-ϕ_{R})}$ in the dihadron production during this process and combines the transverse-momentum-dependent dihadron fragmentation function (DiFF) $H_1^{\perp}$ to deeply analyze its underlying mechanism. Here, the involved DiFF $H_1^{\perp}$ is the analogue of the Collins function for single-hadron production and it describes the fragmentation of a transversely polarized quark at leading twist. Recent studies have shown that the azimuthal asymmetry signal observed by the COMPASS collaboration in the dihadron SIDIS is weak. To reveal the reason for this small signal and to study the asymmetry, we calculate the unknown T-odd DiFF $H_1^{\perp}$ using the spectator model. The spectator model, widely used in SIDIS, describes the internal structure of hadrons and the hadronization mechanism. This model has successfully explained dihadron production in unpolarized and single-polarized processes. During the research process, while maintaining the transverse momentum dependence of the hadron pair, we employ the transverse momentum dependent(TMD) factorization framework, using this method and the model, we first simulate the asymmetry in the COMPASS energy region and compare it with experimental data. Furthermore, we predict the same asymmetry at the HERMES, expecting to provide valuable theoretical references for relevant experimental studies.

hep-ph↗

Global Motion Corresponder for 3D Point-Based Scene Interpolation under Large Motion

Existing dynamic scene interpolation methods typically assume that the motion between consecutive timesteps is small enough so that displacements can be locally approximated by linear models. In practice, even slight deviations from this small-motion assumption can cause conventional techniques to fail. In this paper, we introduce Global Motion Corresponder (GMC), a novel approach that robustly handles large motion and achieves smooth transitions. GMC learns unary potential fields that predict SE(3) mappings into a shared canonical space, balancing correspondence, spatial and semantic smoothness, and local rigidity. We demonstrate that our method significantly outperforms existing baselines on 3D scene interpolation when the two states undergo large global motions. Furthermore, our method enables extrapolation capabilities where other baseline methods cannot.

eess.IV↗