SearcharxivSearch

arXiv subjects

Xuefeng Zhang

Publications and source records attributed to Xuefeng Zhang.

At least 19 recordsLinked to original sources

GSDrive: Reinforcing Driving Policies by Multi-mode Future Trajectory Probing with 3D Gaussian Splatting Environment

End-to-end (E2E) autonomous driving aims to directly map sensory observations to driving actions, but its real-world deployment is hindered by evolving data distributions and the high cost of continual annotation. While combining imitation learning (IL) and reinforcement learning (RL) is a common strategy for policy improvement, conventional RL training relies on delayed, event-based rewards, where policies learn only from catastrophic outcomes such as collisions, leading to premature convergence to suboptimal behaviors. To address these limitations, we propose GSDrive, a framework that uses a differentiable 3D Gaussian Splatting (3DGS) environment for future-aware trajectory probing and reward shaping in E2E driving. GSDrive first learns a multi-mode trajectory probe via IL and then uses RL to evaluate multiple candidate futures in the 3DGS environment, converting their simulated returns into dense shaping rewards for policy optimization. This yields a cyclic hybrid IL-RL training loop, where IL supplies structured future priors and RL provides interactive feedback for iterative refinement. Evaluated on the reconstructed nuScenes dataset, our method outperforms other simulation-based RL approaches in closed-loop experiments. Code is available at https://github.com/ZionGo6/GSDrive.

cs.RO

Observation of Resonance of Kagome Flat Band Doublet

The interplay between local and itinerant electrons underpins many correlated and topological quantum states. Kagome lattices provide an ideal platform by hosting both flat (localized states) and dispersive bands (itinerant states), yet direct spectroscopic evidence of their dynamical coupling has remained elusive. Here we report the long-sought flat band resonance in the quasi-two-dimensional kagome bilayer material CsCr6Sb6. Using angle-resolved photoemission spectroscopy, transport measurements, and combined density functional theory and dynamical mean-field theory, we identify coexisting flat band doublets and dispersive bands near the Fermi energy. Upon cooling, the flat and dispersive bands exhibit a pronounced enhancement of spectral weight and hybridization, directly evidencing flat band resonance. Crucially, this emergence coincides with the onset of short-range antiferromagnetic correlations, contrasting sharply with conventional Kondo lattice behavior. Our findings demonstrate not only the long-sought flat band resonance in kagome materials, but also its unconventional correlation with magnetism.

cond-mat.str-el

Listen, Look, Drive: Coupling Audio Instructions for User-aware VLA-based Autonomous Driving

Vision Language Action (VLA) models promise an open-vocabulary interface that can translate perceptual ambiguity into semantically grounded driving decisions, yet they still treat language as a static prior fixed at inference time. As a result, the model must infer continuously shifting objectives from pixels alone, yielding delayed or overly conservative maneuvers. We argue that effective VLAs for autonomous driving need an online channel in which users can influence driving with specific intentions. To this end, we present EchoVLA, a user-aware VLA that couples camera streams with in situ audio instructions. We augment the nuScenes dataset with temporally aligned, intent-specific speech commands generated by converting ego-motion descriptions into synthetic audios. Further, we compose emotional speech-trajectory pairs into a multimodal Chain-of-Thought (CoT) for fine-tuning a Multimodal Large Model (MLM) based on Qwen2.5-Omni. Specifically, we synthesize the audio-augmented dataset with different emotion types paired with corresponding driving behaviors, leveraging the emotional cues embedded in tone, pitch, and speech tempo to reflect varying user states, such as urgent or hesitant intentions, thus enabling our EchoVLA to interpret not only the semantic content but also the emotional context of audio commands for more nuanced and emotionally adaptive driving behavior. In open-loop benchmarks, our approach reduces the average L2 error by $59.4\%$ and the collision rate by $74.4\%$ compared to the baseline of vision-only perception. More experiments on nuScenes dataset validate that EchoVLA not only steers the trajectory through audio instructions, but also modulates driving behavior in response to the emotions detected in the user's speech.

eess.AS

Tilt-to-length noise subtraction with pointing jitters from closed-loop dynamics for TianQin

TianQin is a proposed space-based mission for gravitational wave detection, employing a constellation of three drag-free satellites in high Earth orbits to form a laser interferometric observatory. A critical technical challenge is mitigating tilt-to-length (TTL) coupling noise, which is expected to be the third dominant noise source after laser frequency and clock noises. This noise is unavoidable in the presence of the residual angular movement of satellites, movable optical subassemblies (MOSAs), and test masses (TMs), and needs to be subtracted after reducing the first two types of noises using time-delay interferometry (TDI). Previous works have shown that TTL coupling coefficients can be estimated from the null TDI channel $\zeta$ and used for noise subtraction in other combinations. However, it was found that correlated MOSA yaw jitters have a negative impact on the TTL calibration, and the effects of realistic residual angular jitters from drag-free and pointing control (DFPC) are yet to be investigated. In this paper, we use closed-loop DFPC simulations to generate more realistic jitters in the science mode and test TTL calibration capability. Our simulations reveal that rotating only one MOSA is more favorable, compared to symmetrically rotating two MOSAs, for enhancing the accuracy of TTL coefficient estimation, while employing only high-frequency data (0.1 - 1 Hz). Moreover, we propose two other methods to further improve estimation accuracy. Firstly, using different null channel combinations, such as $C_3^{14}$, enhances the least squares estimation accuracy even in the case of high correlations in MOSAs' yaw jitters. Secondly, injecting different sinusoidal artificial maneuvers to the six MOSAs also shows improvements. These methods can help TianQin to meet the 0.3 pm/Hz$^{1/2}$ requirement after the TTL noise subtraction.

gr-qc

Detection of Earth's free oscillations utilizing TianQin

The measurement of Earth's free oscillations plays an important role in studying the Earth's large-scale structure. Space technology development presents a potential method to observe these normal modes by measuring inter-satellite distances. However, the disturbance from the Earth's low-degree gravity field makes it challenging for low Earth orbit gravity measurement satellites such as Gravity Recovery and Climate Experiment (GRACE) and TianQin-2 to extract signals from Earth's free oscillations directly. Here, we propose that by taking advantage of the high Earth orbit, the TianQin satellites can effectively avoid this disturbance, enabling direct measurement of Earth's free oscillations. We derive an analytical waveform to describe the response of Earth's free oscillations in TianQin. Based on this waveform, we use Bayesian analysis to extract the normal modes from numerical simulation data and perform parameter estimation. Our findings reveal that for a magnitude 7.9, Wenchuan-like earthquake, the resulting free oscillations will generate a signal that signal-to-noise ratio (SNR) is 73 in TianQin, and approximately 9 different modes can be distinguished. This result shows TianQin can open a new window to examine the Earth's free oscillations and study the Earth's interior and earthquakes independently from ground-based gravity measurement.

physics.geo-ph

A data-driven global ocean forecasting model with sub-daily and eddy-resolving resolution

High-fidelity ocean forecasting at high spatial and temporal resolution is essential for capturing fine-scale dynamical features, with profound implications for hazard prediction, maritime navigation, and sustainable ocean management. While conventional numerical models can generate sub-daily, eddy-resolving forecasts, they demand substantial computational resources and often struggle to maintain predictive skill at such fine scales. Data-driven models offer a promising alternative with significantly higher computational efficiency; however, most are constrained to daily outputs and show a rapid decay in accuracy when extended to sub-daily timescales. Here, we introduce TianHai, the first-of-its-kind global data-driven 6-hour forecasting model, which delivers predictions at 1/12{\deg} eddy-resolving resolution with a vertical extent down to 1,500 m. A key feature of TianHai is the integration of atmospheric forcings through FuXi-Atmosphere, a data-driven atmospheric forecasting system, which enables the explicit representation of air-sea coupling effects. Unlike conventional approaches, TianHai does not rely on numerical atmospheric models or external meteorological forecasts, making it a fully data-driven framework for coupled prediction. Benchmark experiments demonstrate that TianHai delivers state-of-the-art performance in forecasting temperature and salinity profiles, zonal and meridional currents, sea surface temperature, and sea level anomalies for lead times ranging from 1 to 10 days.

physics.ao-ph

Femtosecond low-threshold all-optical switching enabled by giant broadband optical nonlinearity from heteroatom doping

Ultrafast all-optical switching (AOS) is pivotal for advancing integrated photonic devices, from high-speed photonic information processing to next generation all-optical computing and communication networks. However, conventional nonlinear materials suffer from sluggish response time, high power threshold, weak and narrow-bandwidth optical nonlinearities, critically limiting their viability. Here, we report a heteroatom engineering strategy to overcome these limitations by designing zero-dimensional nitrogen-doped carbon quantum dots (N-CQDs) with nonlinear optical performance far exceeding the state-of-the-art. Leveraging spatial self-phase modulation (SSPM) and ultrafast pump-probe technique, we first demonstrate an all-in-one AOS platform, where femtosecond laser pulses serve dual roles as control and signal beams. The AOS simultaneously realizes ultrafast response time (520 fs), ultralow threshold energy (2.2 Wcm-2), and giant nonlinear refraction indexes (10-5 cm2/W) in the wide spectral range (400-1064 nm), yielding performance surpassing state-of-the-art nonlinear carbon materials (i.e. carbon nanotube) by orders of magnitude. Spectroscopic and bandgap analyses attribute these exotic performances to enhanced n-pi interaction enabled by nitrogen doping, which amplifies nonlinear polarization dynamics. Crucially, ultrafast fluorescence spectroscopy reveals a large two-photon absorption cross-section of the N-CQDs, challenging the conventional cognition that broadband SSPM necessitates single-photon excitation. This discovery unveils a multi-channel AOS rooted in synergistic single-photon and two-photon processes.. This work demonstrates a new paradigm for achieving ultrafast, broadband, and energy-efficient AOS by heteroatom doping engineering.

physics.optics

On fine alignment of transmitted beams for TianQin with far-field wavefront error

TianQin is a proposed space-based gravitational wave detector mission that employs inter-satellite laser interferometry. Suppressing measurement noise and achieving high sensitivity require accurate alignment of multiple onboard interferometers after laser link acquisition. However, due to huge armlengths and varying point-ahead angles, the fine alignment of the transmitted beams can be particularly challenging, which needs to take into account both received laser power and far-field wavefront errors. To tackle this issue for TianQin which has small point-ahead angle variations, we propose an efficient alignment strategy that relies on finding the maximum-intensity direction of the transmitted beam as the alignment reference. The direction can be estimated through a quatrefoil scan of the local transmitted beam and the corresponding intensity measurement from the remote satellite. Under TianQin's fixed-value compensation of the point-ahead angles, simulation results reveal that the proposed strategy is capable of aligning the transmitted beams within 20 nrad from the mean value of the point-ahead angles, while the tilt-to-length coupling associated with far-field wavefront error can meet the requirement given a transmitted beam aberration of $\lambda/40$ RMS.

gr-qc

Emergent dynamical Kondo coherence and competing magnetic order in a correlated kagome flat-band metal CsCr6Sb6

Correlated kagome metals host unique electronic states that enable exotic quantum phenomena. In the recently emerged CsCr6Sb6, these manifest through Kondo behavior from localized Cr-3d electrons and unprecedented band flattening near the Fermi level. Yet the intricate interplay among Kondo screening, magnetic frustration, and electronic correlations remains poorly understood-a fundamental gap we address through multifaceted experimental and theoretical approaches. Our angle-resolved photoemission spectroscopy measurements reveal electronic correlation-renormalized flat bands and muon spin relaxation study detect short-range magnetic order at TN ~ 80 K. Complementing these findings, density-functional theory and dynamical mean-field theory calculations identify a coherent-incoherent crossover at TN, with a remarkable restoration of coherence accompanying local moment suppression-an anomalous hallmark of Kondo behavior. Intriguingly, despite strong interlayer antiferromagnetic coupling, the system evades long-range magnetic order due to competing magnetic configurations separated by sub-meV energy differences. These insights establish CsCr6Sb6 as a prototypical platform for investigating dynamical Kondo screening in correlated flat-band systems, opening new avenues to study flat band physics and frustrated magnetism in correlated kagome lattices.

cond-mat.str-el

Emerging kinetic-exchange for the enhanced metallic ferromagnetism in CrGeTe$_3$ under pressure

The microscopic origin of ferromagnetism in correlated materials remains heavily debated, particularly for the competing mechanisms governing insulating versus metallic phases. In this work, we theoretically study the electronic structure evolution of CrGeTe$_{3}$ under pressure and provide a consistent explanation to three unique features of this system, i.e. the semiconducting ferromagnetism at low pressure, the metallic ferromagnetism at high pressure, and the enhanced Curie temperature in the metallic phase. We propose that it is the reduced electronic correlation and enhanced $d$-$p$ hybridization that universally drive the continuous evolution of CrGeTe$_{3}$ under pressure and glue the three distinct experimental observations. Central to our discovery is the dual role of metallicity -- it simultaneously establishes kinetically driven exchange via $d$-$p$ hybridization and enables Stoner-type magnetic instability, with the contribution also from the residual super-exchange. Our analyses reveal that {\it intraband} excitations dominate the pressure-enhanced $\omega_p^2$ and $T_c$ correlation. These findings establish $d$-$p$ hybridization and electronic correlation as the bridge between localized and itinerant magnetism, at least, in CrGeTe$_{3}$.

cond-mat.str-el

Multimodal Fused Learning for Solving the Generalized Traveling Salesman Problem in Robotic Task Planning

Effective and efficient task planning is essential for mobile robots, especially in applications like warehouse retrieval and environmental monitoring. These tasks often involve selecting one location from each of several target clusters, forming a Generalized Traveling Salesman Problem (GTSP) that remains challenging to solve both accurately and efficiently. To address this, we propose a Multimodal Fused Learning (MMFL) framework that leverages both graph and image-based representations to capture complementary aspects of the problem, and learns a policy capable of generating high-quality task planning schemes in real time. Specifically, we first introduce a coordinate-based image builder that transforms GTSP instances into spatially informative representations. We then design an adaptive resolution scaling strategy to enhance adaptability across different problem scales, and develop a multimodal fusion module with dedicated bottlenecks that enables effective integration of geometric and spatial features. Extensive experiments show that our MMFL approach significantly outperforms state-of-the-art methods across various GTSP instances while maintaining the computational efficiency required for real-time robotic applications. Physical robot tests further validate its practical effectiveness in real-world scenarios.

cs.AI

Probing Dark Matter's Gravitational Effects Locally with TianQin

In this study, we explore the potential of using TianQin missions to probe the local gravitational effects of dark matter. The TianQin project plans to launch satellites at both low and high orbits. High-precision orbit determination is expected to aid in detecting Earth's gravity or gravitational waves. By comparing the derived masses in low and high orbits, it is possible to constrain the amount of dark matter between the two spheres, hence placing a local constraint on dark matter's gravitational effect. Our results show the capability of TianQin in detecting the density of dark matter around Earth, with an ultimate sensitivity to a value of $10^{-8}\,\,{\rm kg\,\,m^{-3}}$. This detection limit surpasses the estimated bounds for the solar system and the observation results for our Galaxy by approximately 7 and 14 orders of magnitude, respectively.

gr-qc

Atomistic Simulations of Cation Distribution and Defect Effects on the Performance of Substituted Ferrites

This study investigates Mn-Zn ferrites (nominal composition \ce{Mn_{0.5}Zn_{0.5}Fe2O4}, MZF) substituted with tetravalent (\ce{Si^{4+}}), trivalent (\ce{Co^{3+}}), and divalent (\ce{Ca^{2+}}, \ce{Mg^{2+}}, \ce{Sn^{2+}}) ions. We comprehensively analyze how substitutions at specific tetrahedral and octahedral crystallographic sites modulate the spinel lattice's structural stability, electronic band structure, magnetic anisotropy, and electrical conductivity. Density functional theory (DFT) combined with Boltzmann transport theory is employed to probe the thermoelectric and phonon transport properties of pristine and doped MZF systems. Formation energy calculations indicate that substitutions with \ce{Si^{4+}}, \ce{Ca^{2+}}, and \ce{Mg^{2+}} enhance the thermodynamic stability of MZF, while \ce{Co^{3+}} and \ce{Sn^{2+}} substitutions exhibit slightly higher formation energies, indicating relatively lower stability. Electronic structure analyses confirm all substituted variants retain a finite band gap, preserving their semiconducting nature. Magnetic anisotropy energy (MAE) calculations reveal that ferrites with mixed octahedral/tetrahedral substitutions display a narrower MAE distribution, signifying more uniform magnetic anisotropy. Thermoelectric property analysis at 300 K demonstrates that multivalent ion doping at either crystallographic site reduces electrical conductivity ($\sigma$) while concurrently enhancing the Seebeck coefficient ($S$). This inverse correlation highlights a doping-induced trade-off, likely driven by increased carrier scattering at defect sites and modifications to the electronic density of states near the Fermi level.

cond-mat.mtrl-sci

FuXi-Ocean: A Global Ocean Forecasting System with Sub-Daily Resolution

Accurate, high-resolution ocean forecasting is crucial for maritime operations and environmental monitoring. While traditional numerical models are capable of producing sub-daily, eddy-resolving forecasts, they are computationally intensive and face challenges in maintaining accuracy at fine spatial and temporal scales. In contrast, recent data-driven approaches offer improved computational efficiency and emerging potential, yet typically operate at daily resolution and struggle with sub-daily predictions due to error accumulation over time. We introduce FuXi-Ocean, the first data-driven global ocean forecasting model achieving six-hourly predictions at eddy-resolving 1/12{\deg} spatial resolution, reaching depths of up to 1500 meters. The model architecture integrates a context-aware feature extraction module with a predictive network employing stacked attention blocks. The core innovation is the Mixture-of-Time (MoT) module, which adaptively integrates predictions from multiple temporal contexts by learning variable-specific reliability , mitigating cumulative errors in sequential forecasting. Through comprehensive experimental evaluation, FuXi-Ocean demonstrates superior skill in predicting key variables, including temperature, salinity, and currents, across multiple depths.

cs.LG

HD-PiSSA: High-Rank Distributed Orthogonal Adaptation

Existing parameter-efficient fine-tuning (PEFT) methods for large language models (LLMs), such as LoRA and PiSSA, constrain model updates to low-rank subspaces, limiting their expressiveness and leading to suboptimal performance on complex tasks. To address this, we introduce High-rank Distributed PiSSA (HD-PiSSA), a distributed PEFT approach that initializes orthogonal adapters across different devices and aggregates their delta updates collectively on W for fine-tuning. Unlike Data Parallel LoRA or PiSSA, which maintain identical adapters across all devices, HD-PiSSA assigns different principal components of the pre-trained weights to each GPU, significantly expanding the range of update directions. This results in over 16x higher effective updated ranks than data-parallel LoRA or PiSSA when fine-tuning on 8 GPUs with the same per-device adapter rank. Empirically, we evaluate HD-PiSSA across various challenging downstream tasks, including mathematics, code generation, and multi-task learning. In the multi-task setting, HD-PiSSA achieves average gains of 10.0 absolute points (14.63%) over LoRA and 4.98 points (6.60%) over PiSSA across 12 benchmarks, demonstrating its benefits from the extra optimization flexibility.

cs.LG

A Graph-based Verification Framework for Fact-Checking

Fact-checking plays a crucial role in combating misinformation. Existing methods using large language models (LLMs) for claim decomposition face two key limitations: (1) insufficient decomposition, introducing unnecessary complexity to the verification process, and (2) ambiguity of mentions, leading to incorrect verification results. To address these challenges, we suggest introducing a claim graph consisting of triplets to address the insufficient decomposition problem and reduce mention ambiguity through graph structure. Based on this core idea, we propose a graph-based framework, GraphFC, for fact-checking. The framework features three key components: graph construction, which builds both claim and evidence graphs; graph-guided planning, which prioritizes the triplet verification order; and graph-guided checking, which verifies the triples one by one between claim and evidence graphs. Extensive experiments show that GraphFC enables fine-grained decomposition while resolving referential ambiguities through relational constraints, achieving state-of-the-art performance across three datasets.

cs.CL

Modeling coupled constellation dynamics for TianQin under self-gravity

TianQin is a dedicated geocentric mission for space-based gravitational wave (GW) detection. Among its core technologies, the drag-free and pointing control subsystem (DFPCS) - consisting of suspension, drag-free and pointing controls - keeps the two test masses (TMs) centered and aligned within their housings while maintaining drag-free conditions and precise telescope pointing along the laser-arm directions. This results in orbit-attitude coupled dynamics for the constellation. The coupling is made more prominent due to satellite self-gravity, which requires compensation from DFPCS and generally makes the satellites deviate from pure free-fall orbits. Previous studies assumed that the orbit and attitude dynamics could be decoupled in numerical simulation, neglecting the back-action from the closed-loop control to orbit propagation. To address this, we develop a comprehensive model that can propagate the full 9-body (6 TMs + 3 satellites, orbits and attitudes) dynamics inter-dependently under the inter-satellite pointing and drag-free conditions. This paper is threefold. First, we reassess the applicability of the two TMs and telescope pointing scheme to TianQin using the new model, and confirm the previous conclusion. Second, to meet the constellation stability requirements, it is found that the DC common self-gravity in the flight direction should be minimized, or kept close for the three satellites. Finally, we simulate the long-range light path between two TMs/satellites with a precision of sub-pm/Hz$^{1/2}$, and the results support the decoupling of the closed-loop dynamics and high-precision orbit for computational efficiency. The method is instrumental to other future missions where the orbit-attitude coupling needs careful consideration.

gr-qc

Human-Like Robot Impedance Regulation Skill Learning from Human-Human Demonstrations

Humans are experts in physical collaboration by leveraging cognitive abilities such as perception, reasoning, and decision-making to regulate compliance behaviors based on their partners' states and task requirements. Equipping robots with similar cognitive-inspired collaboration skills can significantly enhance the efficiency and adaptability of human-robot collaboration (HRC). This paper introduces an innovative HumanInspired Impedance Regulation Skill Learning framework (HIImpRSL) for robotic systems to achieve leader-follower and mutual adaptation in multiple physical collaborative tasks. The proposed framework enables the robot to adapt its compliance based on human states and reference trajectories derived from human-human demonstrations. By integrating electromyography (EMG) signals and motion data, we extract endpoint impedance profiles and reference trajectories to construct a joint representation via imitation learning. An LSTM-based module then learns task-oriented impedance regulation policies, which are implemented through a whole-body impedance controller for online impedance adaptation. Experimental validation was conducted through collaborative transportation, two interactive Tai Chi pushing hands, and collaborative sawing tasks with multiple human subjects, demonstrating the ability of our framework to achieve human-like collaboration skills and the superior performance from the perspective of interactive forces compared to four other related methods.

cs.RO