SearcharxivSearch

arXiv subjects

Xingyu Zhang

Publications and source records attributed to Xingyu Zhang.

At least 19 recordsLinked to original sources

Reflex-Informed Neuromuscular Reinforcement Learning for Muscle-Driven Locomotion

Muscle-driven locomotion provides a physically grounded approach to generating realistic human movement. However, achieving both physiological plausibility and adaptability to changes in musculoskeletal capacity and external disturbances remains a fundamental challenge. To address this limitation, we propose a Reflex-Informed Neuromuscular Reinforcement Learning framework for muscle-driven locomotion. Within this framework, a fixed phase-dependent reflex controller serves as the underlying neuromuscular control mechanism, while the reinforcement learning policy produces four biomechanically meaningful residual parameters to modulate key reflex gains and thresholds associated with hip swing, knee support, and ankle propulsion according to the current state. Experimental results demonstrate that the proposed framework generates physiologically plausible locomotion with improved kinematic accuracy and dynamic consistency, as well as better bilateral symmetry and stride-to-stride consistency under nominal walking conditions. The learned policy remains robust under muscle weakness and external perturbations without retraining.

cs.RO

Implementation Possibility of Quantum Simulation for Quantum Molecular Dynamics

In this work, we explore the implementation possibility of quantum simulation for quantum molecular dynamics, in particular for reaction dynamics, though several implementations have already reported through quantum-classical mixed simulations ({\it Acc. Chem. Res.} {\bf 54} (2021), 4229 and {\it J. Phys. Chem. Lett.} {\bf xx} (2026), XXXX). To analyze this aspect, we examine (1) the conjugacy relation between quantum simulator and the target molecular system, (2) the wave function correspondence in quantum algorithm and classical algorithm for multi-dimensional dynamics, (3) problems arisen from real-valued classical algorithms, and finally (4) geometric phase arisen from the separation among the degrees of freedom (DOFs). As is well known, the aforementioned first and second points play fundamental roles in quantum simulation of quantum many-body systems, and the third and fourth points are theoretical issues that might introduce problems in classical and quantum computing. In this work, we mainly focus on the third and fourth points by analysis of the first two points by reviewing previously reported quantum-classical mixed implementations of quantum simulation. We also consider gauge freedom in high-dimensional quantum molecular dynamics that has been introduced recently, and then discuss possibility of advantages and disadvantages of quantum simulation for molecular reaction dynamics.

physics.chem-ph

GUPO: Gradient Uncertainty-aware Policy Optimization for Post-Training Large Language Models

Group Relative Policy Optimization (GRPO) has become a widely used approach for post-training Large Language Models (LLMs) for reasoning. In GRPO, the group gradients induced by different queries within the same mini-batch are directly averaged to form the policy update. However, these group gradients can point in conflicting directions. Our empirical analysis suggests that group-gradient conflicts tend to be associated with less effective policy updates, motivating the need for a reliable aggregated update direction under such conflicts. Standard GRPO aggregation treats the realized group gradients as deterministic contributions and does not account for differences in their reliability during aggregation. To address this issue, we propose Gradient Uncertainty-Aware Policy Optimization (GUPO), which models each group gradient as a random variable under a Bayesian formulation and estimates its probability distribution. GUPO then derives gradient uncertainty using a Dirichlet-based formulation and uses it to calibrate the contribution of each group gradient during aggregation. Extensive experiments on multiple benchmarks demonstrate the effectiveness of GUPO.

cs.LG

Ground-Truth Neighborhood Regularization for Reinforcement Learning Post-Training of Time Series Foundation Models

Time series forecasting (TSF) plays an important role in a wide range of real-world applications. Recently, time series foundation models (TSFMs), pretrained on large-scale datasets, have demonstrated strong generalization capabilities and emerged as an important paradigm for TSF. Reinforcement learning (RL) post-training has consequently attracted growing attention as a means of further improving their performance on downstream tasks. However, we find that, in certain forecast regions, RL post-training may gradually shift the output distributions of TSFMs away from the ground truth, thereby limiting their performance. We refer to this phenomenon as \textbf{suboptimal collapse}. Our analysis suggests that difficulty in initially sampling high-quality trajectories near the ground truth is an important contributing factor to suboptimal collapse. To address this issue, we propose Ground-Truth Neighborhood Regularization (GTN-R) for RL post-training of TSFMs. GTN-R uses the ground truth as a reference for locating high-quality regions and guides the model's probability mass toward the ground-truth neighborhood. This increases the probability of sampling high-quality trajectories, mitigates suboptimal collapse, and improves performance. Moreover, GTN-R can be flexibly integrated into various RL methods for TSFMs. Extensive experiments show its effectiveness.

cs.LG

Cold Dark Matter and Self-Interacting Dark Matter Interpretations of Cloud-9

Recently, the Five-hundred-meter Aperture Spherical Telescope discovered a gas-rich hydrogen cloud near M94 in the $21\,{\rm cm}$ band. Lacking an optical counterpart, this object, dubbed Cloud-9, has been identified as a compelling Reionization Limited \textsc{Hi} Cloud (RELHIC). RELHICs provide exceptionally clean laboratories for probing dark matter, free from the baryonic complexities associated with star formation and feedback. We show that the observed hydrogen column density profile of Cloud-9 is consistent with a gas cloud embedded in either a cuspy halo predicted by the standard cold dark matter (CDM) model or a cored halo produced by self-interacting dark matter (SIDM). In both cases, the halo must have an unusually diffuse central density. The best-fitting CDM halo lies around $7\sigma$ below the cosmological concentration--mass relation, whereas SIDM core-forming halos reduce the tension to only around $3\sigma$. We further identify Cloud-9 analogs in the Concerto suite of cosmological zoom-in simulations with velocity-dependent SIDM, demonstrating that RELHICs provide a promising new probe of dark matter self-interactions.

astro-ph.GA

TacPrint: A Wearable Fingertip Tactile Sensor for Human-to-Robot Contact Reproduction

Human-centric data collection is emerging as a significant paradigm for robot skill acquisition, but seamlessly integrating low-cost, scalable tactile sensing systems that capture fine-grained fingertip interactions without compromising natural operation remains a key challenge. This reduces the reliability of human-to-robot transfer in contact-rich tasks. In this work, we present TacPrint, a wearable fingertip tactile sensor, where protrusions on the inner surface of the silicone skin are aligned one-to-one with 24 capacitive taxels to enable localized capacitive responses. A real-to-sim-to-real pipeline estimates a 35 $\times$ 26 contact-depth map from 24-channel capacitive signals. Against simulation-generated labels, the model achieved a contact-region RMSE of 0.223 $\pm$ 0.161 mm, a weighted-centroid error of 1.213 $\pm$ 2.379 pixels, and an IoU of 0.829 $\pm$ 0.169. With measured capacitive inputs, the network-predicted depth evaluated at the guide-calibrated contact center showed a mean absolute error of 0.085 $\pm$ 0.057 mm across all 40 controlled trials, while the mean contact-position error was 0.250 $\pm$ 0.208 mm across the 37 trials whose reference contact regions were not truncated by the sensing boundary. In human-to-robot replay, tactile-guided compensation increased grasping and wiping success rates from 0% to 91.67% and 90%, respectively. In closed-loop grasping, dense-depth feedback achieved success rates of 87.5% over all tested positions and 85% under edge-contact conditions, compared with 67.5% and 45% for raw-taxel feedback.

cs.RO

Nonreciprocal Relaxation Acceleration

Driven by recent discoveries regarding the quantum Mpemba effect, the anomalous relaxation dynamics of open quantum systems have garnered significant attention. While expediting thermalization to equilibrium has been extensively studied, dynamically accelerating the convergence toward nonequilibrium steady states remains a formidable challenge. In this article, we find a transient engineered nonreciprocal dissipative channel can provide a shortcut that accelerates convergence to the target reciprocal nonequilibrium steady state for the considered two-mode model and initial states. Using interacting bosonic modes, we demonstrate that the temporal activation of a nonreciprocal channel efficiently suppresses prolonged inter-mode energy oscillations, enforcing a rapid, unidirectional thermal dump into the environment. Counterintuitively, we find that this relaxation speedup is robust and independent of the direction of the nonreciprocity. Our results provide a powerful thermodynamic technique for rapid state preparation and cooling in continuous-variable quantum systems, particularly critical for low-temperature quantum information processing.

quant-ph

SIDM and CDM interpretations of the million-solar-mass lensing perturber JVAS B1938+666-$\mathcal{V}$

A $10^6\,M_\odot$ object has recently been inferred from gravitational imaging of the strong-lensing system JVAS B1938+666, exhibiting an unusually dense inner region embedded within an extended envelope, far exceeding expectations for cold dark matter (CDM) halos. Using gravothermal fluid simulations, we show that such a structure arises naturally in self-interacting dark matter (SIDM) halos evolving into a deep core-collapse phase, where a secondary dense central core forms within an extended profile. The resulting density structure closely matches the inferred properties of the lensing object. We also demonstrate that a similar profile could be reproduced in CDM in the presence of an intermediate-mass black hole, but this requires an early-forming progenitor that subsequently loses $5$ orders of magnitude in mass through tidal stripping by the lens galaxy. Whether such a scenario can be realized in realistic cosmological environments remains an open question.

astro-ph.GA

Dirichlet-Guided Group Forecasting for Alleviating Over-smoothing in Time Series Forecasting

Time series forecasting often suffers from over-smoothing, especially when future dynamics are multi-modal. Forecasts may follow the coarse trend of the observed future, but fail to preserve sharp changes, oscillations, turning points, and regime transitions that define plausible dynamic evolution. In this work, we revisit over-smoothing from the perspective of latent dynamical mode compression: under partial observation and single-realization supervision, multiple plausible future modes can be weakened, merged, or averaged during forecasting. Based on this view, we propose Dirichlet-Guided Group Forecasting (DGF), a mode-preserving forecasting framework that explicitly models multiple mode-conditioned predictive distributions and uncertainty over their selection probabilities. DGF uses a Dirichlet-guided hierarchical sampling mechanism and reward-based optimization to encourage forecasts that are accurate, dynamically consistent, and mode-distinct. Extensive experiments on real-world forecasting benchmarks show that DGF reduces over-smoothing while improving forecasting accuracy, diversity, and dynamical consistency.

cs.LG

Hypergraph and Latent ODE Learning for Multimodal Root Cause Localization in Microservices

Root cause localization in cloud native microservice systems requires modeling complex service dependencies, irregular temporal dynamics, and heterogeneous observability data. We present HyperODE RCA, a unified framework that combines hypergraph attention learning, latent ordinary differential equations, and multimodal cross attention fusion for fine grained root cause analysis. The method learns higher order service interactions through differentiable hyperedge construction, captures continuous anomaly evolution from irregular observations with an ODE RNN encoder, and adaptively fuses logs, traces, metrics, entities, and events using context aware modality routing. We further improve robustness with a variational information bottleneck, temporal causal regularization, and invariant risk constraints. Experiments on the Tianchi AIOps benchmark show clear gains over strong baselines in ranking and classification performance, while preserving interpretability through learned hypergraph attention.

cs.LG

UniCreative: Unifying Long-form Logic and Short-form Sparkle via Reference-Free Reinforcement Learning

A fundamental challenge in creative writing lies in reconciling the inherent tension between maintaining global coherence in long-form narratives and preserving local expressiveness in short-form texts. While long-context generation necessitates explicit macroscopic planning, short-form creativity often demands spontaneous, constraint-free expression. Existing alignment paradigms, however, typically employ static reward signals and rely heavily on high-quality supervised data, which is costly and difficult to scale. To address this, we propose \textbf{UniCreative}, a unified reference-free reinforcement learning framework. We first introduce \textbf{AC-GenRM}, an adaptive constraint-aware reward model that dynamically synthesizes query-specific criteria to provide fine-grained preference judgments. Leveraging these signals, we propose \textbf{ACPO}, a policy optimization algorithm that aligns models with human preferences across both content quality and structural paradigms without supervised fine-tuning and ground-truth references. Empirical results demonstrate that AC-GenRM aligns closely with expert evaluations, while ACPO significantly enhances performance across diverse writing tasks. Crucially, our analysis reveals an emergent meta-cognitive ability: the model learns to autonomously differentiate between tasks requiring rigorous planning and those favoring direct generation, validating the effectiveness of our direct alignment approach.

cs.AI

A Primary Unified Geometric Framework of Molecular Reaction Dynamics Based on the Variational Principle

This work describes a geometric framework on molecular reaction dynamics based on the variational principle, where the Schr{\"o}dinger equation must be solved to ``see'' how a reaction occurs. First, the mathematical preliminaries are given by discussing the principle of least action and the mountain pass theorem. Second, we discuss the physical preliminaries, including the principle of equivalence for deriving the kinetic energy operator (KEO) and artificial intelligence (AI) techniques to build the potential energy surface (PES) in general spacetime. Moreover, we simplified electromagnetic interactions in curved spacetime within the molecular system and consequently, we are able to construct the nuclear Hamiltonian in nonzero curvature spacetime. This indicates possibility to introduce gauge fields through the curvature, such as additional term in the nuclear KEO near a conical intersection. Third, the single-particle approximation provides a powful ansatz to solve the Schr{\"o}dinger equation by variational principle. Thus, one can formulate the variational approaches for either electronic structure or quantum dynamics. In this work, based on previous discussions ({\it Phys. Chem. Chem. Phys.} {\bf 27} (2025), 20397) we unified them by a geometric description, where the geometric phase is naturally introduced. Finally, due to optimization characteristic of the present theory, further discussions on the present theory from optimization insight are also given, including two postulates, generative AI techniques, role of perturbation, and Markov process in optimization.

physics.chem-ph

DBMIF: a deep balanced multimodal iterative fusion framework for air- and bone-conduction speech enhancement

The performance of conventional speech enhancement systems degrades sharply in extremely low signal-to-noise ratio (SNR) environments where air-conduction (AC) microphones are overwhelmed by ambient noise. Although bone-conduction (BC) sensors offer complementary, noise-tolerant information, existing fusion approaches struggle to maintain consistent performance across a wide range of SNR conditions. To address this limitation, we propose the Deep Balanced Multimodal Iterative Fusion Framework (DBMIF), a three-branch architecture designed to reconstruct high-fidelity speech through rigorous cross-modal interaction. Specifically, grounded in a multi-scale interactive encoder-decoder backbone, the framework orchestrates an iterative attention module and a cross-branch gated module to facilitate adaptive weighting and bidirectional exchange. To complement this dynamic interaction, a balanced-interaction bottleneck is further integrated to learn a compact, stable fused representation. Extensive experiments demonstrate that DBMIF achieves competitive performance compared with recent unimodal and multimodal baselines in both speech quality and intelligibility across diverse noise types. In downstream ASR tasks, the proposed method reduces the character error rate by at least 2.5 percent compared to competing approaches. These results confirm that DBMIF effectively harnesses the robustness of BC speech while preserving the naturalness of AC speech, ensuring reliability in real-world scenarios. The source code is publicly available at github.com/wyl516w/dbmif.

eess.AS

Closing the Loop: A Control-Theoretic Framework for Provably Stable Time Series Forecasting with LLMs

Large Language Models (LLMs) have recently shown exceptional potential in time series forecasting (TSF), leveraging their inherent sequential reasoning capabilities to model complex temporal dynamics. Existing approaches typically employ an autoregressive generation strategy to adapt LLMs for TSF. However, we identify a theoretical flaw in this paradigm: during inference, the model operates in an open-loop manner, recursively consuming its own generated outputs. This leads to error accumulation, where minor early deviations cascade into significant rollout drift over long horizons. In this paper, we reformulate autoregressive forecasting through the lens of control theory, proposing Feedback-driven LLM (F-LLM), a novel closed-loop framework. Unlike standard methods that passively propagate errors, F-LLM actively stabilizes the trajectory via a learnable residual estimator functioning as a system observer. Furthermore, we provide a mathematical proof that, under explicit contraction assumptions, this closed-loop mechanism guarantees a uniformly bounded step-wise error sequence within the local surrogate dynamics. Extensive experiments demonstrate that F-LLM significantly mitigates error propagation, achieving good performance on time series benchmarks. Our code is publicly available at https://github.com/Zh-XY22/F-LLM.

cs.LG

DexTac: Learning Contact-aware Visuotactile Policies via Hand-by-hand Teaching

For contact-intensive tasks, the ability to generate policies that produce comprehensive tactile-aware motions is essential. However, existing data collection and skill learning systems for dexterous manipulation often suffer from low-dimensional tactile information. To address this limitation, we propose DexTac, a visuo-tactile manipulation learning framework based on kinesthetic teaching. DexTac captures multi-dimensional tactile data-including contact force distributions and spatial contact regions-directly from human demonstrations. By integrating these rich tactile modalities into a policy network, the resulting contact-aware agent enables a dexterous hand to autonomously select and maintain optimal contact regions during complex interactions. We evaluate our framework on a challenging unimanual injection task. Experimental results demonstrate that DexTac achieves a 91.67% success rate. Notably, in high-precision scenarios involving small-scale syringes, our approach outperforms force-only baselines by 31.67%. These results underscore that learning multi-dimensional tactile priors from human demonstrations is critical for achieving robust, human-like dexterous manipulation in contact-rich environments.

cs.RO

Purification Before Fusion: Toward Mask-Free Speech Enhancement for Robust Audio-Visual Speech Recognition

Audio-visual speech recognition (AVSR) typically improves recognition accuracy in noisy environments by integrating noise-immune visual cues with audio signals. Nevertheless, high-noise audio inputs are prone to introducing adverse interference into the feature fusion process. To mitigate this, recent AVSR methods often adopt mask-based strategies to filter audio noise during feature interaction and fusion, yet such methods risk discarding semantically relevant information alongside noise. In this work, we propose an end-to-end noise-robust AVSR framework coupled with speech enhancement, eliminating the need for explicit noise mask generation. This framework leverages a Conformer-based bottleneck fusion module to implicitly refine noisy audio features with video assistance. By reducing modality redundancy and enhancing inter-modal interactions, our method preserves speech semantic integrity to achieve robust recognition performance. Experimental evaluations on the public LRS3 benchmark suggest that our method outperforms prior advanced mask-based baselines under noisy conditions.

eess.AS

V-CORE: Temporally Consistent Video Understanding for Video-LLM

Recent Video Large Language Models (Video-LLMs) have shown strong multimodal reasoning capabilities, yet remain challenged by video understanding tasks that require consistent temporal ordering and causal coherence. Many parameter-efficient Video-LLMs rely on unconstrained bidirectional projectors to model inter-frame interactions, which can blur temporal ordering by allowing later frames to influence earlier representations, without explicit architectural mechanisms to respect the directional nature of video reasoning. To address this limitation, we propose V-CORE, a parameter-efficient framework that introduces explicit temporal ordering constraints for video understanding. V-CORE consists of two key components: (1) Learnable Spatial Aggregation (LSA), which adaptively selects salient spatial tokens to reduce redundancy, and (2) a Causality-Aware Temporal Projector (CATP), which enforces structured unidirectional information flow via block-causal attention and a terminal dynamic summary token acting as a causal sink. This design preserves intra-frame spatial interactions while ensuring that temporal information is aggregated in a strictly ordered manner. With 4-bit QLoRA and a frozen LLM backbone, V-CORE can be trained efficiently on a single consumer GPU. Experiments show that V-CORE achieves strong performance on the challenging NExT-QA benchmark, reaching 61.2% accuracy, and remains competitive across MSVD-QA, MSRVTT-QA, and TGIF-QA, with gains concentrated in temporal and causal reasoning subcategories (+3.5% and +5.2% respectively), directly validating the importance of explicit temporal ordering constraints.

cs.CV

Cross-platform Product Matching Based on Entity Alignment of Knowledge Graph with RAEA model

Product matching aims to identify identical or similar products sold on different platforms. By building knowledge graphs (KGs), the product matching problem can be converted to the Entity Alignment (EA) task, which aims to discover the equivalent entities from diverse KGs. The existing EA methods inadequately utilize both attribute triples and relation triples simultaneously, especially the interactions between them. This paper introduces a two-stage pipeline consisting of rough filter and fine filter to match products from eBay and Amazon. For fine filtering, a new framework for Entity Alignment, Relation-aware and Attribute-aware Graph Attention Networks for Entity Alignment (RAEA), is employed. RAEA focuses on the interactions between attribute triples and relation triples, where the entity representation aggregates the alignment signals from attributes and relations with Attribute-aware Entity Encoder and Relation-aware Graph Attention Networks. The experimental results indicate that the RAEA model achieves significant improvements over 12 baselines on EA task in the cross-lingual dataset DBP15K (6.59% on average Hits@1) and delivers competitive results in the monolingual dataset DWY100K. The source code for experiments on DBP15K and DWY100K is available at github (https://github.com/Mockingjay-liu/RAEA-model-for-Entity-Alignment).

cs.AI