SearcharxivSearch

arXiv subjects

Xinyang Wang

Publications and source records attributed to Xinyang Wang.

At least 19 recordsLinked to original sources

A Frequency-Aware Dynamic Knowledge Distillation Framework: An Effective Tool for Bridging Low- and High-Frequency Seismic Information

Seismic data contain rich information across different frequency bands, with low-frequency components primarily characterizing large-scale geological structures and high-frequency components preserving fine-scale seismic details. Effectively integrating these frequency-dependent components is essential for seismic feature learning to better preserve structural continuity and fine-scale details. Knowledge distillation provides an effective means for transferring informative representations from high-quality data. However, existing distillation-based frameworks usually treat seismic features in a full-band manner, ignoring relationships across frequency bands and thereby limiting the coordinated transfer of low- and high-frequency knowledge. To bridge low- and high-frequency seismic features through knowledge distillation, we propose a frequency-aware dynamic knowledge distillation framework (FADKD-Net), which establishes a teacher-student learning framework and performs frequency-aware knowledge transfer between low- and high-frequency bands. Specifically, FADKD-Net decomposes seismic features into low- and high-frequency components and performs targeted distillation to exploit their complementary information. Low-frequency distillation guides the student model to learn stable structural priors, thereby improving the overall continuity of seismic events. Meanwhile, high-frequency distillation enhances detailed feature modeling and improves the representational capability for complex and small-scale structures. Furthermore, a cross-domain feature alignment strategy is proposed to reduce distributional discrepancies across different surveys and enhance the transferability of the seismic representations learned by FADKD-Net.

physics.geo-ph

The role of thermal photons in a magnetized plasma and their significance in heavy ion collisions

In this contribution, we present thermal photon production mechanisms within a magnetized quark-gluon plasma, utilizing the framework of Landau-level quantization. We examine the specific influences of the magnetic field, chemical potential, and chiral chemical potential on photon yield and polarization. By providing a more comprehensive theoretical description of photon production across the full evolution of heavy-ion collisions, our study offers a promising new avenue for resolving the photon $v_2$ puzzle and potentially detecting the signature of magnetic fields in heavy-ion collision experiments.

hep-ph

Thermal Dilepton Polarization under Rotation or Magnetic Field in Heavy-ion Collisions

Dilepton (Virtual photon) polarization is characterized by anisotropic coefficients $\lambda_{\theta}$, $\lambda_{\phi}$, and $\lambda_{\theta\phi}$, which are expected to be influenced by vorticity and magnetic fields. This work investigates thermal dilepton production in a quark-gluon plasma via the quark-antiquark annihilation process $q\bar{q} \to \gamma^* \to l^+l^-$. Virtual photon polarization can be induced by both the spin polarization of quarks and the anisotropy of their momentum distribution in the medium. By employing the modified quark propagator under an external field, we derive the electromagnetic spectral function in a hot medium. Based on the spin-projection decomposition of the spectral function, the spin density matrix elements of the virtual photon and the anisotropy coefficients for the emitted dileptons are determined. Due to the distinct effects of vorticity and magnetic fields on the quark propagator, the resulting invariant mass spectra of dilepton polarization exhibit characteristic differences. Furthermore, our study reveals the response of dilepton polarization signals to external fields of varying strengths, suggesting dilepton polarization as a complementary and sensitive probe for both vorticity and magnetic fields in relativistic heavy-ion collisions.

hep-ph

Native Video-Action Pretraining for Generalizable Robot Control

The advent of video-action models offers a promising path for robot control. Nevertheless, we argue that repurposing video generative models designed for digital content creation is inherently inadequate for physical environments. To bridge this gap, we present LingBot-VA 2.0, a video-action foundation model built from the ground up for embodiment. Four core design principles showcase its evolution from LingBot-VA. (1) Departing from traditional reconstruction-focused VAEs, we introduce a semantic visual-action tokenizer, which aligns visual representations with both semantics and actions, improving instruction following and action precision in subsequent policy learning. (2) Given the strictly causal nature of temporal dynamics, we adopt a causal pretraining paradigm, training from scratch to circumvent the catastrophic forgetting that frequently occurs when adapting bidirectional architectures. (3) To meet the demands of high-frequency inference, our model employs a sparse MoE backbone, expanding model capacity without compromising efficiency. (4) Real-time closed-loop control is realized through an enhanced asynchronous inference scheme, which predicts future latents in parallel with action execution while re-grounding each rollout on the latest observation via learned forward dynamics. Real-world deployment validates LingBot-VA 2.0 as a robust foundation model, as evidenced by its few-shot generalization across complex manipulation tasks.

cs.RO

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence

Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their primary focus on content creation. For example, their design inherently prioritizes visual fidelity and creativity over computational efficiency and physical realism. In this work, we present LingBot-Video, a DiT-based video pretraining paradigm specifically tailored for embodied intelligence. From the architecture perspective, we adopt the Mixture-of-Experts (MoE), instead of dense, framework to achieve a better trade-off between modeling capacity and inference efficiency, and manage to scale it up from scratch. From the data perspective, we construct a data profiling engine that augments standard internet videos with extensive robot-oriented footage, encompassing manipulation, navigation, and egocentric perspectives, to equip the base model with an intrinsic understanding of actions and world dynamics. From the training perspective, we develop a multi-dimensional reward system to enforce the alignment regarding physical rationality and task completion, going beyond standard criteria such as aesthetics, prompt-following, and motion consistency. Comprehensive evaluations validate its performance and efficiency as a video foundation model. We contribute LingBot-Video as the inaugural large-scale, open-source MoE video foundation model to the community, in a pioneering effort to bridge digital creativity and physical actuation.

cs.CV

From Foundation to Application: Improving VLA Models in Practice

Despite recent progress of VLA foundation models, the disparity between laboratory conditions and real-world applications continues to impede their practical implementation. To bridge this gap, we present LingBot-VLA 2.0, which advances LingBot-VLA through improvements in three functional domains. (1) Generalization across tasks and embodiments. Compared to the previous version, we revamp the data processing pipeline and curate around 60,000 hours of data for pretraining, including 50,000 hours of robot trajectories spanning 20 robot configurations and 10,000 hours of egocentric human videos. (2) Expanded action space in addition to dual-arm hardware platforms. In particular, our system accommodates degrees of freedom for the heads, waists, mobile bases, and dexterous hands, thereby empowering the robots to tackle more complex tasks in practical scenarios. (3) Predictive dynamics modeling for improved temporal reasoning. Specifically, we formulate future prediction as a proxy task, facilitated by a video representation model for semantic priors and a depth estimation model for geometric cues. Evaluations on the GM-100 benchmark, conducted in a generalist setting, validate the beneficial impact of these proposed modifications. Furthermore, benefiting from the expanded pretraining data that covers whole-body degrees of freedom, LingBot-VLA-2.0 demonstrates strong cross-embodiment long-horizon mobile manipulation capability across the two robotic platforms.

cs.RO

CFRNet: Cycle-Consistent Fixed-Point Training for Real-Time Blind Face Restoration on Consumer Embedded NPUs

Blind face restoration on consumer devices has to balance image quality against speed and memory. Strong methods such as GFPGAN and CodeFormer give good perceptual quality, but they rely on large pretrained generative priors and on operators such as attention, codebook lookup, and style modulation that are hard to compile and quantize on the small neural processing units (NPUs) used in consumer hardware. Small convolutional restorers run fast enough, but they tend to over-smooth and to leave artifacts around the eyes, nose, and mouth. We present CFRNet, a 2.0,M-parameter ResNet-style restorer for on-device use at $256\times256$, the common face-crop size on consumer NPUs. The main idea is Cycle-Consistent Fixed-Point Training (CCFP). Instead of training the network for one pass and then running it several times by hand, we train it to act as a fixed-point operator, so that applying it again to a restored face does not change the face. CCFP uses three training losses, namely progressive multi-cycle supervision, an idempotence loss, and a re-degradation cycle loss, and it adds no cost at inference. To compare fairly under our deployment limits, we retrain all baselines from scratch at the same $256\times256$ resolution. On a 300-image test set, CFRNet reaches the best perceptual score (LPIPS 0.250 at three cycles, which is 31% lower than one cycle) and also the best PSNR and SSIM at two cycles. It runs in about 23,ms per cycle in INT8 on a HiSilicon Hi3402 NPU, while the same baselines cannot be compiled to that chip. The cycle count $k$ acts as a simple quality knob that needs no retraining: PSNR is best at $k\!=\!2$ and LPIPS keeps improving up to $k\!=\!3$. We further show that the same idea works with a plain CNN that is even easier to deploy, and we run the model in real time on an in-car driver-monitoring board.

cs.CV

MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding

Multimodal Large Language Models (MLLMs) have demonstrated significant achievements in general visual question answering (VQA) tasks. However, they remain brittle on mechanical engineering drawings, where high annotation density and weak domain knowledge, compounded by unreliable spatial relation reasoning under strict projection rules and geometric constraints, make decisive cues easy to miss and frequently lead to wrong answers. To bridge this gap, we introduce the first comprehensive mechanical drawing understanding dataset, MechVQA, created through a semi-automated construction and quality-control pipeline. MechVQA contains 3.3k high-density pictures with 21K question-answer pairs, spanning 10 different fine-grained tasks across three capability levels: Recognition, Reasoning, and Judging, providing a testbed to evaluate and improve MLLM understanding on real-world mechanical drawings. On top of MechVQA, we then develop the MechVL model through a multi-stage training paradigm, building a strong domain-specialized baseline. Extensive experimental results demonstrate that MechVL outperforms the strongest closed-source baseline by 7.57 percentage points on the MechVQA total score, significantly enhancing mechanical drawing understanding ability and providing a reusable foundation for deploying MLLMs in mechanical design and inspection scenarios.

cs.CV

ASTRA-QA: A Benchmark for Abstract Question Answering over Documents

Document-based question answering (QA) increasingly includes abstract questions that require synthesizing scattered information from long documents or across multiple documents into coherent answers. However, this setting is still poorly supported by existing benchmarks and evaluation methods, which often lack stable abstract references or rely on coarse similarity metrics and unstable head-to-head comparisons. To alleviate this issue, we introduce ASTRA-QA, a benchmark for AbSTRAct Question Answering over documents. ASTRA-QA contains 869 QA instances over academic papers and news documents, covering five abstract question types and three controlled retrieval scopes. Each instance is equipped with explicit evaluation annotations, including answer topic sets, curated unsupported topics, and aligned evidence. Building on these annotations, ASTRA-QA assesses whether answers cover required key points and avoid unsupported content by directly scoring topic coverage and curated unsupported content, enabling scalable evaluation without exhaustive head-to-head comparisons. Experiments with representative Retrieval-Augmented Generation (RAG) methods spanning vanilla, graph-based, and hierarchical retrieval settings show that ASTRA-QA provides reference-grounded diagnostics for coverage, hallucination, and retrieval-scope robustness. Our dataset and code are available at https://xinyangsally.github.io/astra-benchmark.

cs.CL

RG-Consistent (P)NJL Model: Impact of Thermal Cutoff Modifications on Thermodynamics and Net-Baryon Number Fluctuations

In this paper, we investigate the impact of renormalization group (RG) consistency on the chiral phase transition and thermodynamic properties of QCD matter using the RGNJL and RGPNJL models. By implementing a temperature-dependent thermal cutoff $\Lambda_T = k\Lambda_0$, we ensure that thermodynamic quantities converge toward the Stefan-Boltzmann limit at high temperatures, effectively extending the applicability of these effective theories. Our analysis shows that while the RG-consistency condition ($k \rightarrow \infty$) resolves causality violations in the RGNJL model by binding the speed of sound to the conformal limit, the RGPNJL model exhibits a more complex, non-monotonic sensitivity to the parameter $k$. Furthermore, we demonstrate that the RG-improved PNJL framework significantly enhances the description of net-baryon number fluctuations ($\kappa\sigma^2$) relative to lattice QCD data at vanishing chemical potential, though the intensification of these fluctuations at high baryon density highlights a critical sensitivity to the model's parametric constraints. This study provides a rigorous evaluation of the RG-consistency framework's predictive power in mapping the QCD phase diagram and interpreting experimental observables.

hep-ph

Equilibrium Stability and Uniqueness with a Large Number of Commodities and Patient Consumers

We show that a large effective number of commodities can be a source of equilibrium stability and uniqueness: expanding substitution opportunities strengthens aggregate substitution effects. We study finite dated-commodity exchange economies obtained by truncating a countably infinite-horizon environment with discounted, additively separable utilities. In this setting, the effective number of commodities is the discounted count of dated commodities, so greater patience makes distant commodities more relevant. With an appropriate normalization, equilibrium substitution effects accumulate at the rate of the effective number of commodities. When a preference diversification condition holds, equilibrium income effects grow at a lower rate. The condition is satisfied, for example, when agents have sparse or localized taste differences across commodities, or when their taste profiles become sufficiently heterogeneous as the commodity space expands. Hence, whenever the effective number of commodities is sufficiently large, every equilibrium is locally t\^atonnement stable, which in turn implies equilibrium uniqueness.

econ.TH

Chiral Magnetic effect as the anomaly in the transverse axial vector Ward Identity

Through analyzing the quark propagator under the magnetic field, we establish that the axial anomaly originates from an additional Dirac structure in quark propagator induced by the magnetic field. This Dirac structure also allows one to connect the axial anomaly with the topological properties of the system by checking the axial vector Ward identity. For the tree level propagator, we reproduce the result of the anomalous axial current as in the Dirac Hamiltonian approach and kinetic theory. Particularly, we confirm that the chiral magnetic effect (CME) comes from the same term that is in charge of the axial anomaly, specifically, as the anomaly of the transversal axial vector Ward Identity. The identity guarantees that the CME conductivity $C_{\rm CME}$ is a constant as $C_{\rm CME}=\frac{1}{2\pi^2}$, and is robust against the temperature, chemical potential, magnetic field and also interaction. Finally, we verify this numerically by applying the full quark propagator under magnetic field calculated from the functional QCD methods.

hep-th

HaS: Accelerating RAG through Homology-Aware Speculative Retrieval

Retrieval-Augmented Generation (RAG) expands the knowledge boundary of large language models (LLMs) at inference by retrieving external documents as context. However, retrieval becomes increasingly time-consuming as the knowledge databases grow in size. Existing acceleration strategies either compromise accuracy through approximate retrieval, or achieve marginal gains by reusing results of strictly identical queries. We propose HaS, a homology-aware speculative retrieval framework that performs low-latency speculative retrieval over restricted scopes to obtain candidate documents, followed by validating whether they contain the required knowledge. The validation, grounded in the homology relation between queries, is formulated as a homologous query re-identification task: once a previously observed query is identified as a homologous re-encounter of the incoming query, the draft is deemed acceptable, allowing the system to bypass slow full-database retrieval. Benefiting from the prevalence of homologous queries under real-world popularity patterns, HaS achieves substantial efficiency gains. Extensive experiments demonstrate that HaS reduces retrieval latency by 23.74% and 36.99% across datasets with only a 1-2% marginal accuracy drop. As a plug-and-play solution, HaS also significantly accelerates complex multi-hop queries in modern agentic RAG pipelines. Source code is available at: https://github.com/ErrEqualsNil/HaS.

cs.IR

Spectral structure of the Benjamin-Feir instability in deep-water gravity-capillary Stokes waves

We investigate the Benjamin-Feir instability of small-amplitude gravity-capillary Stokes waves in deep water for the full water wave equations. While modulational instability has been classically predicted by formal asymptotic approaches, such as nonlinear Schr\"odinger approximations, a complete spectral description at the level of the Euler equations has remained open. We perform a rigorous Bloch-Floquet spectral analysis of the linearized operator and describe the splitting of the multiple eigenvalues at the origin. In the unstable regime, we identify a pair of eigenvalues with non-zero real part forming the characteristic ``figure-eight'' pattern in the complex plane. As a consequence, we recover sharp instability and stability regions in terms of the surface tension parameter, thereby providing a fully rigorous justification of the classical predictions in the gravity-capillary setting.

math.AP

Determining the NJL Coupling and AMM in Magnetized QCD Matter via Machine Learning

In this study, we investigate the phase structure of magnetized QCD matter by determining the field-dependent parameters of the Nambu-Jona-Lasinio (NJL) model through a physics-informed machine learning framework. Specifically, we focus on extracting the optimal functional forms for the running coupling constant $G(eB)$ and the quark anomalous magnetic moment (AMM) ratio $v_2(eB)$, utilizing lattice QCD-computed quark condensate data as the ``ground truth". By embedding the NJL gap equation as a differentiable physics-constrained module, our neural network pipeline identifies continuous parameter functions that accurately reproduce the inverse magnetic catalysis (IMC) effect. Our results demonstrate that the magnetic field smoothly suppresses both $G$ and $v_2$. This approach not only bridges the gap between effective models and lattice data but also provides new microscopic insights into the response of the QCD vacuum to strong magnetic fields.

hep-ph

Transcending Classical Neural Network Boundaries: A Quantum-Classical Synergistic Paradigm for Seismic Data Processing

In recent years, a number of neural-network (NN) methods have exhibited good performance in seismic data processing, such as denoising, interpolation, and frequency-band extension. However, these methods rely on stacked perceptrons and standard activation functions, which imposes a bottleneck on the representational capacity of deep-learning models, making it difficult to capture the complex and non-stationary dynamics of seismic wavefields. Different from the classical perceptron-stacked NNs which are fundamentally confined to real-valued Euclidean spaces, the quantum NNs leverage the exponential state space of quantum mechanics to map the features into high-dimensional Hilbert spaces, transcending the representational boundary of classical NNs. Based on this insight, we propose a quantum-classical synergistic generative adversarial network (QC-GAN) for seismic data processing, serving as the first application of quantum NNs in seismic exploration. In QC-GAN, a quantum pathway is used to exploit the high-order feature correlations, while the convolutional pathway specializes in extracting the waveform structures of seismic wavefields. Furthermore, we design a QC feature complementarity loss to enforce the feature orthogonality in the proposed QC-GAN. This novel loss function can ensure that the two pathways encode non-overlapping information to enrich the capacity of feature representation. On the whole, by synergistically integrating the quantum and convolutional pathways, the proposed QC-GAN breaks the representational bottleneck inherent in classical GAN. Experimental results on denoising and interpolation tasks demonstrate that QC-GAN preserves wavefield continuity and amplitude-phase information under complex noise conditions.

cs.LG

Unifying Language-Action Understanding and Generation for Autonomous Driving

Vision-Language-Action (VLA) models are emerging as a promising paradigm for end-to-end autonomous driving, valued for their potential to leverage world knowledge and reason about complex driving scenes. However, existing methods suffer from two critical limitations: a persistent misalignment between language instructions and action outputs, and the inherent inefficiency of typical auto-regressive action generation. In this paper, we introduce LinkVLA, a novel architecture that directly addresses these challenges to enhance both alignment and efficiency. First, we establish a structural link by unifying language and action tokens into a shared discrete codebook, processed within a single multi-modal model. This structurally enforces cross-modal consistency from the ground up. Second, to create a deep semantic link, we introduce an auxiliary action understanding objective that trains the model to generate descriptive captions from trajectories, fostering a bidirectional language-action mapping. Finally, we replace the slow, step-by-step generation with a two-step coarse-to-fine generation method C2F that efficiently decodes the action sequence, saving 86% inference time. Experiments on closed-loop driving benchmarks show consistent gains in instruction following accuracy and driving performance, alongside reduced inference latency.

cs.CV

A Consistency-Improved LiDAR-Inertial Bundle Adjustment

Simultaneous Localization and Mapping (SLAM) using 3D LiDAR has emerged as a cornerstone for autonomous navigation in robotics. While feature-based SLAM systems have achieved impressive results by leveraging edge and planar structures, they often suffer from the inconsistent estimator associated with feature parameterization and estimated covariance. In this work, we present a consistency-improved LiDAR-inertial bundle adjustment (BA) with tailored parameterization and estimator. First, we propose a stereographic-projection representation parameterizing the planar and edge features, and conduct a comprehensive observability analysis to support its integrability with consistent estimator. Second, we implement a LiDAR-inertial BA with Maximum a Posteriori (MAP) formulation and First-Estimate Jacobians (FEJ) to preserve the accurate estimated covariance and observability properties of the system. Last, we apply our proposed BA method to a LiDAR-inertial odometry.

cs.RO