SearcharxivSearch

arXiv subjects

Yang Zhong

Publications and source records attributed to Yang Zhong.

At least 19 recordsLinked to original sources

First-Principles Electron-Magnon Coupling with Machine-Learning Hamiltonians: From Band Renormalization to Transport

In analogy to electron-phonon coupling (EPC), electron-magnon coupling (EMC) is expected to shape electronic structure, transport, and possibly unconventional superconductivity in magnetic materials. However, unlike EPC, which is now routinely treated within first-principles frameworks, a quantitative description of EMC, especially for transport, remains elusive because of the lack of theoretical formalism. Consequently, even for elemental iron, EPC-only calculations miss both the magnitude and the $T^2$ component of resistivity. This discrepancy has long been attributed to EMC, although direct computational evidence has been lacking and the underlying transport mechanism remains unresolved. Here we develop a unified first-principles formalism for EMC in collinear magnetic systems within many-body perturbation theory, complemented by machine-learning spinful Hamiltonians that supply quantities not directly accessible from conventional first-principles methods. Our framework enables ab initio transport calculations including EMC effects for the first time. Applied to ferromagnetic $\alpha$-Fe, our approach yields electron spectral functions consistent with previous studies. More importantly, we recover the full $T^2$ component of resistivity with a coefficient in quantitative agreement with measurement and reveal that the $T^2$ component cannot be attributed solely to EMC, as has long been assumed, but is dominated by the strong EPC-EMC interplay. Extending to antiferromagnetic K-doped $\mathrm{BaMn_2As_2}$, our method captures the ARPES-observed magnon-induced kink and a large EMC strength of $\sim 3$ comparable to experimental measurements, demonstrating the generality of the framework. Our work closes a longstanding gap in the quantitative understanding of transport in magnetic systems and provides a predictive foundation for examining magnon-mediated phenomena.

physics.comp-ph

VICBench: A Multi-Language Benchmark for Code Vulnerability Detection

Evaluating security vulnerability detection tools requires benchmark datasets with vulnerability-inducing commits (VICs) - the commits that first introduce vulnerabilities into codebases. VICs are essential for determining the full range of vulnerable software versions. Existing vulnerability datasets suffer from limited programming language coverage, restricted patch complexity, and narrow project scope. Through our dual annotation by human experts and an agentic workflow, we create a benchmark - VICBench - of 100 verified VICs for 100 CVEs across 88 projects in Python, Java, and C++, covering 48 CWE types. VICBench features complex real-world vulnerability fixes averaging 38.6 lines and corresponding VICs of 252.5 lines - significantly larger than prior work. Our evaluation shows that state-of-the-art algorithms V-SZZ and LLM4SZZ achieve only 33.3%-40.1% F1, confirming that using existing approaches still entails significant manual effort. VICBench enables robust evaluation of vulnerability detection approaches.

cs.CR

Nonadiabatic Molecular Dynamics on Real-time Excited-State Surfaces via Machine Learning Hamiltonians

Simulating the coupled, nonequilibrium dynamics of electrons and nuclei is a central challenge in chemistry, physics, and materials science, governing phenomena from photocatalysis to quantum information. The primary bottleneck has been the lack of a general, accurate, and efficient method for modeling the complete excited-state landscape: the potential energy surfaces, forces, and non-adiabatic couplings for multiple electronic states. While machine learning has revolutionized ground-state simulations and shown promise for excited states in molecules, a unified framework that solves the complete multi-state problem for general condensed matter systems has remained elusive. Here we introduce on-the-fly N${^2}$AMD (Neural network NAMD), a machine learning framework that makes on-the-fly NAMD in solids a reality. By employing an equivariant neural network to predict the system Hamiltonian, the framework delivers excited-state energies, forces, and non-adiabatic coupling vectors at a fraction of the cost of ab initio calculations. Crucially, it allows simulations with hybrid functional accuracy, a level of approach previously inaccessible for NAMD. We showcase its capabilities with three topical examples: correcting order-of-magnitude errors in carrier dynamics predicted by conventional procedure in a MoS$_2$/WS$_2$ heterostructure, simulating previously inaccessible photoinduced ferroelectric switching, and capturing real-time polaron formation in TiO$_2$ at the hybrid-functional level. On-the-fly N${^2}$AMD moves beyond the limitations of equilibrium theory, establishing a new paradigm for the predictive, first-principles design of materials operating far from equilibrium.

physics.comp-ph

Causality and Stability of First-Order Relativistic Spin Hydrodynamics with Conserved Charges

We study the causality and stability of first-order relativistic spin hydrodynamics with particle-number conservation. By deriving the complete dispersion relations of linear perturbations around global equilibrium, we find that conserved-charge dynamics modifies the sound sector and introduces additional non-hydrodynamic modes absent in the charge-neutral theory. While the structure of spin relaxation modes remains unchanged, the stability conditions acquire new contributions from charge diffusion and thermodynamic susceptibilities. More importantly, a particle-number-induced mode is shown to violate the causality condition in the short-wavelength limit. We further demonstrate that particle-number conservation does not remove the instability inherent in first-order spin hydrodynamics. These results reveal nontrivial interplay between spin and conserved-charge dynamics and provide important constraints on relativistic spin hydrodynamic theories at finite density.

nucl-th

XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments

Vision-Language-Action (VLA) models drive next-generation autonomous systems, but training them requires scalable, high-quality annotations from complex environments. Current cloud pipelines rely on generic vision-language models (VLMs) that lack geometric reasoning and domain semantics due to their 2D image-text pretraining. To address this mismatch, we propose XEmbodied, a cloud-side foundation model that endows VLMs with intrinsic 3D geometric awareness and interaction with physical cues (e.g., occupancy grids, 3D boxes). Instead of treating geometry as auxiliary input, XEmbodied integrates geometric representations via a structured 3D Adapter and distills physical signals into context tokens using an Efficient Image-Embodied Adapter. Through progressive domain curriculum and reinforcement learning post-training, XEmbodied preserves general capabilities while demonstrating robust performance across 18 public benchmarks. It significantly improves spatial reasoning, traffic semantics, embodied affordance, and out-of-distribution generalization for large-scale scenario mining and embodied VQA.

cs.CV

Causality and stability analysis of relativistic spin hydrodynamics: Insights from a nonvanishing spin density background

We investigate the stability and causality of relativistic spin hydrodynamics in the presence of a nonvanishing spin density background, assuming that the spin chemical potential $\omega^{\mu\nu}$ is of leading order ($\omega^{\mu\nu} \sim \mathcal{O}(1)$) in the gradient expansion and is treated as a finite background in the linear perturbation analysis. It is found that within the first-order spin hydrodynamic framework, a finite spin density background modifies the dispersion relations, and modes propagating along different directions are controlled by distinct transport coefficients. Certain specific modes only appear in the $x$ direction. However, the modes in the large wave-vector limit exhibit acausal behavior. To address this issue, we subsequently adopt the framework of minimal causal spin hydrodynamics and derive the corresponding stability and causality conditions. The spin density background directly determines whether stability and causality can be satisfied simultaneously. In the small wave-vector limit, the results are similar to those of the first-order theory. In the large wave-vector region, however, significant differences emerge: the distinctions between different directions are no longer merely simple substitutions of transport coefficients, but involve more complex combinations. This indicates that the difference between modes in different directions increases with increasing wave vector.

hep-ph

Physics-Informed Long-Range Coulomb Correction for Machine-learning Hamiltonians

Machine-learning electronic Hamiltonians achieve orders-of-magnitude speedups over density-functional theory, yet current models omit long-range Coulomb interactions that govern physics in polar crystals and heterostructures. We derive closed-form long-range Hamiltonian matrix elements in a nonorthogonal atomic-orbital basis through variational decomposition of the electrostatic energy, deriving a variationally consistent mapping from the electron density matrix to effective atomic charges. We implement this framework in HamGNN-LR, a dual-channel architecture combining E(3)-equivariant message passing with reciprocal-space Ewald summation. Benchmarks demonstrate that physics-based long-range corrections are essential: purely data-driven attention mechanisms fail to capture macroscopic electrostatic potentials. Benchmarks on polar ZnO slabs, CdSe/ZnS heterostructures, and GaN/AlN superlattices show two- to threefold error reductions and robust transferability to systems far beyond training sizes, eliminating the characteristic staircase artifacts that plague short-range models in the presence of built-in electric fields.

physics.comp-ph

3rd Place Solution to ICCV LargeFineFoodAI Retrieval

This paper introduces the 3rd place solution to the ICCV LargeFineFoodAI Retrieval Competition on Kaggle. Four basic models are independently trained with the weighted sum of ArcFace and Circle loss, then TTA and Ensemble are successively applied to improve feature representation ability. In addition, a new reranking method for retrieval is proposed based on diffusion and k-reciprocal reranking. Finally, our method scored 0.81219 and 0.81191 mAP@100 on the public and private leaderboard, respectively.

cs.CV

3rd Place Solution to Large-scale Fine-grained Food Recognition

Food analysis is becoming a hot topic in health area, in which fine-grained food recognition task plays an important role. In this paper, we describe the details of our solution to the LargeFineFoodAI-ICCV Workshop-Recognition challenge held on Kaggle. We find a proper combination of Arcface loss[1] and Circle loss[9] can bring improvement to the performance. With Arcface and the combined loss, model was trained with carefully tuned configurations and ensembled to get the final results. Our solution won the 3rd place in the competition.

cs.CV

Efficient E(3)-equivariant framework for universal charge density prediction

Electronic structure is ubiquitously obtained via density functional theory (DFT), where the charge density plays a central role. This work presents EdenGNN (Equivariant Density Graph Neural Network), a machine learning (ML) charge density model for electronic structure. Current universal ML charge density models are hampered by prohibitive computational costs. Furthermore, despite being trained on projector augmented-wave (PAW) based DFT datasets, they predict only the pseudo charge density, which is insufficient to reconstruct the electronic structure. In contrast, EdenGNN overcomes these limitations. It additionally predicts the augmentation occupancies, enabling electronic structure calculations with PAW accuracy. Critically, by employing a basis-expansion formulation with fully trainable radial basis functions and a $\Delta$-learning strategy to capture charge transfer, it is over an order of magnitude faster. Trained on the Materials Project database, our universal model, EdenGNN-Uni, accurately predicts the band structures for the majority of materials across a vast chemical space. These findings establish the ML charge density model as a scalable \textit{ab initio} method for large-scale electronic structure calculations and high-throughput screening.

cond-mat.mtrl-sci

TurboFuzz: FPGA Accelerated Hardware Fuzzing for Processor Agile Verification

Verification is a critical process for ensuring the correctness of modern processors. The increasing complexity of processor designs and the emergence of new instruction set architectures (ISAs) like RISC-V have created demands for more agile and efficient verification methodologies, particularly regarding verification efficiency and faster coverage convergence. While simulation-based approaches now attempt to incorporate advanced software testing techniques such as fuzzing to improve coverage, they face significant limitations when applied to processor verification, notably poor performance and inadequate test case quality. Hardware-accelerated solutions using FPGA or ASIC platforms have tried to address these issues, yet they struggle with challenges including host-FPGA communication overhead, inefficient test pattern generation, and suboptimal implementation of the entire multi-step verification process. In this paper, we present TurboFuzz, an end-to-end hardware-accelerated verification framework that implements the entire Test Generation-Simulation-Coverage Feedback loop on a single FPGA for modern processor verification. TurboFuzz enhances test quality through optimized test case (seed) control flow, efficient inter-seed scheduling, and hybrid fuzzer integration, thereby improving coverage and execution efficiency. Additionally, it employs a feedback-driven generation mechanism to accelerate coverage convergence. Experimental results show that TurboFuzz achieves up to 2.23x more coverage collection than software-based fuzzers within the same time budget, and up to 571x performance speedup when detecting real-world issues, while maintaining full visibility and debugging capabilities with moderate area overhead.

cs.AR

EvaDrive: Evolutionary Adversarial Policy Optimization for End-to-End Autonomous Driving

Autonomous driving faces significant challenges in achieving human-like iterative decision-making, which continuously generates, evaluates, and refines trajectory proposals. Current generation-evaluation frameworks isolate trajectory generation from quality assessment, preventing iterative refinement essential for planning, while reinforcement learning methods collapse multi-dimensional preferences into scalar rewards, obscuring critical trade-offs and yielding scalarization bias.To overcome these issues, we present EvaDrive, a novel multi-objective reinforcement learning framework that establishes genuine closed-loop co-evolution between trajectory generation and evaluation via adversarial optimization. EvaDrive frames trajectory planning as a multi-round adversarial game. In this game, a hierarchical generator continuously proposes candidate paths by combining autoregressive intent modeling for temporal causality with diffusion-based refinement for spatial flexibility. These proposals are then rigorously assessed by a trainable multi-objective critic that explicitly preserves diverse preference structures without collapsing them into a single scalarization bias.This adversarial interplay, guided by a Pareto frontier selection mechanism, enables iterative multi-round refinement, effectively escaping local optima while preserving trajectory diversity.Extensive experiments on NAVSIM and Bench2Drive benchmarks demonstrate SOTA performance, achieving 94.9 PDMS on NAVSIM v1 (surpassing DiffusionDrive by 6.8, DriveSuprim by 5.0, and TrajHF by 0.9) and 64.96 Driving Score on Bench2Drive. EvaDrive generates diverse driving styles via dynamic weighting without external preference data, introducing a closed-loop adversarial framework for human-like iterative decision-making, offering a novel scalarization-free trajectory optimization approach.

cs.LG

A Survey on Vision-Language-Action Models for Autonomous Driving

The rapid progress of multimodal large language models (MLLM) has paved the way for Vision-Language-Action (VLA) paradigms, which integrate visual perception, natural language understanding, and control within a single policy. Researchers in autonomous driving are actively adapting these methods to the vehicle domain. Such models promise autonomous vehicles that can interpret high-level instructions, reason about complex traffic scenes, and make their own decisions. However, the literature remains fragmented and is rapidly expanding. This survey offers the first comprehensive overview of VLA for Autonomous Driving (VLA4AD). We (i) formalize the architectural building blocks shared across recent work, (ii) trace the evolution from early explainer to reasoning-centric VLA models, and (iii) compare over 20 representative models according to VLA's progress in the autonomous driving domain. We also consolidate existing datasets and benchmarks, highlighting protocols that jointly measure driving safety, accuracy, and explanation quality. Finally, we detail open challenges - robustness, real-time efficiency, and formal verification - and outline future directions of VLA4AD. This survey provides a concise yet complete reference for advancing interpretable socially aligned autonomous vehicles. Github repo is available at \href{https://github.com/JohnsonJiang1996/Awesome-VLA4AD}{SicongJiang/Awesome-VLA4AD}.

cs.CV

Unsupervised deep learning model for fast energy layer pre-selection of delivery-efficient proton arc therapy plan optimization of nasopharyngeal carcinoma

Proton arc therapy (PAT) is an emerging and promising modality in radiotherapy, offering improved dose distribution and treatment robustness over intensity-modulated proton therapy. Yet, identifying the optimal energy layer (EL) sequence remains challenging due to the intensive computational demand and prolonged treatment delivery time. This study proposes an unsupervised deep learning model for fast EL pre-selection that minimizes EL switch (ELS) time while maintaining high plan quality. We introduce a novel data representation method, spot-count representation, which encodes the number of proton spots intersecting the target and organs at risk (OAR) in a matrix structured by sorted gantry angles and energy layers. This representation serves as the input of an U-Net style architecture, SPArc_dl, which is trained using a tri-objective function: maximizing spot-counts on target, minimizing spot-counts on OAR, and reducing ELS time. The model is evaluated on 35 nasopharyngeal cancer cases, and its performance is compared to SPArc_particle_swarm (SPArc_ps). SPArc_dl produces EL pre-selection that significantly improves both plan quality and delivery efficiency. Compared to SPArc_ps, it enhances the conformity index by 0.1 (p<0.01), reduces the homogeneity index by 0.71 (p<0.01), lowers the brainstem mean dose by 0.25 (p<0.01), and shortens the ELS time by 37.2% (p < 0.01). The results unintentionally reveal employing unchanged ELS is more time-wise efficient than descended ELS. SPArc_dl's inference time is within 1 second. However, SPArc_dl plan demonstrates limitation in robustness. The proposed spot-count representation lays a foundation for incorporating unsupervised deep learning approaches into EL pre-selection task. SPArc_dl is a fast tool for generating high-quality PAT plans by strategically pre-selecting EL to reduce delivery time while maintaining excellent dosimetric performance.

physics.med-ph

AgentThink: A Unified Framework for Tool-Augmented Chain-of-Thought Reasoning in Vision-Language Models for Autonomous Driving

Vision-Language Models (VLMs) show promise for autonomous driving, yet their struggle with hallucinations, inefficient reasoning, and limited real-world validation hinders accurate perception and robust step-by-step reasoning. To overcome this, we introduce \textbf{AgentThink}, a pioneering unified framework that integrates Chain-of-Thought (CoT) reasoning with dynamic, agent-style tool invocation for autonomous driving tasks. AgentThink's core innovations include: \textbf{(i) Structured Data Generation}, which establishes an autonomous driving tool library to automatically construct structured, self-verified reasoning data explicitly incorporating tool usage for diverse driving scenarios; \textbf{(ii) A Two-stage Training Pipeline}, employing Supervised Fine-Tuning (SFT) with Group Relative Policy Optimization (GRPO) to equip VLMs with the capability for autonomous tool invocation; and \textbf{(iii) Agent-style Tool-Usage Evaluation}, introducing a novel multi-tool assessment protocol to rigorously evaluate the model's tool invocation and utilization. Experiments on the DriveLMM-o1 benchmark demonstrate that AgentThink significantly boosts overall reasoning scores by \textbf{53.91%} and enhances answer accuracy by \textbf{33.54%}, while markedly improving reasoning quality and consistency. Furthermore, ablation studies and robust zero-shot/few-shot generalization experiments across various benchmarks underscore its powerful capabilities. These findings highlight a promising trajectory for developing trustworthy and tool-aware autonomous driving models. Code is available at https://github.com/curryqka/AgentThink.

cs.RO

A Universal Spin-Orbit-Coupled Hamiltonian Model for Accelerated Quantum Material Discovery

The accurate modeling of spin-orbit coupling (SOC) effects in diverse complex systems remains a significant challenge due to the high computational demands of density functional theory (DFT) and the limited transferability of existing machine-learning frameworks. This study addresses these limitations by introducing Uni-HamGNN, a universal SOC Hamiltonian graph neural network that is applicable across the periodic table. By decomposing the SOC Hamiltonian into spin-independent and SOC correction terms, our approach preserves SU(2) symmetry while significantly reducing parameter requirements. Based on this decomposition, we propose a delta-learning strategy to separately fit the two components, thereby addressing the training difficulties caused by magnitude discrepancies between them and enabling efficient training. The model achieves remarkable accuracy (mean absolute error of 0.0025 meV for the SOC-related component) and demonstrates broad applicability through high-throughput screening of the GNoME dataset for topological insulators, as well as precise predictions for 2D valleytronic materials and transition metal dichalcogenide (TMD) heterostructures. This breakthrough eliminates the need for system-specific retraining and costly SOC-DFT calculations, paving the way for rapid discovery of quantum materials.

cond-mat.mtrl-sci

ODverse33: Is the New YOLO Version Always Better? A Multi Domain benchmark from YOLO v5 to v11

You Look Only Once (YOLO) models have been widely used for building real-time object detectors across various domains. With the increasing frequency of new YOLO versions being released, key questions arise. Are the newer versions always better than their previous versions? What are the core innovations in each YOLO version and how do these changes translate into real-world performance gains? In this paper, we summarize the key innovations from YOLOv1 to YOLOv11, introduce a comprehensive benchmark called ODverse33, which includes 33 datasets spanning 11 diverse domains (Autonomous driving, Agricultural, Underwater, Medical, Videogame, Industrial, Aerial, Wildlife, Retail, Microscopic, and Security), and explore the practical impact of model improvements in real-world, multi-domain applications through extensive experimental results. We hope this study can provide some guidance to the extensive users of object detection models and give some references for future real-time object detector development.

cs.CV

Discourse-Driven Evaluation: Unveiling Factual Inconsistency in Long Document Summarization

Detecting factual inconsistency for long document summarization remains challenging, given the complex structure of the source article and long summary length. In this work, we study factual inconsistency errors and connect them with a line of discourse analysis. We find that errors are more common in complex sentences and are associated with several discourse features. We propose a framework that decomposes long texts into discourse-inspired chunks and utilizes discourse information to better aggregate sentence-level scores predicted by natural language inference models. Our approach shows improved performance on top of different model baselines over several evaluation benchmarks, covering rich domains of texts, focusing on long document summarization. This underscores the significance of incorporating discourse features in developing models for scoring summaries for long document factual inconsistency.

cs.CL