SearcharxivSearch

arXiv subjects

Yi Jiang

Publications and source records attributed to Yi Jiang.

At least 19 recordsLinked to original sources

PhysECD: A Physics-Constrained E(3)-Equivariant Framework for Electronic Circular Dichroism Spectrum Prediction

The electronic circular dichroism (ECD) spectrum is a primary experimental probe for assigning the absolute configuration of chiral molecules, yet interpreting a measured spectrum requires time-dependent density functional theory (TDDFT) calculations that can cost hours per molecule and must be repeated for every candidate stereoisomer and conformation. We present PhysECD, a physics-constrained, parity-aware E(3)-equivariant framework that bypasses computationally expensive TDDFT and predicts ECD spectra directly from the 3D structure of an individual conformer. Instead of regressing the spectrum as an opaque sequence, PhysECD predicts the physical quantities that generate it: per-state excitation energies and electric and magnetic transition dipoles. These quantities determine the rotatory strength R -- the dot product of the two dipoles, a pseudoscalar that reverses sign under mirror reflection -- and yield the final spectrum through a differentiable Gaussian-broadening formula derived from the underlying physics. The parity structure of the equivariant features guarantees the correct chiroptical symmetry: reflecting a molecule exactly negates the predicted spectrum. On the CMCDS dataset, PhysECD attains a per-molecule spectral Pearson correlation of 0.642 (mean) / 0.822 (median), substantially exceeding prior learned predictors while remaining physically interpretable. Experiments across multiple backbones further show that the framework is backbone-agnostic, paving the way for real-time assignment of absolute configuration.

physics.chem-ph

Imaging the vacuum fluctuations of a quantum field

Heisenberg uncertainties lead to inevitable fluctuations in the measurement outcomes for quantum-mechanical observables. For quantum fields, these uncertainties result in random spatial structures in snapshots of a field, even when the field is in its ground (vacuum) state. Such `vacuum fluctuations' are at the heart of a wide range of phenomena, from spontaneous decay processes to the Casimir force and Hawking radiation. Their existence is a key manifestation of the quantumness of the physical world, but usually it is only their consequences that are directly observed. Here, we directly observe spatial vacuum fluctuations of a bosonic quantum field. Our experiments are based on a homogeneous planar atomic Bose--Einstein condensate. The condensate comprises two coherently coupled interacting components (spin states), and the quantum field describes its spin degrees of freedom. In the regime where the interactions dominate over the coherent coupling, our system emulates a (massive relativistic) sine-Gordon field. Images of the field reveal simultaneous fluctuations on different length scales, with scale-dependent amplitudes consistent with theoretical predictions for a vacuum state. Observing such fluctuations in the sine-Gordon limit opens many possibilities for laboratory simulations of relativistic fields in regimes that are presently not theoretically tractable.

cond-mat.quant-gas

MazeRunner: Nonlinear Task and Clue Orchestration for LLM-driven Black-Box Automated Penetration Testing

Penetration testing is essential yet resource-intensive. Although large language models (LLMs) show promise for automating security auditing, existing agents mainly execute end-to-end workflows in simplified linear scenarios. Real-world black-box testing is fundamentally nonlinear: the attack graph is initially unknown and must be incrementally inferred from environmental feedback. Observations may reveal multiple attack branches, failures are often ambiguous, and critical clues may span long action horizons. Existing agents therefore tend to become trapped in depth-first exploration, misdiagnose failures, and forget prior evidence. We present MazeRunner, an autonomous penetration testing system built on a three-agent task-and-clue orchestration framework. It separates global orchestration, context-intensive execution, and failure-oriented review while persistently maintaining task states and environmental evidence. This design supports action revision, prerequisite recovery, branch switching, and long-range clue correlation. We evaluate MazeRunner on 10 recently released HTB targets, limiting each system-target run to 20 million LLM tokens and preventing target-specific solution leakage. With Claude Sonnet 4.5, MazeRunner completes 47.7% of annotated subtasks, compared with 36.2% for PentestGPT-V2 and 34.2% for Claude Code. It achieves user-level or higher access on six targets, including root access on two; each same-model baseline reaches user-level access on only two targets and never obtains root access. Execution-trace analysis further shows that MazeRunner explores more attack branches and acquires shells more efficiently.

cs.CR

Tomographic Phase Imaging with Randomized Probe Imaging

We demonstrate tomographic phase imaging of cubic gold nanoparticles with sub-100 nm resolution requiring just a single far-field coherent diffraction pattern per projection. By using randomized probe imaging (RPI), a real-space amplitude and phase image can be reconstructed from a single diffraction pattern. These phase images are then fed into a standard tomographic workflow to retrieve the 3D volume. Nanoscale X-ray phase tomography has often relied on ptychography, which requires slow 2D scanning for each projection. By using RPI, the data collection is greatly sped up at the cost of some resolution. We compare an RPI-tomography volume to a ptychography-tomography volume of the same sample, finding high consistency between the two methods. To the best of our knowledge, this is the first application of RPI to tomography.

physics.optics

Momentum Structure of Superconductivity and Sublattice Effects from Quasiparticle Interference in CsV$_3$Sb$_5$

Quantum interference encoded in the sublattice texture of kagome Bloch wavefunctions has been widely invoked as a route to correlated states, including chiral charge order, unconventional superconductivity, and their possible intertwining in pair-density-wave (PDW) states. Using sub-Kelvin scanning tunneling microscopy, we conducted spectroscopic mapping of the kagome material CsV$_3$Sb$_5$ with high energy resolution and dense energy sampling through the superconducting gap. Quasiparticle interference (QPI) analysis, aided by ab initio and symmetry calculations, reveals an isotropic superconducting gap on the Fermi surfaces derived from V $M_z$-even ($M_z^+$) $d$ orbitals, thereby constraining possible gap symmetries and limiting any gap anisotropy to the remaining V $M_z$-odd ($M_z^-$) and Sb $p_z$ bands. Meanwhile, the CDW-peak-selected d$I$/d$V$ spectra closely track the spatially averaged density of states and show no distinct enhancement restricted to subgap energies, which do not support an additional PDW modulation within our sensitivity. Finally, the selective absence of specific QPI scattering vectors points to a spectroscopic sensitivity to sublattice character on the Fermi surface. Together, these results provide a clearer experimental picture of the low-energy electronic structure relevant to kagome superconductivity in CsV$_3$Sb$_5$.

cond-mat.supr-con

Organizing Principles for Moir\'e Quantum Matter

Moir\'e flat bands in van der Waals bilayers are usually discussed through a small set of mechanisms associated with the $\Gamma$ and $K$ valleys of hexagonal crystals, and more recently with $M$-valleys systems. Here we show that this view is incomplete. The momentum-space location and effective local orbital character of the monolayer's band edge, in conjunction with the moir\'e symmetry and the symmetry representations of the resulting bands, provide a general set of organizing variables for the emergent low-energy moir\'e Hamiltonian. Applying fully relaxed first-principles calculations, band unfolding and symmetry-representation analysis to more than 600 commensurate twisted bilayers spanning all 2D lattice classes, we identify several routes to moir\'e quantum matter beyond the conventional single-orbital paradigm. The resulting flat bands realize trigonal, honeycomb, square, checkerboard and kagome-like Hubbard models with single-orbital, multi-orbital and multi-site Hilbert spaces; spin-orbit-coupled multi-orbital flat bands exhibit symmetry-indicated topology beyond the conventional $K$-valley setting; and nonsymmorphic moir\'e symmetries enforce semimetallic flat-band connectivity. Analogous quasi-one-dimensional flat-band structures are found in $M$-valley hexagonal systems and $X$-valley square or rectangular systems resulting from emergent momentum-space nonsymmorphic symmetries. Separately, coupled multi-valley manifolds with kagome-like connectivity are identified in several systems whose parent band edges lie at non-high-symmetry points. These results establish a valley-orbital-symmetry framework for connecting parent-material electronic structure to emergent moir\'e Hamiltonians relevant to correlated, topological and symmetry-enforced moir\'e phases.

cond-mat.mtrl-sci

MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents

Smart-city airspace is transforming Uncrewed Aerial Vehicles (UAVs) from passive sensing platforms into cyber-physical decision makers that must follow operational rules under degraded observations and ambiguous language. Existing UAV and multimodal benchmarks evaluate perception, navigation, collaboration, and reasoning, but few assess whether physical evidence, protocol constraints, and action risk remain coupled during critical decisions. We introduce MulRobBench, an offline, protocol-conditioned benchmark for Vision-Language-Action (VLA) UAV agents in smart-city environments. MulRobBench integrates real UAV multimodal observations, protocol-level security policies, and action-level cyber-physical safety into a unified evaluation framework. The benchmark contains 3,024 samples spanning 17 task taxonomy nodes and 12 scoring dimensions across four stages: operational context understanding, multimodal evidence arbitration, degradation-aware reasoning, and risk-aware action planning. Evaluation combines semantic scoring with structural diagnostics, including policy compliance, format compliance, unsafe actions, parsing failures, and dimension-level validity. Across 17 multimodal models, the best semantic protocol-decision score reaches only 0.5141, while the best strict mean scoring-dimension accuracy is 0.1599. A controlled 20-anchor modality-ablation study changes 4-15 action selections per model, confirming that both visual and textual inputs influence decisions. Analysis identifies modality-trust selection, constraint extraction, glare, missing data, and operator shorthand as the primary causes of decision instability. MulRobBench provides a reproducible benchmark for trustworthy multimodal UAV decision making under realistic operational constraints.

cs.MA

Towards Automated Formal Verification of zkEVMs Using LLM-Guided Constraint Synthesis

Zero-Knowledge Ethereum Virtual Machines (zkEVMs) secure Ethereum rollups by generating zero-knowledge proofs that guarantee off-chain execution correctness. However, subtle implementation bugs (e.g., incorrect gas accounting) can lead to valid proofs certifying semantically faulty states, thereby silently defeating cryptographic guarantees. Formal verification via SMT solvers can prevent this, but is bottlenecked by specification: current zkEVM development practice lacks automated methods to translate Rust opcode handlers into verification models. Current practices rely on unsustainable manual specifications, while LLM-based approaches suffer from hallucination and lack formal guarantees. To address this, we propose VeriSynth, a framework that synthesizes executable Python/Z3 verification models from Rust zkEVM code. VeriSynth enforces a hybrid paradigm: an LLM acts strictly as a formalization frontend to translate code into symbolic constraints, while an SMT solver serves as the correctness arbiter. To handle complex multi-component state transitions, VeriSynth integrates semantic decomposition, retrieval-grounded prompting, and verification-guided auto-repair into a closed-loop pipeline. We evaluate VeriSynth on the first source-level zkEVM verification benchmark, encompassing both correct and faulty opcode implementations. VeriSynth achieves a bug detection rate of over 90%, substantially outperforming direct and conversational LLM baselines, as well as a production-grade handwritten mutation-testing suite. Ablation studies confirm that each pipeline component is critical to the framework's overall effectiveness.

cs.SE

Quantum geometry and critical temperature enhancement in MgB$_2$ superconductivity

MgB$_2$, a phonon-mediated superconductor with record-high critical temperature $T_c\simeq 39$ K, is revisited to obtain a comprehensive theory of electrons, phonons, and their coupling with minimal ab initio input. We construct compact analytic models for the electronic structure, phonons, and electron-phonon coupling (EPC) of MgB$_2$. We show that strong in-plane B $sp^2$ bonding realizes an obstructed band structure whose natural description is a bond-centered kagome lattice, yielding small quasi-2D $\sigma$-band Fermi-surface cylinders and pronounced quantum-geometric effects. The phonon spectrum is found to closely track that of a graphene-like boron layer, but the heavy intercalated Mg atoms dominate the three acoustic branches and rigidly lift the boron modes into the optical sector, while the in-plane B-B bond-stretching mode exhibits a pronounced softening along $\Gamma$-A. By symmetry, this $\Gamma$-point bond-stretching mode is the only $\Gamma$ phonon that can couple to the $\sigma$ Fermi surface, explaining its dominant contribution to the EPC. Upon electron doping toward the doubly degenerate band edge of the $\sigma$ sheets, we find that a reduced density of states competes with enhanced EPC matrix elements. At light electron doping, ab initio calculations show that the EPC enhancement dominates, leading to an increase in $T_c$ (within the clean doping limit without disorder effects). Using the Gaussian approximation for the EPC tensor, we further show that this enhancement is overwhelmingly quantum geometric in origin, arising from a geometric EPC contribution of the small $\sigma$ Fermi surface peaked at $\Gamma$. Overall, our results provide a transparent, symmetry-based account of superconductivity in MgB$_2$ and suggest that quantum-geometric effects can be essential for shaping doping trends in phonon-mediated superconductors.

cond-mat.supr-con

GlaKG: A Biomarker-Centric Fundus Knowledge Graph for Explainable Glaucoma Diagnosis and Risk Assessment

Glaucoma is a leading cause of irreversible blindness worldwide, yet most automated diagnosis systems rely on opaque deep-learning models that offer little clinical interpretability. We present GlaKG, a biomarker-centric fundus knowledge graph that integrates structural biomarkers, clinically grounded rules, and image features to produce traceable reasoning for glaucoma diagnosis and risk stratification. GlaKG encodes six entity types (Fundus Image, Optic Disc, Neural Rim, Pathology, Diagnosis, Risk Level), eight relation types, and 11 clinically validated rules into a unified graph, so that every prediction is accompanied by an explicit reasoning chain linking biomarker evidence to activated clinical rules. To keep knowledge-based reasoning strictly separate from label information, we adopt a post-processing fusion framework that combines ResNet50 image embeddings with a normalized KG reasoning-chain score via a tunable weight alpha, with all fitting confined to the training split. On a publicly available, AI-annotated fundus dataset, GlaKG reaches F1 = 0.9953 for binary glaucoma classification and 0.930 accuracy with 0.922 weighted F1 for four-class risk stratification; we report openly that the dataset's biomarker annotations are highly label-correlated, and therefore frame these figures as an upper bound attainable with clean structured biomarkers rather than as leakage-free image-only performance. Feature-importance analysis shows KG-derived and biomarker features contributing near-equally (51.1% vs. 48.9%), and the reasoning chain flags borderline cases by exposing low chain scores rather than failing silently. GlaKG's central contribution is therefore a clinically auditable reasoning framework that complements raw predictive performance by explicitly exposing the biomarker evidence and rule activations behind each decision.

cs.CV

Agents' Last Exam

Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains. We argue that this gap is largely an evaluation problem: widely used benchmarks lack sustained performance measurement on real and economically valuable workflows. This paper introduces Agents' Last Exam (ALE), a benchmark designed to evaluate AI agents on long horizon, economically valuable, real world tasks with verifiable outcomes. Developed in collaboration with 250+ industry experts, ALE covers non-physical industries defined with reference to O*NET / SOC 2018 (the U.S. federal occupational taxonomy). It is organized around a task taxonomy with 55 sub fields grouped into 13 industry clusters covering 1K+ tasks. Current results show that the hardest tier remains far from saturated: across mainstream harness and backbone configurations, the average full pass rate is below 1%. ALE is designed as a living benchmark: its task pool grows continuously as new workflows and industries are onboarded. More broadly, ALE is intended not merely as another leaderboard, but as an instrument for closing the gap between benchmark success and GDP relevant impact.

cs.AI

MOSS-Audio Technical Report

MOSS-Audio is a unified audio-language model for speech, environmental sound, and music understanding, supporting audio captioning, time-aware question answering, timestamped transcription, and audio-grounded reasoning. MOSS-Audio couples a dedicated audio encoder with a modality adapter and a large language model: the encoder produces 12.5 Hz temporal representations, the adapter projects them into the decoder space, and the decoder generates autoregressive text outputs. Two design choices are central to the system: DeepStack cross-layer feature injection, which exposes the decoder to acoustic information from multiple encoder depths, and time markers, which provide explicit temporal cues by inserting timestamp markers into the audio-token stream. At the data level, we design an event-preserving audio annotation pipeline that segments raw audio at coherent event boundaries, applies branch-specific annotation to speech, music, and general audio, and merges the results into unified captions for pretraining. The intermediate branch-specific captions are further retained to support the construction of task-oriented SFT data. The model is pretrained on large-scale audio-language data, with time-aware objectives incorporated to support temporal grounding, and then undergoes multi-stage post-training to enhance instruction following and audio-grounded reasoning. We release 4B and 8B variants in both Instruct and Thinking configurations. MOSS-Audio achieves strong performance across general audio understanding, speech captioning, ASR, and timestamped ASR, positioning it as a promising understanding foundation for future voice agents.

cs.SD

Veda: Scalable Video Diffusion via Distilled Sparse Attention

Scaling Diffusion Transformers to generate high-resolution, long videos is constrained by the quadratic cost of self-attention, and existing sparse attention methods degrade under high sparsity. We show empirically that generation quality is determined not by the sparsity ratio itself, but by how well the sparse mask aligns with the tile-wise geometry of full attention. Based on this insight, we propose Veda, a distilled sparse attention framework that formulates tile selection as an explicit reconstruction problem from full attention. Veda integrates statistics-aware tile scoring with head-aware tiling to reduce estimation error and structural mismatch, enabling aggressive sparsity. A hardware-efficient tile-skipping kernel converts theoretical sparsity into practical wall-clock speedups. Experiments on large video diffusion models, including Waver and Wan2.1, demonstrate substantial acceleration with no noticeable degradation in generation quality. To generate 720P 10-second videos on Waver-T2V-12B, Veda achieves a 5.1$\times$ end-to-end speedup and a 10.5$\times$ self-attention speedup, reducing attention overhead from 92% to 50%. Notably, the gains increase with sequence length, indicating that Veda scales favorably with spatiotemporal resolution across models.

cs.CV

Engineering topological flat bands in $\Gamma$-valley moir\'e systems with Ising-type SOC: twisted 1T-ZrS$_2$ and 1T-SnSe$_2$

Twisted moir\'e superlattices hosting topological flat bands provide a platform to explore the interplay between topology and correlations. Here we investigate topological band structures in $\Gamma$-valley moir\'e systems based on 1T-ZrS$_2$ and 1T-SnSe$_2$. Using large-scale ab initio calculations and continuum modelling, we demonstrate that both materials exhibit an approximate spin-$U(1)$ symmetry and host isolated topological moir\'e valence bands, including quantum spin Hall and high spin Chern states. By constructing a hierarchy of $\Gamma$-valley moir\'e continuum models, we show that isolated moir\'e bands carry a trivial $C_3$ symmetry indicator when the low-energy physics is described by a single effective orbital and a single layer-hybridized branch, either bonding or antibonding. Topological bands therefore arise from inter-branch and/or inter-orbital coupling. Moreover, we determine interaction-driven phase diagrams using Hartree--Fock and exact diagonalization, finding various phases tunable by twist angle, interaction strength, and displacement field. We identify specific conditions under which fractional Chern insulators are favored. Together with previous work showing that the moir\'e conduction bands of 1T-ZrS$_2$ and 1T-SnSe$_2$ realize $M$-valley twisting and host quasi-one-dimensional physics, our results establish these systems as ideal platforms for strongly correlated moir\'e physics and provide a systematic framework for understanding topological band structures in $\Gamma$-valley moir\'e materials.

cond-mat.mtrl-sci

Planar master integrals for two-loop NLO electroweak light-fermion contributions to $g g \rightarrow Z H$

For the two-loop next-to-leading-order electroweak (NLO EW) corrections to $gg \rightarrow ZH$, the light-fermion contributions can be classified into eight distinct topologies. Using the canonical differential-equations method, we perform an analytic computation of the master integrals (MIs) associated with the four planar topologies. Canonical bases are constructed using the Magnus-expansion method, and the resulting alphabets consist of algebraic symbol letters involving nontrivial radicals. We develop a systematic framework for identifying the radical structures of the canonical MIs, enabling their organization into suitable subsystems and, whenever possible, their representation in terms of Goncharov polylogarithms (GPLs) up to $\mathcal{O}(\epsilon^4)$. Only a few MIs at $\mathcal{O}(\epsilon^3)$ and $\mathcal{O}(\epsilon^4)$ are instead represented as one-fold integrals over GPLs, due to the presence of nested square roots that obstruct the simultaneous rationalization of all radicals.

hep-ph

SAVOIR: Learning Social Savoir-Faire via Shapley-based Reward Attribution

Social intelligence, the ability to navigate complex interpersonal interactions, presents a fundamental challenge for language agents. Training such agents via reinforcement learning requires solving the credit assignment problem: determining how individual utterances contribute to multi-turn dialogue outcomes. Existing approaches directly employ language models to distribute episode-level rewards, yielding attributions that are retrospective and lack theoretical grounding. We propose SAVOIR (ShApley Value fOr SocIal RL), a novel principled framework grounded in cooperative game theory. Our approach combines two complementary principles: expected utility shifts evaluation from retrospective attribution to prospective valuation, capturing an utterance's strategic potential for enabling favorable future trajectories; Shapley values ensure fair credit distribution with axiomatic guarantees of efficiency, symmetry, and marginality. Experiments on the SOTOPIA benchmark demonstrate that SAVOIR achieves new state-of-the-art performance across all evaluation settings, with our 7B model matching or exceeding proprietary models including GPT-4o and Claude-3.5-Sonnet. Notably, even large reasoning models consistently underperform, suggesting social intelligence requires qualitatively different capabilities than analytical reasoning.

cs.AI

Stratagem: Learning Transferable Reasoning via Trajectory-Modulated Game Self-Play

Games offer a compelling paradigm for developing general reasoning capabilities in language models, as they naturally demand strategic planning, probabilistic inference, and adaptive decision-making. However, existing self-play approaches rely solely on terminal game outcomes, providing no mechanism to distinguish transferable reasoning patterns from game-specific heuristics. We present STRATAGEM, which addresses two fundamental barriers to reasoning transfer: domain specificity, where learned patterns remain anchored in game semantics, and contextual stasis, where static game contexts fail to cultivate progressive reasoning. STRATAGEM selectively reinforces trajectories exhibiting abstract, domain-agnostic reasoning through a Reasoning Transferability Coefficient, while incentivizing adaptive reasoning development via a Reasoning Evolution Reward. Experiments across mathematical reasoning, general reasoning, and code generation benchmarks demonstrate substantial improvements, with particularly strong gains on competition-level mathematics where multi-step reasoning is critical. Ablation studies and human evaluation confirm that both components contribute to transferable reasoning.

cs.AI

M100: An Orchestrated Dataflow Architecture Powering General AI Computing

As deep learning-based AI technologies gain momentum, the demand for general-purpose AI computing architectures continues to grow. While GPGPU-based architectures offer versatility for diverse AI workloads, they often fall short in efficiency and cost-effectiveness. Various Domain-Specific Architectures (DSAs) excel at particular AI tasks but struggle to extend across broader applications or adapt to the rapidly evolving AI landscape. M100 is Li Auto's response: a performant, cost-effective architecture for AI inference in Autonomous Driving (AD), Large Language Models (LLMs), and intelligent human interactions, domains crucial to today's most competitive automobile platforms. M100 employs a dataflow parallel architecture, where compiler-architecture co-design orchestrates not only computation but, more critically, data movement across time and space. Leveraging dataflow computing efficiency, our hardware-software co-design improves system performance while reducing hardware complexity and cost. M100 largely eliminates caching: tensor computations are driven by compiler- and runtime-managed data streams flowing between computing elements and on/off-chip memories, yielding greater efficiency and scalability than cache-based systems. Another key principle was selecting the right operational granularity for scheduling, issuing, and execution across compiler, firmware, and hardware. Recognizing commonalities in AI workloads, we chose the tensor as the fundamental data element. M100 demonstrates general AI computing capability across diverse inference applications, including UniAD (for AD) and LLaMA (for LLMs). Benchmarks show M100 outperforms GPGPU architectures in AD applications with higher utilization, representing a promising direction for future general AI computing.

cs.LG