SearcharxivSearch

arXiv subjects

Zeyu Chen

Publications and source records attributed to Zeyu Chen.

At least 19 recordsLinked to original sources

Statistical Symmetry Release for Equivariant Quantum Learning

Hard symmetry constraints reduce model complexity, but can also erase label information. Statistical symmetry release determines when finite data and quantum measurements justify relaxing such a constraint, which directions to open, and how far to move. We connect global signal detection to local, loss-dependent improvement. A two-copy twirl--swap gate estimates task information in the symmetry-breaking complement with a dimension-independent copy count under paired-state and group-unitary access; reweighting the same records resolves representation sectors. An exact duality distinguishes this Hilbert--Schmidt signal from the larger signal accessible to bounded-outcome readouts. Local improvement is governed by the release gradient and a loss-corrected double-commutator matrix. Simultaneous confidence bounds convert empirical direction selection into certified descent, using either shared Pauli measurements or scalar probes with state-independent truncation bounds. Gaussian testing lower bounds quantify the cost of searching over unknown directions in the calibrated local experiment. Independent validation controls adaptively generated models, and a fast squared-loss bound preserves the approximation--estimation rate of a nested release path. On an eight-qubit Ising model, shared measurements certify release with 6300 times fewer shots than the specified scalar estimator on the tested budget grids. Quotient quantum natural gradient then controls parameter redundancy during training. Together, these results turn symmetry relaxation into a statistically justified model-selection decision.

cs.LO

BRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL

Asynchronous reinforcement learning has become the standard way to scale training for language models, but the resulting policy lag biases the critic toward the stale behavior policy. Existing work on asynchronous LLM training corrects the actor and leaves this bias unaddressed, while the off-policy value correction of classical RL does not carry over to long-horizon agentic tasks, since a short correction horizon leaves the regression target free of the reward and a long one lets the product of importance ratios drift exponentially with the trajectory length. We propose BRACE, an anchored Bellman-residual correction for stale value models. BRACE bounds the correction horizon to a prefix of policy tokens and anchors a constant-weight Monte-Carlo tail beyond it, which separates policy correction from reward propagation. BRACE improves mean@1 on BrowseComp-Plus by $2.4\%$ over the strongest baseline, runs $2.46\times$ faster per step than synchronous training, and remains stable $50$ updates off-policy.

cs.LG

Rendering-in-the-Loop: An Execution-Driven Agent for Interactive Web Development

Multimodal large language models have achieved remarkable progress in front-end web development, generating interactive webpages from multimodal references such as screenshots and interaction videos. However, existing work largely emphasizes visual metrics such as aesthetics and layout similarity, while overlooking the more critical validation of interactive functionality. We present RILA, an execution-driven agent that puts browser rendering in the loop, iteratively editing generated code from runtime interaction feedback. RILA introduces an Action Interaction Verification (AIV) module that replays the reference interaction trajectory on the generated webpage to collect grounded execution-aware observations, and an Execution-aware Rendering Score (ERS) that jointly measures interaction correctness and visual fidelity to guide iterative optimization. We further build an execution-verified data synthesis pipeline that produces diverse, high-quality training data, offering gains complementary to inference-time optimization. On IWR-Bench, RILA consistently improves both interaction and visual fidelity across foundation models. Notably, with our training pipeline, RILA lifts the compact Qwen3.5-9B backbone from 40.40% to 57.52%, surpassing far larger one-shot generators, including the 1T-parameter Kimi-K2.6 (55.61%) and the proprietary GPT-5.5 (55.74%).

cs.CV

Behavioral Memory under Symmetry in One-Way Quantum Automata

Under compact symmetry, observable behavior reduces to an invariant operator algebra, but its dimension is not yet classical memory: some coordinates are dynamically frozen, some invisible to threshold tests, and some already classical. We develop an operator-algebraic theory that separates these effects through three filters. For one automaton, behavior is the Hilbert--Schmidt pairing between prefix-reachable states and suffix-observable effects, whose rank equals the real Hankel rank without controllability or observability assumptions. Maximizing this invariant over a symmetry-constrained dynamical class gives a structural capacity controlled by the symmetry commutant: its center stores isotypic populations frozen by reversible dynamics, its traceless multiplicity blocks carry movable noncommutative coordinates, dissipation removes the unary spectral loss inside those blocks, and covariant mobility releases relative populations subject to component conservation. Operational realization then determines which surviving coordinates force probabilistic states. For a fixed nontrivial invariant readout, full mobility gives an exact dichotomy in worst-case state cost: a commutative invariant algebra costs exactly its dimension, whereas a noncommutative multiplicity block raises the unrestricted cost by exactly one state. Thus noncommutativity has a one-state worst-case classical price. The known four-letter quadratic-plus-one law at trivial symmetry is the fully mobile endpoint of this principle. Schur--Weyl duality further shows that different preserved symmetries on the same tensor-power Hilbert space can change the worst memory scale from polynomial to exponential, while fixed-weight modules give an exact Catalan law at half filling, with structural capacity equal to the Catalan count minus its central-sector correction.

cs.FL

Quantum Natural Gradient on Quotient Spaces

A parametrized quantum circuit reports its state geometry through a quantum Fisher information matrix (QFIM), often singular. A small Fisher value can reflect exact state-preserving redundancy, compression by the circuit chart, or weak intrinsic distinguishability, and these mechanisms call for different numerical treatments. We show that the circuit metric factors as $F=B^{*}MB$, where $B$ is the state-level circuit-to-orbit differential and $M$ is the intrinsic Fisher operator on the reachable orbit. The factorization identifies the exact kernel as $\ker B$, separates coordinate transfer from intrinsic geometry, and yields the condition for a circuit to realize an orbit-level quantum natural-gradient (QNG) direction. When the prescribed redundancy exhausts the Fisher kernel, the Moore--Penrose update is the minimum-norm horizontal lift of the quotient Riemannian gradient. At critical points with a locally diffeomorphic quotient-to-orbit map, chart singular values cancel from the linearized QNG operator while intrinsic anisotropy remains; in the trace-orthonormal full-control generator frame, excitation-gap anisotropy gives $\kappa_{\mathrm{QNG}}=\kappa_{\mathrm{Eucl}}$. Representation theory makes $M$ explicit on highest-weight, Slater, and fermionic-Gaussian orbits, and cominuscule fidelity flow becomes integrable, with conserved principal-defect ratios, cubic Lie-retracted convergence at $\eta=2$, and stability boundary $\eta=4$. Finite data impose a second boundary: an estimated QFIM and its confidence radius alone cannot distinguish an exact zero from a small physical mode, so the estimated spectrum alone cannot license hard projection. Under depolarization, inverse-Fisher scaling amplifies mean updates and fluctuations together and cannot restore update signal-to-noise. A redundant Slater/Givens circuit confirms exact transfer identities and illustrates finite-shot tradeoffs.

quant-ph

Open-Set Visual Text Forensics via Sparse-Constraint Rectified Flow

Rapidly evolving Generative AI enables sophisticated visual text manipulations that increasingly evade current forensic detectors. Existing discriminative models often overfit specific forgery patterns, limiting their generalization to unseen, open-set attacks. To address this challenge, we propose a generative detector that localizes tampering by estimating the local restoration cost required to align a query image with authentic visual-text statistics, rather than by learning forgery-specific decision boundaries. Specifically, we introduce Sparse-Constraint Rectified Flow (SC-RF), a detector-oriented adaptation of Flow Matching for spatially sparse anomaly localization. We further mitigate data scarcity via self-supervised Artifact Injection and preserve high-frequency forensic traces using a pixel-space Forensic-DiT. Extensive experiments on three benchmarks show that our method achieves state-of-the-art performance, surpassing the runner-up by 3.2 and 4.8 percentage points in F1 and IoU, respectively. In particular, the proposed detector demonstrates strong zero-shot performance on challenging unseen text editing patterns. We further provide an auxiliary stress-test analysis showing that local harmonization produced by our model can weaken the statistical cues relied upon by existing detectors, offering a complementary vulnerability-analysis perspective.

cs.CV

Group-Reflective Self-Distillation for Agentic Reinforcement Learning

Reinforcement learning with verifiable rewards (RLVR) is effective for training large language model agents. However, terminal rewards provide only coarse trajectory-level supervision, leaving successful behaviors, recurring mistakes, and incidental choices entangled in the same outcome signal. Existing agentic self-distillation methods enrich sparse supervision with natural-language skills, but skills retrieved externally or extracted from a single trajectory by stronger models may mismatch current experience, exceed the policy's capability, or remain path-specific. We propose Group-Reflective Self-Distillation (GRSD), which derives capability-aligned and outcome-discriminative guidance from the policy's own verified rollouts. For each prompt, the policy reflects on each verified trajectory in an on-policy group, and a stop-gradient snapshot contrasts the resulting reflections from successful and failed rollouts to construct group-level privileged guidance. Conditioned on this guidance, a self-teacher refines turn-level credit assignment by modulating outcome-based advantages while preserving the verifier-determined learning direction. Experiments across multiple agentic environments and model scales demonstrate that GRSD consistently outperforms competitive baselines and generalizes more effectively to unseen tasks.

cs.AI

Observational Evidence for Anisotropic Metal Excess around Galaxies

The exchange of matter and energy between galaxies and their surroundings drives the cosmic baryon cycle, yet mapping metal transport remains an observational challenge. While simulations predict that galactic winds escape anisotropically along minor axes, evidence for chemical enrichment in neighboring galaxies is limited. We analyze 1,433 galaxy pairs from the Dark Energy Spectroscopic Instrument survey and detect a gas-phase metallicity excess of 14.6% $\pm$ 3.7% to 24.2% $\pm$ 2.6% in neighbors aligned with the minor axis of massive primary at projected separations of 15--60 kpc. This signal, qualitatively consistent with IllustrisTNG simulation, varies from a marginal detection (>92% confidence) at 15--30 kpc to a significant signal (>98% confidence) at 30--60 kpc. In this work, we show that this anisotropic metallicity excess is consistent with a scenario of enrichment via galactic outflows, providing empirical constraints on feedback models and complementing other environmental processes.

astro-ph.GA

Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning

Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, but the resulting stale, off-policy data can destabilize optimization and ultimately cause policy collapse. Existing methods typically retain or discard tokens based solely on the magnitude of their importance ratios, applying the same threshold uniformly across token positions. In this work, we reveal that the natural scale of the importance ratio varies systematically with token entropy. Under asynchronous dynamics, this entropy-ratio scaling dictates two distinct phenomena: at low entropy, the inherent train-inference discrepancy is drastically amplified into substantial sampling noise; at high entropy, in-flight weight updates naturally induce pronounced, legitimate exploratory deviations. Consequently, magnitude-only correction inadvertently admits the amplified noise while strictly masking out the essential exploration triggered by in-flight updates. To address this, we propose the Entropy-Scaled Trust Region (ESTR), which scales each token's off-policy deviation by its local entropy, requiring no auxiliary forward passes or explicit version-switch detection. Across long-horizon agentic tasks and mathematical reasoning benchmarks, ESTR consistently outperforms existing asynchronous methods and achieves the best train-inference consistency. It reaches $37.34$ avg@1 on BrowseComp-Plus and $95.69$ on multi-turn GSM8K, matching synchronous GRPO while achieving a $2.6\times$ speedup.

cs.AI

Environmental Imprints on the Assembly of the Cool Gas around Bright Cluster Galaxies

Galaxy clusters represent extreme cosmic laboratories where environmental processes dramatically reshape their constituent galaxies, yet their effect on the gaseous halos of central galaxies remains poorly constrained. Here we present the first statistical mapping of cool gas around massive brightest cluster galaxies (BCGs) at $z\approx0.55$. Using Mg II absorption in stacked sight-line spectra from over a million background quasars observed by the Dark Energy Spectroscopic Instrument, we compare BCGs to a matched sample of field galaxies and trace the radial profile from 40 kpc to 15 Mpc. Our analysis reveals a striking dual environmental signature: within 200 kpc, the circumgalactic medium (CGM) around BCGs is significantly suppressed compared to that of field galaxies, while at larger radii (200 kpc to 10 Mpc) a pronounced excess of cool gas emerges. This clear transition from suppression in the core to enhancement on such large scales delineates a novel observed pattern for gas regulation by the dense environment. It suggests that clusters may not only strip gas in the core but also facilitate its accumulation in the outskirts. Our results provide key observational constraints on theoretical models of environmental processing in and around the most massive dark matter halos.

astro-ph.GA

RustMizan: A Compilable, Contamination-Aware Benchmarking Framework for Rust Vulnerabilities

LLM agents are increasingly applied to vulnerability analysis, but existing benchmarks have not kept pace. They typically rely on small non-compilable snippets, focus on binary classification (vulnerable or not), and do not account for the risk that publicly-released datasets are part of model training corpora. We introduce RustMizan, a benchmarking framework for Rust vulnerability analysis that addresses these gaps. RustMizan contains compilable code variants at the crate, file, and function levels, with annotations for binary vulnerability detection, CWE classification, and function- and line-level localization. A paired mutation framework produces semantics-preserving code mutants for contamination testing and robustness probing. Across four frontier models in an agentic setup with command-line access, binary classification sits in the 56-65% range, but line localization F1 stays near 20%, and adversarial cues drop line F1 by about 27%.

cs.CR

ACPO: Asymmetric Credit Policy Optimization via Mode-Local Entropy Surrogate

Outcome-supervised reinforcement learning scales to verifiable reasoning tasks, but trajectory-level rewards assign the same outcome signal to all sampled tokens, overlooking their unequal contributions to the reasoning process. Entropy provides a natural indicator of the model's decision state, yet using it for token-level credit assignment presents two key challenges: long-tail probabilities in large vocabularies corrupt both entropy values and gradients, and uncertainty carries distinct semantics across positive- and non-positive-advantage trajectories. We propose Asymmetric Credit Policy Optimization (ACPO), which replaces global entropy with the complement of the top-token probability as a mode-local proxy. Guided by gradient analysis, ACPO incorporates mismatch routing and saturation correction to shape policy updates into the desired asymmetric form, emphasizing uncertain decisions on positive trajectories while penalizing confident regions on failed ones. Theoretically, ACPO locally preserves the advantage direction while bounding surrogate error. Experiments on mathematical and coding reasoning benchmarks, including AIME 2025 and HumanEval Pro, show that ACPO consistently outperforms both entropy-aware methods (e.g., 80/20, GTPO) and strong outcome-supervised RL baselines (e.g., DAPO, SAPO).

cs.LG

ECHO: Prune To Act, Trace To Learn With Selective Turn Memory In Agentic RL

Long-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. Context-management methods make such rollouts feasible by simplifying past interactions through deletion, folding, or memory editing. However, when useful history is collapsed into compressed states, the reconstructed context may no longer reveal which earlier observations support a successful final answer. This creates a mismatch between bounded-context acting and outcome-based reinforcement learning: the policy acts on reconstructed context, while the learner lacks source-level provenance for assigning credit to the evidence that mattered. We propose ECHO, a selective turn-memory framework for traceable context reconstruction in Agentic RL. ECHO compresses each completed environment turn into a compact source-indexed memory record, reconstructs bounded policy contexts by selecting useful records, and reuses the selected source indices to route positive outcome credit to the final trajectory segment, reused evidence turns, memory findings, and memory-selection actions. On BrowseComp-Plus, ECHO reaches 43.4% held-out accuracy, outperforming GRPO at 28.9% and the rolling-summary baseline SUPO at 36.1%, while using fewer turns and lower trajectory volume than SUPO. The trained policy also improves zero-shot generalization across multi-objective QA, code generation, and deep information-seeking benchmarks on both dense and MoE backbones.

cs.LG

Learning from the Self-future: On-policy Self-distillation for dLLMs

On-policy self-distillation (OPSD) has proven effective for post-training large language models (LLMs), yet its application to diffusion LLMs (dLLMs) remains unexplored. Existing OPSD methods are inherently autoregressive-centric. They inject privileged information via left-to-right prefix conditioning with token-level divergence supervision, a design that fundamentally conflicts with the arbitraryorder generation of dLLMs. We introduce d-OPSD, the first OPSD framework tailored for dLLMs. Our approach makes two core contributions. First, we reframe self-teacher construction by using self-generated answers as suffix conditioning, enabling the student model to learn from "self future-experience" rather than privileged prefixes. Second, we shift supervision from token-level to step-level, aligning training with the iterative denoising process of dLLMs. Experiments across four reasoning benchmarks show that d-OPSD consistently outperforms RLVR and SFT baselines with superior sample efficiency, requiring only around 10% of the optimization steps by RLVR and opening a promising pathway for dLLM posttraining. The code is available at https://github.com/xingzhejun/d-OPSD.

cs.CL

Cooler Phases of the Circumgalactic Medium Are More Centrally Concentrated: Constraints from Multiphase Absorption Lines

We present a systematic study of the multiphase circumgalactic medium (CGM) around galaxies and quasars, traced by Ca II $\lambda\lambda3934,3969$, Mg II $\lambda\lambda2796,2803$, and C IV $\lambda\lambda1548,1550$, using the Year 1 dataset from the Dark Energy Spectroscopic Instrument. These three doublets trace CGM gas across a range of temperatures, from cold to warm phases, and we employ a stacking technique to measure the corresponding absorption signals using background sources. We show that CGM structure is strongly phase-dependent: ions tracing progressively cooler gas exhibit increasingly steep radial profiles in equivalent width ($W_i$). These trends are broadly consistent with predictions from cosmological simulations, supporting a phase-stratified CGM in which cooler gas is more centrally concentrated. Specifically, halos of emission-line galaxies exhibit a strong radial transition from cool to warm gas, whereas halos of quasars show a more uniform distribution, likely regulated by active galactic nuclei feedback; in contrast, the cold gas traced by Ca II in low-redshift galaxies is tightly confined to inner regions. We further demonstrate that the radial scaling $W_i \propto D^{\alpha}$ is primarily set by host stellar mass, particularly for the cool-phase medium, suggesting efficient heating processes in massive halos. By jointly leveraging multiple absorption tracers from observations and simulations, we map the CGM from cold to warm phases and place new constraints on the baryon cycle governing galaxy evolution.

astro-ph.GA

CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook

Multimodal representation alignment is pivotal for large language models and robotics. Traditional methods are often hindered by cross-modal information discrepancies and data scarcity, leading to suboptimal alignment spaces that overlook modality-unique features. We propose CodeBind, a framework that optimizes multimodal representation spaces through a modality-shared-specific codebook design. By incrementally aligning target and bridging modalities, CodeBind bypasses the need for fully paired data. Unlike traditional hard alignment, CodeBind decomposes features into shared components for semantic consistency and specific components for modality-unique details. This design utilizes a compositional vector quantization scheme, where a shared codebook bridges modality gaps and modality-specific codebooks mitigate representation bias by preventing dominant modalities from overshadowing others. Validated across nine modalities (text, image, video, audio, depth, thermal, tactile, 3D point cloud, EEG), CodeBind achieves state-of-the-art performance in multimodal classification and retrieval tasks.

cs.CV

Beyond Detection: A Structure-Aware Framework for Scene Text Tracking

Modern visual object trackers show impressive results on general targets, yet their performance drops substantially when dealing with scene text. Although currently underexplored, tracking text in videos is essential for dynamic text manipulations such as segmentation, removal, and editing. To fill this gap, this paper formalizes this specific task as Scene Text Tracking and presents the first systematic work for it. We identify three primary challenges in this task: 1) severe geometric distortions from perspective shifts, 2) high visual ambiguity across different instances, and 3) high sensitivity to fine-grained structural details. To address these issues, we propose SymTrack, a unified detection-free framework with synergistic dual-branch design. It integrates a Cross-Expert Calibration mechanism to reduce semantic bias, along with a Predictive Token Rectification mechanism to correct structural imbalances, complemented by an Adaptive Inference Engine that stabilizes predictions under motion constraints. Considering the lack of dedicated benchmarks for this task, we utilize three datasets from video text spotting to construct a benchmark with high-quality annotations. Extensive experiments demonstrate that SymTrack sets the new state-of-the-art on all three benchmarks, outperforming previous best trackers by up to 11.97\% AUC on $ \text{BOVText}_{\text{SOT}} $. Overall, our work promotes efficient and thorough text tracking, paving the way toward more generalized video text manipulation.

cs.CV