SearcharxivSearch

arXiv subjects

Yi Chen

Publications and source records attributed to Yi Chen.

At least 19 recordsLinked to original sources

Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents

Self-evolving runtime harnesses can substantially improve the capabilities of large language model (LLM) agents and provide a promising paradigm for optimizing agent execution. Existing harness evolution methods typically rely on iterative search, repeatedly evaluating and revising candidate harnesses based on execution feedback from task instances. While this paradigm enables continuous harness optimization, it incurs substantial time overhead due to repeated agent executions and code modifications, and may overfit to observed tasks and specific failure patterns, resulting in degraded generalization to unseen tasks. We identify the lack of principled failure diagnosis as a key bottleneck in harness evolution: an observed failure can reflect either model-specific deficiencies or systematic harness deficiencies, and directly optimizing against individual failures can lead to unnecessary model-specific accommodation. We therefore propose Ecdysis, an efficient and effective framework that distinguishes model-specific accommodation from harness-level repair and biases adaptation toward systematic harness deficiencies by identifying recurring cross-task failure patterns. Ecdysis adopts a batch-level cross-instance failure aggregation paradigm to jointly analyze failure evidence from multiple task instances and further introduces Failure-Driven Collaborative Refinement to diagnose failure causes and iteratively refine harness modification specifications. By combining cross-instance failure analysis with multi-role diagnosis, Ecdysis enables more effective harness evolution with lower training time. Experiments show that Ecdysis achieves up to a 1.84x speedup in harness training compared with existing harness evolution methods, while improving the reasoning accuracy of the resulting harnesses by 18.56%.

cs.SE

Bridging the Gap: A Longitudinal Analysis of Extended Identifiers in the Post-Cookie Era

As third-party cookies fade because of browser restrictions, the online advertising ecosystem is turning to extended identifiers (EIDs) as an alternative. EIDs are persistent user identifiers, such as hashed email addresses, that are employed to link users across domains and devices. This paper presents a 41-month longitudinal study examining EID usage in over 145 million HTTP header bidding requests sent to six major supply-side platforms (SSPs) from 616,539 websites. Our findings show that EIDs are widely used and are becoming increasingly prevalent in the digital advertising ecosystem, reaching 83.76% of studied websites by May 2025. Our analysis of the 18 popular EID providers that account for 99.42% of all transmitted EIDs in our dataset raises concerns about the readiness of EIDs as an alternative to third-party cookie tracking. In terms of accuracy, only one identity provider consistently recognizes and identifies that the visitor is a self-identified bot crawler, and many providers regularly transmit multiple EIDs for the same visitor. We also identify privacy concerns with EIDs, as 12 of the providers create persistent EIDs that can identify the same user across visits, websites, devices, and months. Finally, we found that 16 providers transmit EIDs on EU websites without user consent.

cs.NI

A Task Force on Strong Coupling Determinations from Event Shapes

The strong coupling constant $\alpha_s$ is a fundamental parameter of the Standard Model. Its precise determination is essential for accurately predicting, studying, and understanding processes at the Large Hadron Collider and future experiments such as the Future Circular Collider. Event shape and correlator observables measured at electron-positron colliders provide one of the cleanest environments for extracting $\alpha_s$, thanks to their sensitivity to $\alpha_s$ and the availability of high-precision data from the Large Electron-Positron Collider. More broadly, such observables provide an ideal setting to develop and test our understanding of the perturbative and non-perturbative elements of Quantum Chromodynamics, which will underpin the field's precision and discovery frontiers for decades to come. Despite these advances, significant discrepancies persist between different determinations of $\alpha_s$ from event shapes, both in the extracted central values and estimated uncertainties. This document motivates the establishment of a dedicated Task Force to coordinate a community-wide effort addressing these open questions. We report on the first two-day meeting held at CERN in November 2025, summarizing the scientific discussion and documenting the experimental analyses identified as priorities during the meeting, as well as the concrete list of tasks to be carried out by the theory community in preparation for future meetings.

hep-ph

Reinforcement Learning Enhanced LLM Agents for Complex Vehicle Routing Problems

Vehicle Routing Problems (VRPs) are fundamental combinatorial optimization problems with widespread applications in various scenarios. The advanced optimization solvers can effectively solve such problems. However, modeling complex VRP variants for solvers often requires substantial domain expertise, which limits the accessibility of advanced optimization technologies. In this paper, we propose Reinforcement Learning Enhanced LLMAgents(RLEA), a multi-agent framework designed to automate the modeling of complex VRPs. RLEA introduces a lightweight neural Planner trained with Soft Q-learning to efficiently orchestrate the actions of LLM-based agents. In addition, we equip the system with an evolutionary memory module and retrieval-augmented generation, enabling the agent to leverage both accumulated experience and external solver knowledge during program generation and refinement for solving VRPs. We evaluated 48 distinct VRP variants across various solvers. The experimental results demonstrate that RLEA outperforms the previous state-of-the-ar method, achieving a 16.67% higher success rate while significantly reducing runtime errors. These results validate that integrating reinforcement learning with LLM-based reasoning is highly effective for automated optimization modeling. The appendix is available at: https://doi.org/10.5281/zenodo.19134435.

cs.AI

Cut-ViT: Task-Specific Model Pruning via Gram Anchoring Subspace Consistency

Pruning visual foundation models has attracted considerable attention. However, existing methods focus on rigid point-to-point token alignment on a single dataset for pruning, suffering from two limitations: i) robustness degradation, and ii) task-specificity deficiency. To address these limitations, we propose a task-specific pruning pipeline, named Cut-ViT. Specifically, we first construct gram anchoring matrices from both spatial and semantic perspectives, and perform the subspace decomposition to extract the corresponding subspace bases. Basis-agnostic and residual constraints are then adopted to align the gram subspaces between the native and pruned DINOv3 models along spatial and channel dimensions, enabling subnetworks to inherit robust feature representations of native DINOv3. Furthermore, we design spectral entropy adaptation, which quantifies the information density of feature manifolds along spatial and channel dimensions, thereby adapting the pruning objective to specific downstream tasks. Experiments show that Cut-ViT requires approximately one minute on a single A100 GPU to obtain subnetworks at various sparsity levels, using only 20.9% of the time and 45.5% of the GPU memory compared with previous methods, while achieving SOTA performance on six tasks across nine datasets.

cs.CV

Distance Is Not Enough: Forget-Retain Alignment Gap Predicts LLM Relearning Robustness

Machine unlearning aims to make a model forget specific data, yet unlearned LLMs often fail to stay unlearned: brief fine-tuning can revive removed knowledge. Existing robustness predictors rely on global weight-space displacement, but distance alone can be misleading when random or destructive updates collapse performance. We argue that relearning robustness depends on update structure: robust unlearning should affect forget-critical weights while sparing retain-critical ones. We introduce the Forget-Retain Alignment Gap (FRAG), a training-free predictor that scores an update's forget-retain alignment without running a relearning attack, and separates selective from dense updates more reliably than global distance. Building on the forget-critical, retain-sparing principle, Forget-Retain Pruning (FRP) improves relearning robustness. Our results suggest that weight selectivity better explains robustness than distance alone. Code is available at https://github.com/Yi1-Chen/FRAG.

cs.AI

First measurement of the one-point charge correlator in $e^+e^-$ collisions at $\sqrt{s} = 91.2$ GeV with DELPHI Open Data

The chiral structure of the $Z$ couplings imprints a parity-odd flow of electric charge on hadronic $Z$ decays. The related forward-backward asymmetries, a key set of observables in the electroweak precision program, were measured at LEP and SLC using the jet charge. The one-point charge correlator offers a complementary route, measuring the hadronic charge flow directly as a function of polar angle relative to the incoming electron-beam axis, without reference to jets or a reconstructed quark direction, following the formalism developed in a companion paper. We report its first measurement, using $61~\mathrm{pb}^{-1}$ of archival DELPHI Open Data recorded at $\sqrt{s} = 91.2$~GeV in 1994 and 1995. Detector effects are corrected in two stages. The first is derived from fully simulated samples, and the second bounds the residual charge-misreconstruction difference between data and simulation using a measurement in $e^+e^-\to\tau^+\tau^-$ events. The measured charge correlator exhibits the characteristic parity-odd $\sin(2\theta)$ modulation and agrees with the \textsc{PYTHIA}~8.3 prediction. The measurement demonstrates the experimental feasibility of the observable and establishes strategies for controlling associated detector effects, paving the way for a new program of charge-flux measurements, both in archival $e^+e^-$ data and at future colliders.

hep-ex

Analysis note: one-point charge correlator with DELPHI Open Data

We present the first measurement of the one-point charge correlator, the angular flux of electric charge in hadronic final states, using DELPHI Open Data collected at LEP-1 at $\sqrt{s} = 91.2$~GeV during 1994 and 1995. The data, corrected for detector effects, exhibit a clear $\sin(2\theta)$ modulation, consistent with the parity-violating hadronic charge flow that the chiral structure of the $Z$ couplings imprints on the final state. The measurement demonstrates the experimental feasibility of the observable and establishes strategies for controlling associated detector effects, thereby motivating a new program to measure charge-flux observables. This note documents the experimental details supporting the companion experimental paper and the joint theory--experiment Letter.

hep-ex

Observing Macroscopic Consequences of Electroweak Anomalies with Archival DELPHI Data

In this Letter, we emphasize that asymmetries in charge flux produced in the decays of on-shell Z-bosons provide a macroscopic manifestation of electroweak anomalies in the Standard Model (SM). We propose that these can be cleanly observed using charge correlators, providing a new formulation of forward-backward asymmetry measurements that is particularly well suited for precision studies of hadronic decays. Using archival DELPHI data, we perform a first measurement of the one-point charge correlator of electromagnetic charge flux on hadrons, and cleanly observe the macroscopic imprint of the underlying anomaly. Our analysis illustrates the potential of charge correlators as precision electroweak observables, and motivates a renewed effort to resolve longstanding tensions in hadronic asymmetry measurements.

hep-ph

Emergent trans-moir\'e orbitals and topology in rhombohedral graphene

The fractional quantum anomalous Hall effect (FQAHE) exhibited in fractional Chern insulators has recently been demonstrated in twisted MoTe2 and rhombohedral graphene/hBN moir\'e superlattices, promising new routes toward topological quantum computation. Central to realizing this promise is the understanding of the underlying microscopic mechanism. This, however, remains elusive in the case of rhombohedral graphene, with the crux being its two seemingly paradoxical conditions: a pronounced small-twist-angle ({\theta}) moir\'e interface, yet only when electrons are kept distant from it. Here, by scanning tunnelling microscopic imaging with both conditions fulfilled, we capture dramatic electronic structure reshaping in rhombohedral hexalayer graphene by unforeseen 'trans-moir\'e orbitals', which emerge on the other, distant side of the moir\'e interface but nevertheless enforce the moir\'e periodicity at all measured fillings. We visualize a hierarchy of spatially and energetically distinct trans-moir\'e orbitals which doped electrons must sequentially occupy--the lowest-energy orbital, expectedly responsible for the FQAHE at small fillings, carries a hollow-cage-like shape. Remarkably, these trans-moir\'e orbitals vanish at {\theta} {\gtrsim} 1{\deg}, and so do QAHE plateaus in similar devices. Simulations reveal an interaction-driven charge-redistribution mechanism which shapes the trans-moir\'e orbitals and corresponding Chern minibands. With our findings providing the missing microscopic link, the paradoxical conditions find a natural explanation: electrons are not simply kept distant from a small-{\theta} moir\'e interface; they are forced into topological trans-moir\'e orbitals, forged precisely under such conditions. Our microscopic diagnostics unlocks a wide range of possible 'synthetic' FQAHE platforms.

cond-mat.mes-hall

Bridging Balancing Weights and Augmentation in Covariate-adjusted Analyses with Time-to-Event Endpoints: Theory and Practical Recommendations

Covariate adjustment improves the efficiency of treatment-effect analyses in randomized clinical trials, provided the adjustment targets the correct quantity. For time-to-event endpoints, two marginal targets are of primary interest: the log-rank test for the presence of a treatment effect and the marginal hazard ratio for its magnitude. Existing covariate adjustment approaches reach these targets by different ways. Augmentation adjusts the log-rank score by regressing derived outcomes on the baseline covariates within each arm. Weighting instead reweights the two arms to balance the covariates before the survival comparison is formed: inverse probability weighting does so through a fitted propensity model, while calibration weighting solves directly for weights that match covariate means. In this manuscript, we first develop balancing weighting for time-to-event endpoints, covering both calibration weights (stable balancing weights and entropy balancing) and propensity score weights, and prove that any balancing-regular weighting is first-order equivalent to the augmented log-rank score and to the root of the marginal Cox score. All three routes therefore deliver the same estimator to first order, and calibration reaches it without fitting any model. The weighted procedures thereby inherit the validity and guaranteed efficiency gain of the augmentation approach. In addition, we show that the efficiency gain grows with the prognostic strength of the adjustment covariates, while the practical caveat lies in variance estimation, for which we give recommendations to guard against finite-sample Type I error inflation. We further confirm our results through simulation studies and an analysis of the REWIND cardiovascular trial.

stat.ME

pyHB: an open-source automatic-differentiation-enhanced semi-analytical solver for nonlinear dynamics

The Harmonic Balance (HB) method is widely used to compute and analyze the periodic responses of nonlinear systems. However, its application to high-dimensional complex systems is limited by the burden of handling the partial derivatives of the nonlinearities. This work presents pyHB, an open-source, automatic-differentiation-enhanced semi-analytical framework that integrates the complete HB workflow for general user-defined nonlinear systems. The proposed formulation exploits localized nonlinearities and applies PyTorch-based automatic differentiation (AD) only to the reduced nonlinear force, thereby avoiding the need for user-supplied derivatives of the nonlinear force and maintaining controllable GPU memory usage. Weighted arc-length continuation, sparse matrix assembly, a blocked solution strategy for the augmented continuation equations, and Floquet-based stability analysis are incorporated within a modular architecture that separates model definition from reusable numerical procedures. Hence, pyHB can provide a complete landscape of the nonlinear system's periodic response based solely on the user-defined dynamical equations. Four examples, including a quasi-zero-stiffness isolator, a nonlinear piezoelectric energy harvester, a 284 degrees of freedom (DOFs) aeroengine model, and a 2000 DOFs Bernoulli beam, demonstrate the ability of pyHB to trace stable and unstable solution branches and capture subharmonic resonance, combination resonance, and mixed-order electromechanical responses. Notably, in the Bernoulli beam example with 202000 HB unknowns, the AD-enhanced solver requires approximately 0.44s per continuation point, achieving several-hundred-fold speedup compared to the Newmark-$\beta$ method and remaining 637.8MB of additional RAM and 243.5MB of GPU memory. The proposed pyHB provides a general, one-stop benchmark platform for HB-based nonlinear dynamics analysis.

cs.MS

Radiative corrections in neutral-current (anti)neutrino elastic scattering at $\text{GeV}$ energies I: Nucleon targets

We introduce radiative corrections in neutral-current (anti)neutrino-nucleon elastic scattering at $\text{GeV}$ energies within the effective field theory framework. We factorize cross sections into soft and hard functions, clarify the (anti)neutrino flavor dependence at both amplitude and cross-section levels, and improve the quantum chromodynamics (QCD) contributions to low-energy neutral-current processes. The radiative corrections at the single-nucleon level reach a magnitude comparable to the contributions from strange quarks. We also compare our results with the experimental data from BNL E734 and MiniBooNE collaborations, finding excellent agreements with the experimental data.

hep-ph

Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations

On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly understood. We present a systematic study examining the role, pathologies, and regulations of OPD. We first clarify the role of OPD as an exploration catalyst: it steers the student toward correct reasoning paths via dense token-level guidance, without expanding capability ceiling. We confirm this by showing that prompt diversity matters more than per-problem sampling numbers, and critically, that the effectiveness of OPD hinges entirely on the quality of its guiding signal. This dependency exposes two pathologies that derail exploration. The Student-Teacher Mismatch occurs when a large teacher-student distributional gap causes the guiding signal to misalign with task correctness, steering exploration in counterproductive directions. Length Exploitation arises when the aggregated token-level objective creates length-dependent shortcuts, allowing the student to game the reward landscape through response truncation or redundant padding, exploring degenerate length modes rather than reasoning strategies. To tame these pathologies, we investigate lightweight signal regulations: advantage clipping and log-scale compression, ensuring exploration is guided by faithful signals. Experiments across seven benchmarks demonstrate that these regulations alleviate length exploitation and enable effective distillation, stably surpassing OPD variants and RLVR baselines, thereby confirming that well-regulated signal quality, rather than mere teacher scale, governs successful exploration in OPD.

cs.CL

Just-In-Time Scene Graph Growth: Combating Perceptual Saturation in Long-Horizon Robotics

While 3D Scene Graphs (3DSGs) provide crucial structured representations for embodied agents, conventional Ahead-of-Time, build-everything-then-filter pipelines conflict with the real-time, low-latency demands of edge platforms, inducing a perceptual saturation effect via severe observation redundancy. To resolve this, we present JITOMA (Just-In-Time On-demand Memory Activation), a closed-loop framework that unifies task reasoning, perception, and memory into a just-in-time growth process. Instead of exhaustively mapping the entire environment, JITOMA leverages a top-down task heatmap at the frontend to filter continuous observations, routing minimal streams to maintain a global foundation of low-cost, dormant anchors. Upon a cognitive query, the backend Large Language Model (LLM) parses the robotic intent to dynamically awaken task-relevant anchors, triggering resource-intensive operations -- such as dense node captioning and functional inference -- exclusively within the activated local subgraph. To evaluate these dynamic capabilities and study perceptual saturation trade-offs, we introduce JITOMA-Bench, a comprehensive suite for long-horizon multi-tasking and complex multi-step reasoning. Extensive experiments demonstrate that JITOMA substantially reduces active graph size and captioning latency, while maintaining stable processing time under long-horizon task switching.

cs.CV

VINE: Taming Generative Control Policies for Reinforcement Learning

Flow-matching policies have emerged as an effective policy parameterization for robot learning. They iteratively generate actions from noise, enabling highly expressive modeling of complex and multimodal action distributions. However, prior works observed that scaling these policies with value-gradient reinforcement learning (RL) often leads to training instability. Existing methods attribute this instability to iterative generation and therefore avoid end-to-end value-gradient optimization by sacrificing iterative generation, high expressiveness, or value-gradient optimization. Contrary to prior belief, we show the instability does not stem from iterative generation itself, but from the vanilla sampling strategy originally designed for behavior cloning, which becomes brittle under value-gradient RL. Motivated by this insight, we propose VINE, an RL-oriented sampling method that enables stable end-to-end value-gradient optimization for flow-matching policies. Instead of following a single flow trajectory, VINE reconstructs a new interpolation state at every denoising step, creating a stable differentiable path for value-gradient propagation while remaining compatible with the original flow-matching denoising process. As a result, VINE preserves the expressiveness and iterative generation of flow-matching without sacrificing end-to-end value-gradient optimization. Despite performing end-to-end backpropagation through all ten denoising steps, VINE achieves stable policy improvement and consistently outperforms state-of-the-art RL methods on the OGBench offline RL benchmark and real-world robotic manipulation task. Videos are available on our website: https://agibottech.github.io/vine.

cs.RO

Analysis note: Long-range near-side correlation in $e^+e^-$ with $W$-boson-pair events at 183-209 GeV with ALEPH archived data

Events characterized by a high multiplicity of charged particles have been a central focus in the study of collective behavior across both large and small collision systems. A previous measurement of two-particle angular correlations in $e^+e^-$ collisions at center-of-mass energies up to $\sqrt{s} = 209$ GeV, using LEP2 data, revealed discrepancies with Monte Carlo (MC) predictions at high multiplicity, suggesting the possible emergence of long-range near-side correlations even in the simplest collision system. Unlike at lower energies, where quark-antiquark production dominates, $W^+W^-$ processes become increasingly important at multiplicities above 30. On the one hand, the observed excess in long-range correlations may reflect the more complex color-string configurations arising from $W^+W^-$ production. On the other hand, it can simply arise from the higher final-state multiplicity made possible by the increased collision energy, independent of the underlying production mechanism. To discriminate between these competing interpretations, we present a measurement of two-particle angular correlations in $e^+e^-$ collisions at $\sqrt{s} = 183-209$ GeV, with a focus on enhancing the contribution from $W^+W^-$ processes. The analysis uses data collected by the ALEPH detector during the LEP2 program. Correlation functions are evaluated across a broad range of pseudorapidities and full azimuth, in bins of charged-particle multiplicity. A ridge-like modulation is seen for multiplicity above 50, deviating from the MC reference. In addition, the correlation functions are further decomposed into a Fourier series, and the resulting harmonic coefficients $v_n$ are compared with predictions from the archived Monte Carlo sample. For multiplicity starting from 30, the signed $v_2$-like proxy goes from negative to positive, also deviating from the MC baseline.

hep-ex

ElementCheck: Complexity-Aware Long-Form Text Factuality Evaluation via Sentence Elements

Existing long-form factuality evaluation relies on the decompose-retrieve-verify pipeline. However, the pipeline suffers from noise from claim decomposition and fixed verification granularity, resulting in unreliable results. We propose ElementCheck, a complexity-aware framework that verifies long-form outputs via sentence elements. Instead of uniformly decomposing sentences into atomic sub-claims, ElementCheck extracts entity pairs that are explicitly linked through verifiable connections in the original sentence as elements, and organizes these into an element graph. The graph topology provides a structural signal for estimating sentence complexity, enabling direct verification for simple sentences and targeted element-level refinement and verification for complex ones. To support fine-grained evaluation, we construct a new benchmark \textbf{FastFact-Sent} by mapping isolated claims from FastFact-Bench back to their source sentences. Experiments on FastFact-Sent and two domain-specific benchmarks show ElementCheck consistently improves factuality verification across five backbone models while maintaining a favorable accuracy-cost trade-off. Further analyses demonstrate that complexity-aware verification reduces unnecessary re-verification and maintains stability across different backbones. The code is available at \href{https://github.com/gudehhh666/elementcheck.git}{Here}.

cs.CL