SearcharxivSearch

arXiv subjects

Yingqi Wang

Publications and source records attributed to Yingqi Wang.

8 recordsLinked to original sources

Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents

Self-improving agents accumulate capability by repeatedly rewriting procedural policies, controllers, or heuristic rules. They typically rely on self-authored tests or metrics to decide whether to accept subsequent edits. The agent controls both the optimized object and its verifier. As a result, self-assigned scores can remain near perfect while real deployment performance degrades or stays low. We study this problem through the verifier--deployment gap. This gap refers to the discrepancy between an agent's self-authored verification signal and a sealed deployment evaluation that the agent cannot observe or access. We ask how self-authored verification fails under iterative policy-and-test rewriting, how the failure changes with capability, and how little exogenous trust is sufficient to prevent real regressions from being deployed. To address this problem, we introduce a Sealed Exogenous Acceptance Loop (SEAL). SEAL retains self-authored tests but compares each candidate with the incumbent through a fixed harness-side audit. The agent cannot author or inspect the audit, receives only accept/reject, and the whole incumbent state is retained after a clear regression. Our experiments show that this problem often appears in heuristic learning settings. These settings require trial-and-error discovery of the target objective. We further find that failures of self-written verification are stratified by capability. Weaker agents tend to damage previously acquired strategies behind easy self-tests. Stronger agents are more stable, but they still mismeasure the deployment distribution. Standard self-written constraints do not reliably close this gap. In contrast, SEAL outperforms unprotected baselines across six models and three random seeds. Reliable self-improvement need not abandon self-verification, but it requires at least one deployment-acceptance signal outside the agent's control.

cs.CL

From Words to Worlds: Benchmarking Cross-Cultural Cultural Understanding in Machine Translation

Culture-expressions, such as idioms, slang, and culture-specific items (CSIs), are pervasive in natural language and encode meanings that go beyond literal linguistic form. Accurately translating such expressions remains challenging for machine translation systems. Despite this, existing benchmarks remain fragmented and do not provide a systematic framework for evaluating translation performance on culture-loaded expressions. To address this gap, we introduce CulT-Eval, a benchmark designed to evaluate how models handle different types of culturally grounded expressions. CulT-Eval comprises over 7,959 carefully curated instances spanning multiple types of culturally grounded expressions, with a comprehensive error taxonomy covering culturally grounded expressions. Through extensive evaluation of large language models and detailed analysis, we identify recurring and systematic failure modes that are not adequately captured by existing automatic metrics. Accordingly, we propose a complementary evaluation metric that targets culturally induced meaning deviations overlooked by standard MT metrics. The results indicate that current models struggle to preserve culturally grounded meaning and to capture the cultural and contextual nuances essential for accurate translation. Our benchmark and code are available at https://anonymous.4open.science/r/CulT-Eval-E75D/.

cs.CL

MAS-FIRE: Fault Injection and Reliability Evaluation for LLM-Based Multi-Agent Systems

As LLM-based Multi-Agent Systems (MAS) are increasingly deployed for complex tasks, ensuring their reliability has become a pressing challenge. Since MAS coordinate through unstructured natural language rather than rigid protocols, they are prone to semantic failures (e.g., hallucinations, misinterpreted instructions, and reasoning drift) that propagate silently without raising runtime exceptions. Prevailing evaluation approaches, which measure only end-to-end task success, offer limited insight into how these failures arise or how effectively agents recover from them. To bridge this gap, we propose MAS-FIRE, a systematic framework for fault injection and reliability evaluation of MAS. We define a taxonomy of 15 fault types covering intra-agent cognitive errors and inter-agent coordination failures, and inject them via three non-invasive mechanisms: prompt modification, response rewriting, and message routing manipulation. Applying MAS-FIRE to three representative MAS architectures, we uncover a rich set of fault-tolerant behaviors that we organize into four tiers: mechanism, rule, prompt, and reasoning. This tiered view enables fine-grained diagnosis of where and why systems succeed or fail. Our findings reveal that stronger foundation models do not uniformly improve robustness. We further show that architectural topology plays an equally decisive role, with iterative, closed-loop designs neutralizing over 40% of faults that cause catastrophic collapse in linear workflows. MAS-FIRE provides the process-level observability and actionable guidance needed to systematically improve multi-agent systems.

cs.SE

Mechanistic Insights into Non-Adiabatic Interband Transitions on a Semiconductor Surface Induced by Hydrogen Atom Collisions

To understand the recently observed mysterious non-adiabatic energy transfer for hyperthermal H atom scattering from a semiconductor surface, Ge(111)c(2*8), we present a mixed quantum-classical non-adiabatic molecular dynamics model based on time-dependent evolution of Kohn-Sham orbitals and a classical path approximation. Our results suggest that facile non-adiabatic transitions occur selectively at the rest atom site, featuring excitation of valance band electrons to the conduction band, but not at the adatom site. This drastic site specificity can be attributed to the changes of the local band structure upon energetic H collisions at different surface sites, leading to transient near-degeneracies and significant couplings between occupied and unoccupied orbitals at the rest atom, but not at the adatom. These insights shed valuable light on the collisional induced non-adiabatic dynamics at semiconductor surfaces.

physics.chem-ph

Dynamics-aware Adversarial Attack of Adaptive Neural Networks

In this paper, we investigate the dynamics-aware adversarial attack problem of adaptive neural networks. Most existing adversarial attack algorithms are designed under a basic assumption -- the network architecture is fixed throughout the attack process. However, this assumption does not hold for many recently proposed adaptive neural networks, which adaptively deactivate unnecessary execution units based on inputs to improve computational efficiency. It results in a serious issue of lagged gradient, making the learned attack at the current step ineffective due to the architecture change afterward. To address this issue, we propose a Leaded Gradient Method (LGM) and show the significant effects of the lagged gradient. More specifically, we reformulate the gradients to be aware of the potential dynamic changes of network architectures, so that the learned attack better "leads" the next step than the dynamics-unaware methods when network architecture changes dynamically. Extensive experiments on representative types of adaptive neural networks for both 2D images and 3D point clouds show that our LGM achieves impressive adversarial attack performance compared with the dynamic-unaware attack methods. Code is available at https://github.com/antao97/LGM.

cs.CV

Chaos with Gaussian invariant distribution by quantum-noise random phase feedback

We experimentally present a random phase feedback based on quantum noise to generate a chaotic laser with Gaussian invariant distribution. The quantum noise from vacuum fluctuations is acquired by balanced homodyne detection and injected into a phase modulator to form a random phase feedback. An optical switch using high-speed intensity modulator is employed to reset the chaotic states repeatedly and the time evolutions of intensity statistical distributions of the chaotic states stemming from the initial noise are measured. By the quantum-noise random phase feedback, the transient intensity distributions of the chaotic outputs are improved from asymmetric invariant distributions to Gaussian invariant distributions, and the Gaussian invariant distribution indicates a randomly perturbed dynamical transition from microscopic initial noise to macroscopic stochastic fluctuation. The effects of phase feedback bandwidth and modulation depth on the invariant distributions are investigated experimentally. The chaotic time-delay signature and mean permutation entropy are suppressed to 0.036 and enhanced to 0.999 using the random phase feedback, respectively. The high-quality chaotic laser with Gaussian invariant distribution can be a desired random source for ultrafast random number generation and secure communication.

physics.optics

Vital Sign Monitoring in Dynamic Environment via mmWave Radar and Camera Fusion

Contact-free vital sign monitoring, which uses wireless signals for recognizing human vital signs (i.e, breath and heartbeat), is an attractive solution to health and security. However, the subject's body movement and the change in actual environments can result in inaccurate frequency estimation of heartbeat and respiratory. In this paper, we propose a robust mmWave radar and camera fusion system for monitoring vital signs, which can perform consistently well in dynamic scenarios, e.g., when some people move around the subject to be tracked, or a subject waves his/her arms and marches on the spot. Three major processing modules are developed in the system, to enable robust sensing. Firstly, we utilize a camera to assist a mmWave radar to accurately localize the subjects of interest. Secondly, we exploit the calculated subject position to form transmitting and receiving beamformers, which can improve the reflected power from the targets and weaken the impact of dynamic interference. Thirdly, we propose a weighted multi-channel Variational Mode Decomposition (WMC-VMD) algorithm to separate the weak vital sign signals from the dynamic ones due to subject's body movement. Experimental results show that, the 90${^{th}}$ percentile errors in respiration rate (RR) and heartbeat rate (HR) are less than 0.5 RPM (respirations per minute) and 6 BPM (beats per minute), respectively.

eess.SP

Intercell Moiré Exciton Complexes in Electron Lattices

Excitons, Coulomb-bound electron-hole pairs, play a fundamental role in both optical excitation and correlated phenomena in solids. When an exciton interacts with other quasi-particles, few- and many-body excited states, such as trions, exciton Fermi-polarons, Mahan excitons can appear. Here, we report a new interaction between exciton and charges enabled by unusual quantum confinement in 2D moiré superlattices, which results in novel exciton many-body ground states composed of moiré excitons and correlated electron lattices. Unique to H-stacked (or 60o-twisted) WS2/WSe2 heterobilayer, we found that the interlayer atomic registry and moiré structural reconstruction leads to an interlayer moiré exciton (IME) whose hole in one layer is surrounded by its partner electron's wavefunction spread among three adjacent moiré traps in the other layer. This 3D excitonic structure can enable large in-plane electrical quadrupole moments in addition to the vertical dipole. Upon doping, the electric quadrupole facilitates the binding of IME to the charges in neighboring moiré cells, forming an intercell charged exciton complex. The exciton complex is unveiled by the IME photoluminescence energy jumps when the electron lattices form at both fractional and integer-filled moiré minibands, with replica-like spectral features between successive integer moiré fillings. Our work provides the framework in understanding and engineering emergent exciton many-body states in correlated moiré charge orders.

cond-mat.mes-hall