SearcharxivSearch

arXiv subjects

Tingyu Li

Publications and source records attributed to Tingyu Li.

13 recordsLinked to original sources

Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents

Self-evolving runtime harnesses can substantially improve the capabilities of large language model (LLM) agents and provide a promising paradigm for optimizing agent execution. Existing harness evolution methods typically rely on iterative search, repeatedly evaluating and revising candidate harnesses based on execution feedback from task instances. While this paradigm enables continuous harness optimization, it incurs substantial time overhead due to repeated agent executions and code modifications, and may overfit to observed tasks and specific failure patterns, resulting in degraded generalization to unseen tasks. We identify the lack of principled failure diagnosis as a key bottleneck in harness evolution: an observed failure can reflect either model-specific deficiencies or systematic harness deficiencies, and directly optimizing against individual failures can lead to unnecessary model-specific accommodation. We therefore propose Ecdysis, an efficient and effective framework that distinguishes model-specific accommodation from harness-level repair and biases adaptation toward systematic harness deficiencies by identifying recurring cross-task failure patterns. Ecdysis adopts a batch-level cross-instance failure aggregation paradigm to jointly analyze failure evidence from multiple task instances and further introduces Failure-Driven Collaborative Refinement to diagnose failure causes and iteratively refine harness modification specifications. By combining cross-instance failure analysis with multi-role diagnosis, Ecdysis enables more effective harness evolution with lower training time. Experiments show that Ecdysis achieves up to a 1.84x speedup in harness training compared with existing harness evolution methods, while improving the reasoning accuracy of the resulting harnesses by 18.56%.

cs.SE

Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization

Safety defenses for large language models (LLMs) have been extensively studied, with existing approaches focusing on attack detection and refusal mechanisms. Such fixed-form direct refusal strategies may introduce the risk of prefix injection attacks. Recent work has explored a new direction that leverages humor as an indirect refusal mechanism to mitigate over-refusal in jailbreak scenarios and reduce prefix injection risks. However, this approach implicitly assumes that humorous responses are safe. Whether humorization itself introduces safety risks remains unexplored. To address this issue, we conduct an exploratory study involving over 30,000 real-world agent interaction records and 45 stand-up comedians, revealing practical safety concerns in LLM-based content humorization. Motivated by these findings, we propose \textsc{HumorSafe}, a novel framework for evaluating latent safety risk propagation during humorization. \textsc{HumorSafe} enables LLMs to learn harmful humorization patterns and use them to transform benign content into humorous content with safety risks. Across five frontier LLMs, we find that LLMs can introduce stereotypes and toxicity during humorization. We further propose \textsc{HumorPIA}, a prompt injection attack that exploits latent risks in humor-based defenses. \textsc{HumorPIA} preserves the appearance of safe humorous refusal while covertly injecting harmful content, allowing latent risks to evade existing detection mechanisms. Experiments show that it increases toxicity by 3.14$\times$ while maintaining an apparent safety rate of 97.8\% even under defense settings. Our findings highlight a gap in existing LLM safety evaluations under humorized settings.

cs.CR

Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement

Interleaved thinking, where a unified multimodal model alternates between textual reasoning and visual generation, has shown promise on spatial and physical tasks. However, in complex long-chain scenarios, we identify a fundamental failure mode: generated images diverge from the textual context while subsequent text ignores the visual evidence, causing the two modalities to alternate without genuinely informing each other. We term this Modal Isolation and attribute it to compounding information loss at modality boundaries. We decompose each reasoning cycle into atomic operations and define modality transition loss, quantifying cross-modal hallucination (text-to-image) and visual utilization deficit (image-to-text) at each boundary. We propose MoTiF (Modality Tiransition Fidelity), a two-stage training framework that directly optimizes these transitions: Reflective SFT trains the model to detect and recover from erroneous visual outputs; Flow-GRPO improves image generation fidelity via reinforcement learning. All training signals in MoTiF derive from transition-level fidelity rather than end-task accuracy. Across four visual puzzle benchmarks, this transition-level supervision substantially improves both cross-modal coherence and final task accuracy. The results demonstrate that effective interleaved reasoning requires explicit structural supervision at modality boundaries, not merely scaling or end-task optimization.

cs.CV

A New Origin of the Big Bang from Dark-Sector-Induced Vacuum Decay and Its Gravitational-Wave Signal

We propose a novel scenario for the onset of the thermal Big Bang. In this framework, the inflaton transfers its energy exclusively into a dark sector, leaving the Standard Model (SM) sector temporarily trapped in a false vacuum. As the Hubble expansion rate rapidly decreases, the SM phase transition eventually completes, and the standard thermal Big Bang era commences upon the thermalization of the highly energetic bubble walls. We demonstrate that the large Lorentz boost of these bubble walls, combined with their Hubble-scale macroscopic size, generates distinctive gravitational-wave signatures from the SM vacuum decay. This stochastic gravitational-wave background provides a powerful new probe of the early Universe's expansion history, with a present-day energy density fraction that can reach $\Omega_{\text{GW}} \sim 3\times10^{-8}$.

hep-ph

AcademiClaw: When Students Set Challenges for AI Agents

Benchmarks within the OpenClaw ecosystem have thus far evaluated exclusively assistant-level tasks, leaving the academic-level capabilities of OpenClaw largely unexamined. We introduce AcademiClaw, a bilingual benchmark of 80 complex, long-horizon tasks sourced directly from university students' real academic workflows -- homework, research projects, competitions, and personal projects -- that they found current AI agents unable to solve effectively. Curated from 230 student-submitted candidates through rigorous expert review, the final task set spans 25+ professional domains, ranging from olympiad-level mathematics and linguistics problems to GPU-intensive reinforcement learning and full-stack system debugging, with 16 tasks requiring CUDA GPU execution. Each task executes in an isolated Docker sandbox and is scored on task completion by multi-dimensional rubrics combining six complementary techniques, with an independent five-category safety audit providing additional behavioral analysis. Experiments on six frontier models show that even the best achieves only a 55\% pass rate. Further analysis uncovers sharp capability boundaries across task domains, divergent behavioral strategies among models, and a disconnect between token consumption and output quality, providing fine-grained diagnostic signals beyond what aggregate metrics reveal. We hope that AcademiClaw and its open-sourced data and code can serve as a useful resource for the OpenClaw community, driving progress toward agents that are more capable and versatile across the full breadth of real-world academic demands. All data and code are available at https://github.com/GAIR-NLP/AcademiClaw.

cs.AI

Gravitational Waves and Primordial Black Holes produced by Dark Meta Stable Vacuum Decay

Inspired by string theory and cosmological constant problem, it is plausible that the Universe's vacuum structure is characterized by a landscape of metastable vacua. The existence of dark matter and dark energy further suggests that the dark sector may inhabit its own "dark landscape". If the dark vacuum is metastable, bubbles of lower-energy phases can nucleate at an approximately constant rate. Because the Hubble expansion rate is monotonically non-increasing with cosmic time, such nucleation can eventually lead to percolation and completion of a dark-sector phase transition. In this work, we investigate the phenomenological consequences of this transition, focusing on the resulting stochastic gravitational-wave background and the potential formation of primordial black holes. We find that the gravitational wave spectrum peaks at $k_{\mathrm{peak}}=3.1 H_{\mathrm{PT}}$, with an amplitude $\Omega_{\mathrm{GW}}^{\mathrm{peak}}\simeq1.5 \Omega_\gamma(\Delta\rho/\rho_{\mathrm{tot}})^2$. Furthermore, the formation of primordial black holes is suppressed due to $\Delta N_{\mathrm{eff}}$ constraint.

hep-ph

Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language Reasoning

Recent vision-language models (VLMs) achieve remarkable reasoning through reinforcement learning (RL), which provides a feasible solution for realizing continuous self-evolving large vision-language models (LVLMs) in the era of experience. However, RL for VLMs requires abundant high-quality multimodal data, especially challenging in specialized domains like chemistry, earth sciences, and multimodal mathematics. Existing strategies such as synthetic data and self-rewarding mechanisms suffer from limited distributions and alignment difficulties, ultimately causing reward hacking: models exploit high-reward patterns, collapsing policy entropy and destabilizing training. We propose DoGe (Decouple to Generalize), a dual-decoupling framework that guides models to first learn from context rather than problem solving by refocusing on the problem context scenarios overlooked by synthetic data methods. By decoupling learning process into dual components (Thinker and Solver), we reasonably quantify the reward signals of this process and propose a two-stage RL post-training approach from freely exploring context to practically solving tasks. Second, to increase the diversity of training data, DoGe constructs an evolving curriculum learning pipeline: an expanded native domain knowledge corpus and an iteratively evolving seed problems pool. Experiments show that our method consistently outperforms the baseline across various benchmarks, providing a scalable pathway for realizing self-evolving LVLMs.

cs.AI

Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI Collaboration

Agents built on large language models (LLMs) have excelled in turn-by-turn human-AI collaboration but struggle with simultaneous tasks requiring real-time interaction. Latency issues and the challenge of inferring variable human strategies hinder their ability to make autonomous decisions without explicit instructions. Through experiments with current independent System 1 and System 2 methods, we validate the necessity of using Dual Process Theory (DPT) in real-time tasks. We propose DPT-Agent, a novel language agent framework that integrates System 1 and System 2 for efficient real-time simultaneous human-AI collaboration. DPT-Agent's System 1 uses a Finite-state Machine (FSM) and code-as-policy for fast, intuitive, and controllable decision-making. DPT-Agent's System 2 integrates Theory of Mind (ToM) and asynchronous reflection to infer human intentions and perform reasoning-based autonomous decisions. We demonstrate the effectiveness of DPT-Agent through further experiments with rule-based agents and human collaborators, showing significant improvements over mainstream LLM-based frameworks. DPT-Agent can effectively help LLMs convert correct slow thinking and reasoning into executable actions, thereby improving performance. To the best of our knowledge, DPT-Agent is the first language agent framework that achieves successful real-time simultaneous human-AI collaboration autonomously. Code of DPT-Agent can be found in https://github.com/sjtu-marl/DPT-Agent.

cs.AI

Identifying average causal effect in regression discontinuity design with auxiliary data

Regression discontinuity designs are widely used when treatment assignment is determined by whether a running variable exceeds a predefined threshold. However, most research focuses on estimating local causal effects at the threshold, leaving the challenge of identifying treatment effects away from the cutoff largely unaddressed. The primary difficulty in this context is that the treatment assignment is deterministically defined by the running variable, violating the commonly assumed positivity assumption. In this paper, we introduce a novel framework for identifying the average causal effect in regression discontinuity designs. Our approach assumes the existence of an auxiliary variable for which the running variable can be seen as a surrogate, and an additional dataset that consists of the running variable and the auxiliary variable alongside the traditional regression discontinuity design setup. Under this framework, we propose three estimation methods for the ATE, which resembles the outcome regression, inverse propensity weighted and doubly robust estimators in classical causal inference literature. Asymptotically valid inference procedures are also provided. To demonstrate the practical application of our method, simulations are conducted to show the good performance of our methods; besides, we use the proposed methods to assess the causal effects of vitamin A supplementation on the severity of autism spectrum disorders in children, where a positive effect is found but with no statistical significance.

stat.ME

Dark Photon Dark Matter and Low-Frequency Gravitational Wave Detection with Gaia-like Astrometry

Astrometric surveys offer us a method to search for elusive cosmic signatures, such as ultralight dark photon dark matter and gravitational waves, by observing the deflection to the apparent positions of the stars. The detection capabilities of such surveys rapidly decrease at low frequencies, because the signals become hardly distinguishable from the background motion of stars. In this work, we find that the background motion can be well described by a linear model over time, based on which we propose a linear background subtraction scheme. Compared to the conventional quadratic subtraction, the advantage of linear subtraction emerges within the frequency range below $6 \times 10^{-9}~{\rm Hz}$. Taking dark photons with purely gravitational interactions, dark photons with additional $U(1)_{B}$ or $U(1)_{B-L}$ gauge interactions, and low-frequency gravitational waves as examples, we illustrate that the linear subtraction scheme can result in an enhancement of more than one order of magnitude in the exclusion limits of Gaia-like experiments in the low-frequency range.

hep-ph

Characterizing current noise of commercial constant-current sources by using of an optically-pumped rubidium atomic magnetometer

This paper introduces a method for characterizing the current noise of commercial constant-current sources(CCSs) using a free-induction-decay(FID) type optically-pumped rubidium atomic magnetometer driven by a radio-frequency(RF) magnetic field. We convert the sensitivity of the atomic magnetometer into the current noise of CCS by calibrating the coil constant. At the same time, the current noise characteristics of six typical commercial low-noise CCSs are compared. The current noise level of the KeySight Model B2961A is the lowest among the six tested CCSs, which is 36.233 0.022 nA / Hz1/2 at 1-25 Hz and 133.905 0.080 nA / Hz1/2 at 1-100 Hz respectively. The sensitivity of atomic magnetometer is dependent on the current noise level of the CCS. The CCS with low noise is of great significance for high-sensitivity atomic magnetometer. The research provides an important reference for promoting the development of high precision CCS, metrology and basic physics research.

physics.atom-ph

Large-Scale Integrated Flexible Tactile Sensor Array for Sensitive Smart Robotic Touch

In the long pursuit of smart robotics, it has been envisioned to empower robots with human-like senses, especially vision and touch. While tremendous progress has been made in image sensors and computer vision over the past decades, the tactile sense abilities are lagging behind due to the lack of large-scale flexible tactile sensor array with high sensitivity, high spatial resolution, and fast response. In this work, we have demonstrated a 64x64 flexible tactile sensor array with a record-high spatial resolution of 0.9 mm (equivalently 28.2 pixels per inch), by integrating a high-performance piezoresistive film (PRF) with a large-area active matrix of carbon nanotube thin-film transistors. PRF with self-formed microstructures exhibited high pressure-sensitivity of ~385 kPa-1 for MWCNTs concentration of 6%, while the 14% one exhibited fast response time of ~3 ms, good linearity, broad detection range beyond 1400 kPa, and excellent cyclability over 3000 cycles. Using this fully integrated tactile sensor array, the footprint maps of an artificial honeybee were clearly identified. Furthermore, we hardware-implemented a smart tactile system by integrating the PRF-based sensor array with a memristor-based computing-in-memory chip to record and recognize handwritten digits and Chinese calligraphy, achieving high classification accuracies of 98.8% and 97.3% in hardware, respectively. The integration of sensor networks with deep learning hardware may enable edge or near-sensor computing with significantly reduced power consumption and latency. Our work could pave the road to building large-scale intelligent sensor networks for next-generation smart robotics.

cond-mat.mtrl-sci

Intercomparison Study of Time and Frequency Transfer between VLBI and Other Techniques (GPS, ETS8(TCE), TW(DPN) and DMTD)

We carried out the intercomparison experiments between VLBI and other techniques to show the capability of VLBI time and frequency transfer by using the current geodetic VLBI technique and facilities as the summary of the experiments that we carried out since 2007. The results from the two different types of experiments show that the VLBI is more stable than GPS but is slightly noisier than two new two-way techniques (TW(DPN), ETS8(TCE)), and VLBI can measure the correct time difference as same as ETS8(TCE).

astro-ph.IM