SearcharxivSearch

arXiv subjects

Dong Yan

Publications and source records attributed to Dong Yan.

At least 19 recordsLinked to original sources

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominantly adopt independent evaluation. Consequently, the behavior of self-evolving agents in realistic streaming settings, where agents adapt to diverse and complex task streams, remains poorly understood. To address this gap, we introduce AgentStream, a unified framework that evaluates self-evolving agents spanning diverse evolution components by organizing agentic benchmarks into a configurable task stream and instantiating the \texttt{Isolated}, \texttt{Sequential}, and \texttt{Interleaved} streaming scenarios at test time, which progressively vary the scope and domain composition of the stream. Over these scenarios, we combinatorially evaluate five representative self-evolving methods across three frontier foundation models, disentangling how model capability, method architecture, and streaming scenario jointly shape self-evolution. Our results show that self-evolution reliability varies across streaming scenarios, the benefit of self-evolution is gated by model capability and non-monotonic in model strength, and no single method dominates across models and scenarios. These findings offer concrete guidance for selecting self-evolving methods across models and streaming scenarios. Overall, we advocate that self-evolving agents should be evaluated under realistic task streams rather than isolated single-task settings.

cs.AI

Asymmetric Nonlinear Return Extrapolation and Optimal Portfolio Choice under Stochastic Volatility

We extend the return extrapolation framework of Atmaz (2022) to incorporate two behaviorally realistic features absent from the linear benchmark: saturation in belief updating and asymmetry between gains and losses. We introduce a smooth, nonlinear, asymmetric extrapolation function and characterize the optimal portfolio of a CRRA investor under Heston (1993) stochastic volatility as the sum of a sentiment-distorted myopic demand, a variance hedging demand, and a sentiment hedging demand. The resulting semilinear Hamilton-Jacobi-Bellman equation is solved by two independent numerical methods, a finite-difference ADI scheme with time-step policy iteration and a deep learning-driven iterative scheme. The model generates four investor-level behavioral anomalies: asymmetric responses to gains and losses, attenuated reactions at extremes, excess trading volume, and welfare loss rising with the strength of extrapolation, each of which maps onto documented empirical patterns. Its central finding is that saturation acts as an endogenous correction mechanism: at the same local slope at the origin, the asymmetric nonlinear extrapolator carries a smaller welfare loss than a linear one.

q-fin.PM

Chiral Quantum Entanglement Transfer with Giant Atoms

We investigate entanglement transfer in a multi-giant-atom waveguide system. By tailoring chiral spontaneous emission and exploiting dark-state dynamics, the setup enables perfect, unidirectional sequential or selective transfer of quantum states and their associated entanglement. The distance between two entangled atoms, i.e., the entanglement length, can be dynamically adjusted, allowing robust conversion between long-range and short-range entanglement during propagation. The system inherently converges to a dark state, guaranteeing high-fidelity directional transfer. When the additional phase is modulated as a periodic piecewise function, spatially separated giant atoms exhibit stable, nearly lossless state exchange and maintain steady entanglement even under non-Markovian conditions. This behavior mimics conventional braided architectures without suffering from propagation delays or spatial restrictions. Our proposal offers a scalable pathway for continuous long-distance entanglement transport and resilient state exchange in quantum networks.

quant-ph

FunctionEvolve: Structure-Guided Symbolic Regression with LLMs

Symbolic regression aims to uncover explicit scientific laws from data. Recent methods use LLMs to guide mutation from background text, which is more directed than random genetic programming. However, exact symbolic recovery requires both semantic guidance and explicit structure, so that domain-informed search are carried out through valid symbolic representation. Current LLM-driven systems remain structure-blind: they select among opaque candidates, lack explicit mechanisms for local mutation, and rely on brittle coefficient fitting that can undervalue correct skeletons. We propose FunctionEvolve, an evolutionary framework using expression trees to organize the whole search: structural summaries promote diverse parent selection, local tree edits preserve useful subexpressions, and structure-aware fitting decomposes, constrains, and simplifies coefficients for more reliable scoring. It uses only elementary function families, without additional domain-specific rules limiting generalization. On the 129-task synthetic subset of LLM-SRBench, FunctionEvolve with \emph{Claude Opus 4.6} recovers 107 exact forms, reaching 82.9% SA@50, 4.5x above same-backbone baselines, and 55.8% SA@1, 3.6x above the strongest previously published top-1 result. Ablations show that structure-visible search is central to reliable recovery, with LLM-guided refinements and structure-aware coefficient optimization serving as essential proposal and scoring mechanisms. We also audit the benchmark and show that collinearity in its materials-science subset creates identifiability issues.

cs.LG

Controlling entanglement by phase engineering in giant-atom waveguide

We investigate the entanglement dynamics of two giant atoms coupled to a common waveguide. By introducing additional phase modulation at each coupling point, every photon propagation path is jointly controlled by two distinct coupling phases, enabling precise and flexible manipulation of the entanglement evolution. This phase engineering induces destructive interference among different paths, leading to entanglement dynamics in nested giant atoms that become equivalent to those of small atoms, as well as dynamical equivalence between separated and braided configurations. Furthermore, the proposed scheme significantly enhances the robustness of entanglement against variations in the phase shift, offering a practical route to generate stable entanglement and enabling quantum devices with programmable propagation and controllable memory effects.

quant-ph

What If Consensus Lies? Selective-Complementary Reinforcement Learning at Test Time

Test-Time Reinforcement Learning (TTRL) enables Large Language Models (LLMs) to enhance reasoning capabilities on unlabeled test streams by deriving pseudo-rewards from majority voting consensus. However, existing TTRL methods rely exclusively on positive pseudo-labeling strategies. Such reliance becomes vulnerable under challenging scenarios where answer distributions are highly dispersed, resulting in weak consensus that inadvertently reinforces incorrect trajectories as supervision signals. In this paper, we propose SCRL (Selective-Complementary Reinforcement Learning), a robust test-time reinforcement learning framework that effectively mitigates label noise amplification. SCRL develops Selective Positive Pseudo-Labeling, which enforces strict consensus criteria to filter unreliable majorities. Complementarily, SCRL introduces Entropy-Gated Negative Pseudo-Labeling, the first negative supervision mechanism in TTRL, to reliably prune incorrect trajectories based on generation uncertainty. Extensive experiments on multiple reasoning benchmarks demonstrate that SCRL achieves substantial improvements over baselines, while maintaining robust generalization and training stability under constrained rollout budgets. Our code is available at https://github.com/Jasper-Yan/SCRL.

cs.LG

Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs

Recent studies have shown that large language models (LLMs) can infer private user attributes (e.g., age, location, gender) from user-generated text shared online, enabling rapid and large-scale privacy breaches. Existing anonymization-based defenses are coarse-grained, lacking word-level precision in anonymizing privacy-leaking elements. Moreover, they are inherently limited as altering user text to hide sensitive cues still allows attribute inference to occur through models' reasoning capabilities. To address these limitations, we propose a unified defense framework that combines fine-grained anonymization (TRACE) with inference-preventing optimization (RPS). TRACE leverages attention mechanisms and inference chain generation to identify and anonymize privacy-leaking textual elements, while RPS employs a lightweight two-stage optimization strategy to induce model rejection behaviors, thereby preventing attribute inference. Evaluations across diverse LLMs show that TRACE-RPS reduces attribute inference accuracy from around 50\% to below 5\% on open-source models. In addition, our approach offers strong cross-model generalization, prompt-variation robustness, and utility-privacy tradeoffs. Our code is available at https://github.com/Jasper-Yan/TRACE-RPS.

cs.CR

Chaos in the near-horizon dynamics of the dyonic $\rm{AdS_4}$-Reissner-Nordstr\"{o}m black hole

We investigate the chaos in the dynamics of a probe massless particle confined by the harmonic potential near the horizon of the dyonic $\rm{AdS_4}$-Reissner-Nordstr\"om black hole. The total energy of the particle, chemical potential and magnetic field in this system serving as independently adjustable parameters tune nonlinearity and phase-space structure. By analyzing the trajectories on the Poincar\'e section and evaluating the Lyapunov exponents, we obtain the dynamical phase diagrams of the chaos and find their counteracting regulatory role: at low energy, chaos is enhanced and the Lyapunov exponent $\lambda_L$ violates its upper bound (i.e. surface gravity) in the extremal black hole limit(combined paramete $\Gamma=3$); at high energy, the same extremal limit suppresses chaos, with $\lambda_L$ dropping to zero and a regular dynamical corridor emerging along $\Gamma=3$ in the dynamical phase diagrams. These results establish a direct mapping between black hole thermodynamics and microscopic chaos, offering new insights into the AdS/QCD correspondence and nonlinear dynamics in strongly curved spacetimes.

hep-th

Unidirectional reflection lasing based on destructive interference and Bragg scattering modulation in defective atomic lattice

The novel and ingenious scheme we propose for achieving unidirectional reflection lasing (URL) involves integrating a one-dimensional (1D) defective atomic lattice with a coherent gain atomic system. Its physical essence lies in the fact that the right-side reflectivity is drastically reduced due to the destructive interference between primary and secondary reflections, whereas on the left-side primary reflection is effectively suppressed and the secondary reflection is efficiently enhanced, ultimately reaching the lasing threshold. Through numerical results and further analyses, we have elucidated how to precisely tailor the lattice parameters and coupling fields to control destructive interference point (DIP), thereby realizing URL and enabling its active modulation. Our scheme is experimentally feasible and not only effectively circumvents the stringent conditions faced in directly realizing URL, providing a new pathway, but also beneficial for integrating active photonic devices into compact quantum networks and may improve the efficiency of optical information transmission.

physics.optics

Unidirectional spectral singularity lasing in a defective atomic lattice

We propose an efficient scheme for achieving mode-tunable unidirectional reflection lasing (URL) by establishing a coherent gain atomic system to amplify the probe field and ingeniously designing the one-dimensional (1D) defective atomic lattice. This lattice not only replaces the resonant cavity to provide a distributed feedback mechanism but also breaks the spatial symmetry of the probe susceptibility. Correspondingly, the URL can be characterized by a non-Hermitian degenerate spectral singularity (NHDSS), where the two eigenvalues of the inverse scattering matrix are engineered to satisfy $\lambda _{{S^{-1}}}^{+}\simeq \lambda _{{S^{-1}}}^{-}\rightarrow 0$. This intriguing NHDSS depends on the probe susceptibility and the Bragg condition, both of which can be modulated by adjusting the external optical field and lattice structure, rendering the scheme experimentally feasible. Our approach achieves both nonreciprocity and lasing oscillation in a single system, significantly enhancing the efficiency of optical information transmission and facilitating the integration of active photonic devices into compact quantum networks.

physics.optics

High-Power Dual-Channel Field Chamber for High-Frequency Magnetic Neuromodulation

Several novel methods, including magnetogenetics and magnetoelectric stimulation, use high frequency alternating magnetic fields to precisely manipulate neural activity. To quantify the behavioral effects of such interventions in a freely moving mouse, we developed a dual-channel magnetic chamber, specifically designed for rate-sensitive magnetothermal-genetic stimulation, and adaptable for other uses of alternating magnetic fields. Through an optimized coil design, the system allows independent control of two spatially orthogonal uniform magnetic fields delivered at different frequencies within a 10 cm x 10 cm x 6 cm chamber. The two channels have nominal frequencies of 50 and 550 kHz with peak magnetic field strengths of 88 and 12.5 mT, achieved with resonant coil drives having peak voltages of 1.6 and 1.8 kV and currents of 1.0 and 0.26 kA, respectively. Additionally, a liquid cooling system enables magnetic field generation for second-level duration, and an observation port and camera allow video capture of the animal's behavior within the chamber. The system generates high-amplitude magnetic fields across two widely separated frequency channels with negligible interference (< 1%). Relatively uniform magnetic field distribution (+/-10% across 94% of the chamber volume) is maintained throughout the chamber, and temperature increase of the inner side of the coil enclosure during the operation is limited to < 0.35 {\deg}C/s to ensure in vivo safety. Using cobalt-doped and undoped iron oxide nanoparticles, we demonstrate channel-specific heating rates of 3.5 {\deg}C/s and 1.5 {\deg}C/s, respectively, validating frequency-selectivity. Both channels can run continuously for four seconds stably.

eess.SY

Portfolio selection with exogenous and endogenous transaction costs under a two-factor stochastic volatility model

In this paper, we investigate a portfolio selection problem with transaction costs under a two-factor stochastic volatility structure, where volatility follows a mean-reverting process with a stochastic mean-reversion level. The model incorporates both proportional exogenous transaction costs and endogenous costs modeled by a stochastic liquidity risk process. Using an option-implied approach, we extract an S-shaped utility function that reflects investor behavior and apply its concave envelope transformation to handle the non-concavity. The resulting problem reduces to solving a five-dimensional nonlinear Hamilton-Jacobi-Bellman equation. We employ a deep learning-based policy iteration scheme to numerically compute the value function and the optimal policy. Numerical experiments are conducted to analyze how both types of transaction costs and stochastic volatility affect optimal investment decisions.

q-fin.MF

Mission Impossible: Feedback-Guided Dynamic Interactive Planning for Improving Reasoning on LLMs

Recent advancements in language agents have led to significant improvements in multi-hop reasoning tasks. However, existing approaches often struggle with handling open-domain problems, which require massive information retrieval due to their reliance on a fixed sequence of actions. To address this, we propose Feedback-Guided Dynamic Interactive Planning (FGDIP), a novel framework tailored to enhance reasoning in LLMs by utilizing dynamic and adaptive strategies for information exploration in open-domain multi-hop reasoning tasks. Our approach begins by identifying key entities relevant to the problem, which serve as the initial nodes in the reasoning process. From these initial nodes, we then generate reasoning child nodes with the process being refined through a combination of historical error analysis and real-time feedback, which allows the framework to dynamically adjust and optimize its reasoning strategies. By integrating depth-first search with an innovative node generation technique, our framework adapts based on both prior error paths and concurrently generated nodes at the same hierarchical level. This dynamic strategy effectively expands the search space while ensuring the reasoning process systematically converges toward accurate solutions. Experimental results show that FGDIP achieved up to 54.47% F1 score on the HotpotQA dataset and 70.05% on the StrategyQA dataset, surpassing the best baseline by 5.03% and 7.25% respectively, highlighting its versatility and potential to enhance language agents in multi-hop reasoning tasks.

cs.CL

A deep learning-driven iterative scheme for high-dimensional HJB equations in portfolio selection with exogenous and endogenous costs

In this paper, we first conduct a study of the portfolio selection problem, incorporating both exogenous (proportional) and endogenous (resulting from liquidity risk, characterized by a stochastic process) transaction costs through the utility-based approach. We also consider the intrinsic relationship between these two types of costs. To address the associated nonlinear two-dimensional Hamilton-Jacobi-Bellman (HJB) equation, we propose an innovative deep learning-driven policy iteration scheme with three key advantages: i) it has the potential to address the curse of dimensionality; ii) it is adaptable to problems involving high-dimensional control spaces; iii) it eliminates truncation errors. The numerical analysis of the proposed scheme, including convergence analysis in a general setting, is also discussed. To illustrate the impact of these two types of transaction costs on portfolio choice, we conduct through numerical experiments using three typical utility functions.

q-fin.MF

Pricing American options with exogenous and endogenous transaction costs

We study an American option pricing problem with liquidity risks and transaction fees. As endogenous transaction costs, liquidity risks of the underlying asset are modeled by a mean-reverting process. Transaction fees are exogenous transaction costs and are assumed to be proportional to the trading amount, with the long-run liquidity level depending on the proportional transaction costs rate. Two nonlinear partial differential equations are established to characterize the option values for the holder and the writer, respectively. To illustrate the impact of these transaction costs on option prices and optimal exercise prices, we apply the alternating direction implicit method to solve the linear complementarity problem numerically. Finally, we conduct model calibration from market data via maximum likelihood estimation, and find that our model incorporating liquidity risks outperforms the Leland model significantly.

q-fin.MF

Unidirectional lasing via vacuum induced coherent in defective atomic lattice

We skillfully utilized vacuum induced coherence to amplify the probe light, and then successfully achieved both nonreciprocal reflection and lasing oscillation in a single physical system by leveraging the distributed feedback and spatial symmetry breaking effect of the one-dimensional defective atomic lattice. This innovative scheme for realizing unidirectional reflection lasing (URL) is based on both non-Hermitian degeneracy and spectral singularity (NHDSS, means $\lambda_{+}^{-1}\simeq\lambda_{-}^{-1}\rightarrow0$). Therefore, we analyze the modulation of parameters such as the lattice structure and external optical fields in this system to find NHDSS point, and further verified the conditions for its occurrence by solving the transcendental equation of susceptibility satisfying the NHDSS point, as well as analyzed its physical essence. Our mechanism is not only beneficial for the integration of photonic devices in quantum networks, but also greatly improves the efficiency of optical information transmission.

physics.optics

Nonreciprocal Entanglement by Dynamically Encircling a Nexus

Nonreciprocal entanglement, characterized by inherently robust operation, is a cornerstone for quantum information processing and communications. However, it remains a great challenge to achieve nonreciprocal entanglement characterized by stability and robustness against environmental fluctuations. Here, we propose a universal nonlinear mechanism to engineer magnetic-free nonreciprocity in dissipative optomechanics by utilizing bistability, a phenomenon ubiquitous across nonlinear physical systems. By dynamically encircling the nexus of bistability, a cusp converged by the bistable surfaces, we obtain nonreciprocal displacement and then utilize it to achieve robust nonreciprocal entanglement. Owing to the unique landscape of bistability, our nonreciprocal displacement and entanglements exhibit stability and robustness through closed-loop operations. Our work presents a foundational framework for leveraging nonlinearity to achieve nonreciprocal quantum information processing. It paves new avenues for exploring nonreciprocal quantum information processing and designing backaction-immune quantum metrology with nonlinearity.

quant-ph

STAIR: Improving Safety Alignment with Introspective Reasoning

Ensuring the safety and harmlessness of Large Language Models (LLMs) has become equally critical as their performance in applications. However, existing safety alignment methods typically suffer from safety-performance trade-offs and the susceptibility to jailbreak attacks, primarily due to their reliance on direct refusals for malicious queries. In this paper, we propose STAIR, a novel framework that integrates SafeTy Alignment with Itrospective Reasoning. We enable LLMs to identify safety risks through step-by-step analysis by self-improving chain-of-thought (CoT) reasoning with safety awareness. STAIR first equips the model with a structured reasoning capability and then advances safety alignment via iterative preference optimization on step-level reasoning data generated using our newly proposed Safety-Informed Monte Carlo Tree Search (SI-MCTS). We further train a process reward model on this data to guide test-time searches for improved responses. Extensive experiments show that STAIR effectively mitigates harmful outputs while better preserving helpfulness, compared to instinctive alignment strategies. With test-time scaling, STAIR achieves a safety performance comparable to Claude-3.5 against popular jailbreak attacks. Relevant resources in this work are available at https://github.com/thu-ml/STAIR.

cs.CL