SearcharxivSearch

arXiv subjects

Xiaoguang Li

Publications and source records attributed to Xiaoguang Li.

At least 19 recordsLinked to original sources

Experience Funnel: A State-Policy Alternating Loop for Self-Evolving Agents

Autonomous agents powered by large language models (LLMs) continuously accumulate experience through interaction, creating an opportunity to improve future behavior through self-evolution. A fundamental challenge is how to transform abundant, task-specific interaction experience into reusable model competence without sacrificing the ability to adapt rapidly to newly observed evidence. Explicit textual states, such as skills and agent harnesses, provide fast, human-readable and editable adaptation, but incur persistent dependence on external context; parametric policies provide compact and reusable competence, but are substantially slower to update. We present \textit{Experience Funnel}, a self-evolving framework that couples fast state adaptation with slow policy consolidation in an alternating loop. Interaction trajectories are first distilled into an explicit textual state, where newly acquired experience can be rapidly incorporated and validated. The framework then selectively identifies state-enabled behavior that remains useful across state revisions and consolidates it into the policy through transition-aware distillation. The updated state--policy pair subsequently generates new rollouts, providing fresh evidence for the next round of state adaptation and policy consolidation. Experiments across diverse agent benchmarks show that \textit{Experience Funnel} consistently improves agent capability over state-only evolution and policy-internalization approaches, while progressively converting useful explicit experience into autonomous policy competence.

cs.CL

Induced electromotive force of a thin metal rod in the alternating electromagnetic field of Helmholtz coil: experimental results and theoretical analysis

We apply a thin metal rod (copper rod) as a probe in the alternating magnetic field generated by a Helmholtz coil, and measure the variation of the induced electromotive force (EMF) on the metal rod at different radial positions of Helmholtz coil. Experimental results show that the induced EMF is zero when the center of the metal rod passes through the center of the cylindrical magnetic field inside the Helmholtz coil. When the metal rod is displaced from the center of field to different radial positions, the induced EMF gradually increases from zero, reaches a maximum at a certain position, and then decreases monotonically as the radial distance continues to increase. At the position where the induced EMF reaches its maximum, the metal rod intersects the radial cross-section of the internal magnetic field of the Helmholtz coil at two points, with a small central portion of the rod located inside the Helmholtz coil and the two end portions outside the coil. To explain the experimental phenomena, we construct four simplified models of the magnetic field distribution based on the actual field distribution of the Helmholtz coil to quantitatively investigate the radial variation of the induced EMF along the metal rod. Our theoretical results show that the four models yield similar results and can all qualitatively explain the experimental data curves, particularly reproducing well the variation trend of the induced EMF and the position of the extremum point. The result of piecewise function fitting model is quantitatively in good agreement with the experimental data. This work is also very much helpful and instructive for undergraduate-level students.

physics.class-ph

OptiMAS: Automatically Optimize Multi-Agent System

Automated evolution of Multi-Agent Systems (MAS) holds significant potential for reducing the manual effort required to design and optimize LLM-based agent architectures. However, extant search-based paradigms face a fundamental trade-off, where an expanded optimization scope exacerbates evolutionary instability, while discrete branch-and-discard search isolates insights across lineages. To address these limitations, we propose a continuous, data-driven optimization paradigm built upon a unified ReAct-based infrastructure that reconciles a broad optimization scope with operational stability. Under this paradigm, we present OptiMAS, a task-agnostic agentic optimizer that leverages textual interaction trajectories and task feedback as loss signals for end-to-end MAS evolution. Equipped with a novel dual-track memory mechanism, OptiMAS sustains performance improvement over extended optimization horizons. Evaluation on four heterogeneous agentic benchmarks with three varying scale and accessibility LLM backbones, demonstrates that OptiMAS consistently achieves competitive or superior accuracy relative to both domain-specialized hand-crafted systems and existing evolutionary methods. Our work establishes a practical milestone toward robust, automated MAS evolution.

cs.MA

Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents

Deep search agents operate over trajectories spanning dozens of steps, yet standard reinforcement learning provides only a single outcome reward per trajectory, which is far too sparse for effective credit assignment. On-policy self-distillation (OPSD) addresses this by using the model's own logits as dense token-level teachers, but extending it to search agents introduces a fundamental tension: the teacher, having access to privileged information such as the correct answer, produces a distribution that differs systematically from the student's exploration-based reasoning, and naive distillation causes the student to inherit this information asymmetry rather than learn better search strategies. We resolve this tension through two contributions. First, we construct Evidence Anchors, which are concise, step-level evidence snippets extracted from the web, as privileged information that captures key reasoning steps without revealing the entire answer path. Second, we propose Step-Level Self-Distilled Policy Optimization (SSPO), which converts teacher-student disagreement into step-level advantage weights within GRPO, applied exclusively to incorrect trajectories. This design decouples what to update from how much to update: the outcome reward determines the direction of policy change, while the teacher modulates its magnitude at each step. Correct trajectories are left untouched, preserving their diversity. On Qwen3-8B, SSPO consistently outperforms GRPO across BrowseComp, GAIA, and FRAMES, surpassing or matching GRPO trained with twice as many gradient steps while adding only about 5 percent overhead per step from a single additional forward pass.

cs.LG

iSMART: An Iterative Sampling-and-Regression Technique for Solving Martingale-Based PDEs

We propose the {\bf i}terative {\bf S}a{\bf M}pling-{\bf A}nd-{\bf R}egression {\bf T}echnique (iSMART) for high-dimensional martingale-based partial differential equations (PDEs) in this paper. By leveraging the $L^2$-projection property of conditional expectation and adopting the stop-gradient technique, iSMART reformulates the continuous martingale condition derived from PDEs into a sequence of tractable sampling-regression problems within an iterative framework. This approach relies solely on standard SDE path simulation and plain squared-error loss minimization, completely bypassing the need for adversarial optimization or nested expectation estimation in previous methods. iSMART accommodates linear, semi-linear, and fully nonlinear martingale-based PDEs within a unified iterative procedure. In particular, for fully nonlinear Hamilton-Jacobi-Bellman (HJB) equations, a freezing-and-compensating technique is introduced to strategically shift a portion of the nonlinearity into the SDE drift, thereby improving the convergence behavior of the iterations. Numerous numerical experiments on linear reaction-diffusion equations with sharp gradients, semilinear Burgers-type equations, and fully nonlinear HJB equations demonstrate the accuracy, efficiency, and robustness of the proposed approach in various high dimensions.

math.NA

TXL Fusion: A Hybrid Machine Learning Framework Integrating Chemical Heuristics and Large Language Models for Topological Materials Discovery

Topological materials, including topological insulators (TIs) and topological semimetals (TSMs), offer promising platforms for quantum, spintronic, and low-dissipation electronic technologies. Their discovery, however, remains constrained by the high cost of first-principles calculations and the slow, resource-intensive nature of experimental validation. Here, we introduce TXL Fusion, a hybrid machine-learning framework that integrates chemically inspired heuristics, physically interpretable numerical descriptors, and large language model (LLM)-derived semantic embeddings for topological-materials classification and discovery. By combining space-group symmetry, electron-count and orbital descriptors, composition-derived topological heuristics, and physics-aware semantic representations, TXL Fusion classifies materials into trivial, TSM, and TI categories with improved overall performance and enhanced minority-class TI recognition relative to conventional descriptor-based baselines. The model further serves as a high-throughput pre-screening tool for external discovery spaces, rapidly prioritizing candidate TSMs before expensive first-principles or experimental validation. Representative TXL-prioritized candidates were subsequently supported by density functional theory (DFT) calculations, demonstrating the practical value of the framework for reducing discovery cost. By uniting symbolic chemical rules, statistical learning, and language-based representations, TXL Fusion provides a scalable and interpretable strategy for accelerating the discovery of next-generation topological and quantum materials.

cond-mat.mtrl-sci

Quantum-inspired Chemical Rule for Discovering Topological Materials

Topological materials exhibit unique electronic structures that underpin both fundamental quantum phenomena and next-generation technologies, yet their discovery remains constrained by the high computational cost of first-principles calculations and the slow, resource-intensive nature of experimental synthesis. Recent machine-learning approaches, such as the heuristic topogivity rule, offer a data-driven pre-screening tool by quantifying each element's intrinsic tendency toward topological behavior. Here, we develop a hybrid quantum-classical neural network (HQCNN) that extends this rule into a quantum-inspired formulation. Within this framework, the HQCNN maps compositional descriptors to quantum probability amplitudes, naturally introducing pairwise inter-element correlations inaccessible to classical heuristics. The physical validity of these correlations is substantiated by constructing an equivalent complex-valued neural network (CVNN), confirming both the consistency and interpretability of the formulation. Retaining the simplicity of chemical reasoning while embedding quantum-native features, our quantum-inspired rule enables efficient and generalizable topological classification. High-throughput screening combined with first-principles (DFT) validation reveals five previously unreported topological compounds, demonstrating the enhanced predictive power and physical insight afforded by quantum-inspired heuristics.

cond-mat.mtrl-sci

KG2Code: Bridging Knowledge Graphs and Large Language Models via Executable Code for Question Answering

Recent research has explored the integration of knowledge graphs (KGs) with large language models (LLMs) to enhance their performance on downstream knowledge-intensive tasks, particularly knowledge graph question answering (KGQA). Existing approaches primarily combine LLMs with KGs through retrieval-augmented generation (RAG)-based, agent-based, and SPARQL-based methods. Although these methods have achieved notable success, they still suffer from several limitations, including structural information loss, unfaithful reasoning, and limited flexibility and generalization. To address these challenges, this paper proposes KG2Code, a novel approach that transforms knowledge graphs into a code-based representation, preserving structural semantics while naturally aligning with the code-aware pretraining of modern LLMs. Based on KG2Code, KG2Code-QA is further introduced as a KGQA framework that formulates KGQA as a code generation task. This formulation enables the generation of verifiable reasoning traces and executable code, thereby substantially mitigating the impact of hallucinations. In addition, an automated pipeline is developed to construct a large-scale, high-quality code corpus for effectively training open-source LLMs on KG2Code-QA. After training, LLMs are able to perform KGQA in zero-shot scenarios. Extensive experiments demonstrate that the proposed approach significantly outperforms existing KG-enhanced LLM methods for KGQA, while exhibiting strong generalization to unseen KGs. The code and data are available at Github.

cs.AI

MTGenRec: An Efficient Distributed Training System for Generative Recommendation Models in Meituan

Recommendation is crucial for both user experience and company revenue in Meituan as a leading lifestyle company, and generative recommendation models (GRMs) are shown to produce quality recommendations recently. However, existing systems are limited by insufficient functionality support and inefficient implementations for training GRMs in industrial scenarios. As such, we introduce MTGenRec as an efficient and scalable system for GRM training. Specifically, to handle real-time insertions/deletions of sparse embeddings, MTGenRec employs dynamic hash tables to replace static ones. To improve training efficiency, MTGenRec conducts dynamic sequence balancing to address the computation load imbalances among GPUs and adopts feature ID deduplication alongside automatic table merging to accelerate embedding lookup. Extensive experiments show that MTGenRec improves training throughput by $1.6\times -- 2.4\times$ while achieving good scalability when running over 100 GPUs. MTGenRec has been deployed for many applications in Meituan and is now handling hundreds of millions of requests on a daily basis. On the delivery platform, we observe a 1.22% growth in user order volume and a 1.31% enhancement in online PV_CTR.

cs.DC

Sharp Exponent of Stable Standing Waves for the Perturbated Hartree Equation

This paper is concerned with the stability of standing waves for the mass-critical Hartree equation with a focusing perturbation by the variational method. The profile decomposition theory is employed to prove the attainability of the cross constrained variational problem, and then the comparison of two cross constrained variational problems is derived. The sharp criteria of blowup, the orbital stability, and strong instability of standing waves without any frequency constraint are obtained. This improves the cross constrained variational argument proposed by Zhang (2005).

math.AP

DLLM Agent: See Farther, Run Faster

Diffusion large language models (DLLMs) have emerged as an alternative to autoregressive (AR) decoding with appealing efficiency and modeling properties, yet their implications for agentic multi-step decision making remain underexplored. We ask a concrete question: when the generation paradigm is changed but the agent framework and supervision are held fixed, do diffusion backbones induce systematically different planning and tool-use behaviors, and do these differences translate into end-to-end efficiency gains? We study this in a controlled setting by instantiating DLLM and AR backbones within the same agent workflow (DeepDiver) and performing matched agent-oriented fine-tuning on the same trajectory data, yielding diffusion-backed DLLM Agents and directly comparable AR agents. Across benchmarks and case studies, we find that, at comparable accuracy, DLLM Agents are on average over 30% faster end to end than AR agents, with some cases exceeding 8x speedup. Conditioned on correct task completion, DLLM Agents also require fewer interaction rounds and tool invocations, consistent with higher planner hit rates that converge earlier to a correct action path with less backtracking. We further identify two practical considerations for deploying diffusion backbones in tool-using agents. First, naive DLLM policies are more prone to structured tool-call failures, necessitating stronger tool-call-specific training to emit valid schemas and arguments. Second, for multi-turn inputs interleaving context and action spans, diffusion-style span corruption requires aligned attention masking to avoid spurious context-action information flow; without such alignment, performance degrades. Finally, we analyze attention dynamics across workflow stages and observe paradigm-specific coordination patterns, suggesting stronger global planning signals in diffusion-backed agents.

cs.CL

UIS-Digger: Towards Comprehensive Research Agent Systems for Real-world Unindexed Information Seeking

Recent advancements in LLM-based information-seeking agents have achieved record-breaking performance on established benchmarks. However, these agents remain heavily reliant on search-engine-indexed knowledge, leaving a critical blind spot: Unindexed Information Seeking (UIS). This paper identifies and explores the UIS problem, where vital information is not captured by search engine crawlers, such as overlooked content, dynamic webpages, and embedded files. Despite its significance, UIS remains an underexplored challenge. To address this gap, we introduce UIS-QA, the first dedicated UIS benchmark, comprising 110 expert-annotated QA pairs. Notably, even state-of-the-art agents experience a drastic performance drop on UIS-QA (e.g., from 70.90 on GAIA and 46.70 on BrowseComp-zh to 24.55 on UIS-QA), underscoring the severity of the problem. To mitigate this, we propose UIS-Digger, a novel multi-agent framework that incorporates dual-mode browsing and enables simultaneous webpage searching and file parsing. With a relatively small $\sim$30B-parameter backbone LLM optimized using SFT and RFT training strategies, UIS-Digger sets a strong baseline at 27.27\%, outperforming systems integrating sophisticated LLMs such as O3 and GPT-4.1. This demonstrates the importance of proactive interaction with unindexed sources for effective and comprehensive information-seeking. Our work not only uncovers a fundamental limitation in current agent evaluation paradigms but also provides the first toolkit for advancing UIS research, defining a new and promising direction for robust information-seeking systems. The dataset has been released at: https://huggingface.co/datasets/UIS-Digger/UIS-QA.

cs.AI

Alternating Subspace Method for Sparse Recovery of Signals

This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible. Numerous renowned algorithms for tackling the compressed sensing problem employ an alternating strategy, which typically involves data matching in one module and denoising in another. We present a novel approach, the Alternating Subspace Method (ASM), which integrates the principles of the greedy methods (e.g., the orthogonal matching pursuit type methods) and the splitting methods (e.g., the approximate message passing type methods). Crucially, ASM enhances the splitting method by achieving fidelity in a subspace-restricted fashion. \textcolor{black}{We reveal that such a restriction strategy guarantees global convergence via proximal residual control and establish its local geometric convergence on the LASSO problem.} Numerical experiments on the LASSO, channel estimation, and dynamic compressed sensing problems demonstrate its high convergence rate and its capacity to incorporate different prior distributions. Overall, the proposed method is promising in terms of efficiency, accuracy, and flexibility, and has the potential to be competitive in different sparse recovery applications.

cs.IT

Gradually Excavating External Knowledge for Implicit Complex Question Answering

Recently, large language models (LLMs) have gained much attention for the emergence of human-comparable capabilities and huge potential. However, for open-domain implicit question-answering problems, LLMs may not be the ultimate solution due to the reasons of: 1) uncovered or out-of-date domain knowledge, 2) one-shot generation and hence restricted comprehensiveness. To this end, this work proposes a gradual knowledge excavation framework for open-domain complex question answering, where LLMs iteratively and actively acquire external information, and then reason based on acquired historical knowledge. Specifically, during each step of the solving process, the model selects an action to execute, such as querying external knowledge or performing a single logical reasoning step, to gradually progress toward a final answer. Our method can effectively leverage plug-and-play external knowledge and dynamically adjust the strategy for solving complex questions. Evaluated on the StrategyQA dataset, our method achieves 78.17% accuracy with less than 6% parameters of its competitors, setting new SOTA for ~10B-scale LLMs.

cs.CL

Robust Single-message Shuffle Differential Privacy Protocol for Accurate Distribution Estimation

Shuffler-based differential privacy (shuffle-DP) is a privacy paradigm providing high utility by involving a shuffler to permute noisy report from users. Existing shuffle-DP protocols mainly focus on the design of shuffler-based categorical frequency oracle (SCFO) for frequency estimation on categorical data. However, numerical data is a more prevalent type and many real-world applications depend on the estimation of data distribution with ordinal nature. In this paper, we study the distribution estimation under pure shuffle model, which is a prevalent shuffle-DP framework without strong security assumptions. We initially attempt to transplant existing SCFOs and the naïve distribution recovery technique to this task, and demonstrate that these baseline protocols cannot simultaneously achieve outstanding performance in three metrics: 1) utility, 2) message complexity; and 3) robustness to data poisoning attacks. Therefore, we further propose a novel single-message \textit{adaptive shuffler-based piecewise} (ASP) protocol with high utility and robustness. In ASP, we first develop a randomizer by parameter optimization using our proposed tighter bound of mutual information. We also design an \textit{Expectation Maximization with Adaptive Smoothing} (EMAS) algorithm to accurately recover distribution with enhanced robustness. To quantify robustness, we propose a new evaluation framework to examine robustness under different attack targets, enabling us to comprehensively understand the protocol resilience under various adversarial scenarios. Extensive experiments demonstrate that ASP outperforms baseline protocols in all three metrics. Especially under small $ε$ values, ASP achieves an order of magnitude improvement in utility with minimal message complexity, and exhibits over threefold robustness compared to baseline methods.

cs.CR

ReThinker: Scientific Reasoning by Rethinking with Guided Reflection and Confidence Control

Expert-level scientific reasoning remains challenging for large language models, particularly on benchmarks such as Humanity's Last Exam (HLE), where rigid tool pipelines, brittle multi-agent coordination, and inefficient test-time scaling often limit performance. We introduce ReThinker, a confidence-aware agentic framework that orchestrates retrieval, tool use, and multi-agent reasoning through a stage-wise Solver-Critic-Selector architecture. Rather than following a fixed pipeline, ReThinker dynamically allocates computation based on model confidence, enabling adaptive tool invocation, guided multi-dimensional reflection, and robust confidence-weighted selection. To support scalable training without human annotation, we further propose a reverse data synthesis pipeline and an adaptive trajectory recycling strategy that transform successful reasoning traces into high-quality supervision. Experiments on HLE, GAIA, and XBench demonstrate that ReThinker consistently outperforms state-of-the-art foundation models with tools and existing deep research systems, achieving state-of-the-art results on expert-level reasoning tasks.

cs.AI

ELAIPBench: A Benchmark for Expert-Level Artificial Intelligence Paper Understanding

While large language models (LLMs) excel at many domain-specific tasks, their ability to deeply comprehend and reason about full-length academic papers remains underexplored. Existing benchmarks often fall short of capturing such depth, either due to surface-level question design or unreliable evaluation metrics. To address this gap, we introduce ELAIPBench, a benchmark curated by domain experts to evaluate LLMs' comprehension of artificial intelligence (AI) research papers. Developed through an incentive-driven, adversarial annotation process, ELAIPBench features 403 multiple-choice questions from 137 papers. It spans three difficulty levels and emphasizes non-trivial reasoning rather than shallow retrieval. Our experiments show that the best-performing LLM achieves an accuracy of only 39.95%, far below human performance. Moreover, we observe that frontier LLMs equipped with a thinking mode or a retrieval-augmented generation (RAG) system fail to improve final results-even harming accuracy due to overthinking or noisy retrieval. These findings underscore the significant gap between current LLM capabilities and genuine comprehension of academic papers.

cs.AI

DeepDiver: Adaptive Search Intensity Scaling via Open-Web Reinforcement Learning

Information seeking demands iterative evidence gathering and reflective reasoning, yet large language models (LLMs) still struggle with it in open-web question answering. Existing prompting and supervised fine-tuning (SFT) methods remain fixed by prompt rules or training corpora, and are usually benchmarked only on well-structured wiki sources, limiting real-world adaptability. We introduce WebPuzzle, a 24k-sample training and 275-sample test benchmark that evaluates information seeking on the live internet, across both wiki and open-domain queries. Leveraging 7k WebPuzzle instances, we develop DeepDiver, a reinforcement-learning (RL) framework that cultivates Search Intensity Scaling (SIS)-an emergent ability to escalate search frequency and depth instead of settling on overconfident, under-evidenced answers. With SIS, Qwen2.5-7B-Instruct and Pangu-7B-Reasoner attain performance on real-web tasks comparable to the 671B-parameter DeepSeek-R1. We detail DeepDiver's curriculum from cold-start SFT to a well designed RL procedure, and show that its seeking policy generalized from closed-ended queries to open-ended generation such as long-form writing. Our results advance adaptive information seeking in LLMs and provide a rigorous benchmark for future work.

cs.CL