SearcharxivSearch

arXiv subjects

Yuqian Wang

Publications and source records attributed to Yuqian Wang.

9 recordsLinked to original sources

StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models

As large language models are increasingly used in data-scarce and evolving task scenarios, few-shot in-context learning (ICL) has become a key paradigm for task adaptation. However, direct ICL often uses a small set of examples without explicitly abstracting task rules, making it sensitive to example construction. In contrast, human learners often reduce such sensitivity by first summarizing task rules from examples and then applying them to new instances. To evaluate this ability, we propose StrategyBench, which selects strategy-inducible tasks from BIG-Bench, constructs reference strategies, and defines evaluation metrics along two dimensions: strategy quality and downstream utility. We further analyze strategy induction from three perspectives: task variation, model configuration, and adaptation setting, covering category-wise differences, generator-executor choices, demonstration design, and SFT-based adaptation. Experiments show that explicit strategy utility differs substantially across task categories and depends on both strategy generation and execution conditions. The benchmark is released at: https://anonymous.4open.science/r/StrategyBench-D53C.

cs.AI

STAMP: Provenance-Guided Credit Assignment for Deep Search Agents

Reinforcement learning for deep-search agents has largely focused on trajectory-level scoring -- outcome correctness, citation-aware rewards, and evidence coverage. Yet the actions that expose supporting documents receive no targeted credit, a gap we call the reward-credit mismatch. We propose STAMP, in which a reference-based verifier judges whether each cited document supports an entity or relation in a training-time evidence graph, and first-exposure attribution traces each supported citation back to the action that first surfaced it. This step credit is injected through sign-preserving advantage modulation, which redistributes advantage across steps without changing the trajectory-level reward or the relative ranking of trajectories within each group. On BrowseComp, BrowseComp-ZH, and xbench-DS, STAMP improves the GRPO baseline by +2.0/+5.5/+3.0 points under matched SFT initialization, training data, and search tools, and composes with both outcome-only and citation-rubric base rewards. Component ablations confirm that the provenance-based credit signal and the sign-preserving advantage modulation each contribute to the gains.

cs.AI

Mind2Report: Expert-Level Commercial Report Synthesis via Cognitive Deep Research Agent

Synthesizing informative commercial reports from massive and noisy web sources is critical for high-stakes business decisions. Although recent deep research agents (DRAs) achieve notable progress, their reports remain limited in quality, reliability, and coverage. These mainly stem from ambiguous intents that cause search drift, retrieved web content that rapidly saturates the context window, and single-pass synthesis that limits report comprehensiveness. In this work, we propose Mind2Report, a cognitive deep research agent that emulates commercial analysts to synthesize expert-level reports. Mind2Report first probes fine-grained commercial intent to establish a structured outline, then recursively explores web sources and distills validated evidence into research memory to preserve context efficiency. Meanwhile, the research memory and outline continuously co-evolve, refining the report structure to avoid rigid initial planning. Finally, Mind2Report iteratively synthesizes the report based on the evolving outline and accumulated evidence. Together, these designs enable reliable and context-efficient long-horizon commercial deep research. To rigorously evaluate commercial DRAs, we further construct QRC-Eval, comprising 200 real-world commercial tasks and a holistic evaluation framework covering report quality, reliability, and coverage. Extensive experiments demonstrate that Mind2Report consistently outperforms leading proprietary and open-source DRAs, while ablations verify the effectiveness of each module and further analyze the challenges they address. We expect this work to advance the development of commercial deep research agents.

cs.CL

DoPI: Doctor-like Proactive Interrogation LLM for Traditional Chinese Medicine

Enhancing interrogation capabilities in Traditional Chinese Medicine (TCM) diagnosis through multi-turn dialogues and knowledge graphs presents a significant challenge for modern AI systems. Current large language models (LLMs), despite their advancements, exhibit notable limitations in medical applications, particularly in conducting effective multi-turn dialogues and proactive questioning. These shortcomings hinder their practical application and effectiveness in simulating real-world diagnostic scenarios. To address these limitations, we propose DoPI, a novel LLM system specifically designed for the TCM domain. The DoPI system introduces a collaborative architecture comprising a guidance model and an expert model. The guidance model conducts multi-turn dialogues with patients and dynamically generates questions based on a knowledge graph to efficiently extract critical symptom information. Simultaneously, the expert model leverages deep TCM expertise to provide final diagnoses and treatment plans. Furthermore, this study constructs a multi-turn doctor-patient dialogue dataset to simulate realistic consultation scenarios and proposes a novel evaluation methodology that does not rely on manually collected real-world consultation data. Experimental results show that the DoPI system achieves an accuracy rate of 84.68 percent in interrogation outcomes, significantly enhancing the model's communication ability during diagnosis while maintaining professional expertise.

cs.AI

Machine Learning Assisted Long-Range Wireless Power Transfer

Near-field magnetic resonance wireless power transfer (WPT) technology has garnered significant attention due to its broad application prospects in medical implants, electric vehicles, and robotics. Addressing the challenges faced by traditional WPT systems in frequency optimization and sensitivity to environmental disturbances, this study innovatively applies the gradient descent optimization algorithm to enhance a system with topological characteristics. Experimental results demonstrate that the machine learning-optimized Su-Schrieffer-Heeger (SSH)-like chain exhibits exceptional performance in transfer efficiency and system robustness. This achievement integrates non-Hermitian physics, topological physics, and machine learning, opening up new avenues and showcasing immense potential for the development of high-performance near-field wave functional devices.

physics.app-ph

ChatSOP: An SOP-Guided MCTS Planning Framework for Controllable LLM Dialogue Agents

Dialogue agents powered by Large Language Models (LLMs) show superior performance in various tasks. Despite the better user understanding and human-like responses, their **lack of controllability** remains a key challenge, often leading to unfocused conversations or task failure. To address this, we introduce Standard Operating Procedure (SOP) to regulate dialogue flow. Specifically, we propose **ChatSOP**, a novel SOP-guided Monte Carlo Tree Search (MCTS) planning framework designed to enhance the controllability of LLM-driven dialogue agents. To enable this, we curate a dataset comprising SOP-annotated multi-scenario dialogues, generated using a semi-automated role-playing system with GPT-4o and validated through strict manual quality control. Additionally, we propose a novel method that integrates Chain of Thought reasoning with supervised fine-tuning for SOP prediction and utilizes SOP-guided Monte Carlo Tree Search for optimal action planning during dialogues. Experimental results demonstrate the effectiveness of our method, such as achieving a 27.95% improvement in action accuracy compared to baseline models based on GPT-3.5 and also showing notable gains for open-source models. Dataset and codes are publicly available.

cs.CL

High-order Uncertain Differential Equation and Its Application to Nuclear Reactors

High-order uncertain differential equation (HUDE) was introduced in literature. But the present method to solve a HUDE is incorrect. In this paper, we will rigorously prove some comparion theorems of high-order differential equations, and present a method to solve a family of HUDE, including parameter estimation and hypothesis test. Then an application to nuclear reactor kinetics is given to illustrate the method.

math.AP

Significant enhancement of magnetic shielding effect by using the composite metamaterial composed of mu-near-zero media and ferrite

The magnetic shield plays an important role in magnetic near-field control. However, the requirements of efficient, ultrathin, lightweight and cheap are still the challenges. Here, we firstly propose a composite metamaterial in which the mu-near-zero media is covered with a ferrite slab. We verify that this structure can enhance the shielding effectiveness in a small area. Furthermore, we optimize the magnetic path by changing the bulk ferrite slab into a patterned slab. In this way, significant shielding effectiveness enhancement can be achieved in a large area. Experimental results show that the maximum shielding effectiveness (SE) of the composite metamaterial with a patterned ferrite is 20.56 dB, which is nearly 19 dB higher than that of a single ferrite slab with the same thickness of the composite metamaterial. The results on the composite metamaterial would be very useful in the applications involving magnetic shielding.

physics.app-ph

Experimental Realization of Near-Field Photonic Routing with All-Electric Metasources

The spatially confined evanescent microwave photonics have been proved to be highly desirable in broad practical scenarios ranging from robust information communications to efficient quantum interactions. However, the feasible applications of these photonics modes are limited due to the lack of fundamental understandings and feasible directional coupling approaches at sub-wavelengths. Here, we experimentally demonstrate the efficient near-field photonic routing achieved in waveguides composed of two kinds of single-negative metamaterials. Without mimicking the polarization features, we propose all-electric near-field metasource in subwavelength scale and exemplify its near-field functions like Janus, Huygens and spin sources, corresponding to time-reversal, parity-time and parity symmetries of its inner degree of freedom. Our work furthers the understandings about optical near-field symmetry and feasible engineering approaches of directional couplings, which would pave the way for promising integrated mircrowave photonics devices.

physics.optics