SearcharxivSearch

arXiv subjects

Yuefeng Huang

Publications and source records attributed to Yuefeng Huang.

6 recordsLinked to original sources

Dep-Search: Learning Dependency-Aware Reasoning Traces with Persistent Memory

Large Language Models (LLMs) have demonstrated remarkable capabilities in complex reasoning tasks, particularly when augmented with search mechanisms that enable systematic exploration of external knowledge bases. The field has evolved from traditional retrieval-augmented generation (RAG) frameworks to more sophisticated search-based frameworks that orchestrate multi-step reasoning through explicit search strategies. However, existing search frameworks still rely heavily on implicit natural language reasoning to determine search strategies and how to leverage retrieved information across reasoning steps. This reliance on implicit reasoning creates fundamental challenges for managing dependencies between sub-questions, efficiently reusing previously retrieved knowledge, and learning optimal search strategies through reinforcement learning. To address these limitations, we propose Dep-Search, a dependency-aware search framework that advances beyond existing search frameworks by integrating structured reasoning, retrieval, and persistent memory through GRPO. Dep-Search introduces explicit control mechanisms that enable the model to decompose questions with dependency relationships, retrieve information when needed, access previously stored knowledge from memory, and summarize long reasoning contexts into reusable memory entries. Through extensive experiments on seven diverse question answering datasets, we demonstrate that Dep-Search significantly enhances LLMs' ability to tackle complex multi-hop reasoning tasks, achieving substantial improvements over strong baselines across different model scales.

cs.CL

ACEBench: Who Wins the Match Point in Tool Usage?

Large Language Models (LLMs) have demonstrated significant potential in decision-making and reasoning, particularly when integrated with various tools to effectively solve complex problems. However, existing benchmarks for evaluating LLMs' tool usage face several limitations: (1) limited evaluation scenarios, often lacking assessments in real multi-turn dialogue contexts; (2) narrow evaluation dimensions, with insufficient detailed assessments of how LLMs use tools; and (3) reliance on LLMs or real API executions for evaluation, which introduces significant overhead. To address these challenges, we introduce ACEBench, a comprehensive benchmark for assessing tool usage in LLMs. ACEBench categorizes data into three primary types based on evaluation methodology: Normal, Special, and Agent. "Normal" evaluates tool usage in basic scenarios; "Special" evaluates tool usage in situations with ambiguous or incomplete instructions; "Agent" evaluates tool usage through multi-agent interactions to simulate real-world, multi-turn dialogues. We conducted extensive experiments using ACEBench, analyzing various LLMs in-depth and providing a more granular examination of error causes across different data types.

cs.CL

ToolACE-DEV: Self-Improving Tool Learning via Decomposition and EVolution

The tool-using capability of large language models (LLMs) enables them to access up-to-date external information and handle complex tasks. Current approaches to enhancing this capability primarily rely on distilling advanced models by data synthesis. However, this method incurs significant costs associated with advanced model usage and often results in data compatibility issues, led by the high discrepancy in the knowledge scope between the advanced model and the target model. To address these challenges, we propose ToolACE-DEV, a self-improving framework for tool learning. First, we decompose the tool-learning objective into sub-tasks that enhance basic tool-making and tool-using abilities. Then, we introduce a self-evolving paradigm that allows lightweight models to self-improve, reducing reliance on advanced LLMs. Extensive experiments validate the effectiveness of our approach across models of varying scales and architectures.

cs.CL

Advancing and Benchmarking Personalized Tool Invocation for LLMs

Tool invocation is a crucial mechanism for extending the capabilities of Large Language Models (LLMs) and has recently garnered significant attention. It enables LLMs to solve complex problems through tool calls while accessing up-to-date world knowledge. However, existing work primarily focuses on the fundamental ability of LLMs to invoke tools for problem-solving, without considering personalized constraints in tool invocation. In this work, we introduce the concept of Personalized Tool Invocation and define two key tasks: Tool Preference and Profile-dependent Query. Tool Preference addresses user preferences when selecting among functionally similar tools, while Profile-dependent Query considers cases where a user query lacks certain tool parameters, requiring the model to infer them from the user profile. To tackle these challenges, we propose PTool, a data synthesis framework designed for personalized tool invocation. Additionally, we construct \textbf{PTBench}, the first benchmark for evaluating personalized tool invocation. We then fine-tune various open-source models, demonstrating the effectiveness of our framework and providing valuable insights. Our benchmark is public at https://github.com/hyfshadow/PTBench.

cs.CL

Heroin addiction hijacks the Nucleus Accumbens: craving and reactivity to naturalistic stimuli

Drug-related cues hijack attention away from alternative reinforcers in drug addiction, inducing craving and motivating drug-seeking. However, the neural correlates underlying this biased processing, its expression in the real-world, and its relationship to cue-induced craving are not fully established, especially in opioid addiction. Here we tracked inter-brain synchronization in the Nucleus Accumbens (NAc), a hub of motivational salience, while heroin-addicted individuals and healthy control subjects watched the same engaging heroin-related movie. Strikingly, the left NAc was synchronized during drug scenes in the addicted individuals and non-drug scenes in controls, predicting scene- and movie-induced heroin craving in the former. Our results open a window into the neurobiology underlying shared drug-biased processing of naturalistic stimuli and cue-induced craving in opiate addiction as they unfold in the real world.

q-bio.NC

Hunting micrometer-sized graphene flakes on gold substrate

Gold is widely used as the substrate material in many graphene devices, due to its superior optoelectronic properties and chemical stability. However, there has been little experimental investigation on the optical contrast of graphene films on Au substrates. Here we report accurate measurement of the optical contrast spectra of few-layer graphene flakes on bulk Au. We used a high-resolution optical microscopy with a 100x magnification objective, accurately determining the thickness of flakes as small as one micrometer in lateral size, which are highly desired in many applications. The results are in excellent agreement with theoretical calculations and confirmed by Raman and AFM measurements. Furthermore, we demonstrate that the optical contrast spectroscopy is sensitive enough to detect the adsorption of a sub-monolayer airborne hydrocarbon molecules, which can reveal whether graphene is con-taminated and opens the opportunity to develop miniaturized and ultrasensitive molecular sensors.

cond-mat.mtrl-sci