SearcharxivSearch

arXiv subjects

Qiancheng Xu

Publications and source records attributed to Qiancheng Xu.

13 recordsLinked to original sources

TInR: Exploring Tool-Internalized Reasoning in Large Language Models

Tool-Integrated Reasoning (TIR) has emerged as a promising direction by extending Large Language Models' (LLMs) capabilities with external tools during reasoning. Existing TIR methods typically rely on external tool documentation during reasoning. However, this leads to tool mastery difficulty, tool size constraints, and inference inefficiency. To mitigate these issues, we explore Tool-Internalized Reasoning (TInR), aiming at facilitating reasoning with tool knowledge internalized into LLMs. Achieving this goal presents notable requirements, including tool internalization and tool-reasoning coordination. To address them, we propose TInR-U, a tool-internalized reasoning framework for unified reasoning and tool usage. TInR-U is trained through a three-phase pipeline: 1) tool internalization with a bidirectional knowledge alignment strategy; 2) supervised fine-tuning warm-up using high-quality reasoning annotations, and 3) reinforcement learning with TInR-specific rewards. We comprehensively evaluate our method across in-domain and out-of-domain settings. Experiment results show that TInR-U achieves superior performance in both settings, highlighting its effectiveness and efficiency.

cs.CL

KNN-SSD: Enabling Dynamic Self-Speculative Decoding via Nearest Neighbor Layer Set Optimization

Speculative Decoding (SD) has emerged as a widely used paradigm to accelerate the inference of large language models (LLMs) without compromising generation quality. It works by efficiently drafting multiple tokens using a compact model and then verifying them in parallel using the target LLM. Notably, Self-Speculative Decoding proposes skipping certain layers to construct the draft model, which eliminates the need for additional parameters or training. Despite its strengths, we observe in this work that drafting with layer skipping exhibits significant sensitivity to domain shifts, leading to a substantial drop in acceleration performance. To enhance the domain generalizability of this paradigm, we introduce KNN-SSD, an algorithm that leverages K-Nearest Neighbor (KNN) search to match different skipped layers with various domain inputs. We evaluated our algorithm in various models and multiple tasks, observing that its application leads to 1.3x-1.6x speedup in LLM inference.

cs.CL

Agent-as-a-Judge

LLM-as-a-Judge has revolutionized AI evaluation by leveraging large language models for scalable assessments. However, as evaluands become increasingly complex, specialized, and multi-step, the reliability of LLM-as-a-Judge has become constrained by inherent biases, shallow single-pass reasoning, and the inability to verify assessments against real-world observations. This has catalyzed the transition to Agent-as-a-Judge, where agentic judges employ planning, tool-augmented verification, multi-agent collaboration, and persistent memory to enable more robust, verifiable, and nuanced evaluations. Despite the rapid proliferation of agentic evaluation systems, the field lacks a unified framework to navigate this shifting landscape. To bridge this gap, we present the first comprehensive survey tracing this evolution. Specifically, we identify key dimensions that characterize this paradigm shift and establish a developmental taxonomy. We organize core methodologies and survey applications across general and professional domains. Furthermore, we analyze frontier challenges and identify promising research directions, ultimately providing a clear roadmap for the next generation of agentic evaluation.

cs.CL

Towards Dynamic Theory of Mind: Evaluating LLM Adaptation to Temporal Evolution of Human States

As Large Language Models (LLMs) increasingly participate in human-AI interactions, evaluating their Theory of Mind (ToM) capabilities - particularly their ability to track dynamic mental states - becomes crucial. While existing benchmarks assess basic ToM abilities, they predominantly focus on static snapshots of mental states, overlooking the temporal evolution that characterizes real-world social interactions. We present \textsc{DynToM}, a novel benchmark specifically designed to evaluate LLMs' ability to understand and track the temporal progression of mental states across interconnected scenarios. Through a systematic four-step framework, we generate 1,100 social contexts encompassing 5,500 scenarios and 78,100 questions, each validated for realism and quality. Our comprehensive evaluation of ten state-of-the-art LLMs reveals that their average performance underperforms humans by 44.7\%, with performance degrading significantly when tracking and reasoning about the shift of mental states. This performance gap highlights fundamental limitations in current LLMs' ability to model the dynamic nature of human mental states.

cs.CL

PEToolLLM: Towards Personalized Tool Learning in Large Language Models

Tool learning has emerged as a promising direction by extending Large Language Models' (LLMs) capabilities with external tools. Existing tool learning studies primarily focus on the general-purpose tool-use capability, which addresses explicit user requirements in instructions. However, they overlook the importance of personalized tool-use capability, leading to an inability to handle implicit user preferences. To address the limitation, we first formulate the task of personalized tool learning, which integrates user's interaction history towards personalized tool usage. To fill the gap of missing benchmarks, we construct PEToolBench, featuring diverse user preferences reflected in interaction history under three distinct personalized settings, and encompassing a wide range of tool-use scenarios. Moreover, we propose a framework PEToolLLaMA to adapt LLMs to the personalized tool learning task, which is trained through supervised fine-tuning and direct preference optimization. Extensive experiments on PEToolBench demonstrate the superiority of PEToolLLaMA over existing LLMs.

cs.CL

Hybrid Machine Learning and Physics-based Modelling of Pedestrian Pushing Behaviours at Bottlenecks

In high-density crowds, close proximity between pedestrians makes the steady state highly vulnerable to disruption by pushing behaviours, potentially leading to serious accidents. However, the scarcity of experimental data has hindered systematic studies of its mechanisms and accurate modelling. Using behavioural data from bottleneck experiments, we investigate pedestrian heterogeneity in pushing tendencies, showing that pedestrians tend to push under high-motivation and in wider corridors. We introduce a spatial discretization method to encode neighbour states into feature vectors, serving together with pedestrian pushing tendencies as inputs to a random forest model for predicting pushing behaviours. Through comparing speed-headway relationships, we reveal that pushing behaviours correspond to an aggressive space-utilization movement strategy. Consequently, we propose a hybrid machine learning and physics-based model integrating pushing tendencies heterogeneity, pushing behaviours prediction, and dynamic movement strategies adjustment. Validations show that the hybrid model effectively reproduces experimental crowd dynamics and fits to incorporate additional behaviours.

physics.soc-ph

Enhancing Tool Retrieval with Iterative Feedback from Large Language Models

Tool learning aims to enhance and expand large language models' (LLMs) capabilities with external tools, which has gained significant attention recently. Current methods have shown that LLMs can effectively handle a certain amount of tools through in-context learning or fine-tuning. However, in real-world scenarios, the number of tools is typically extensive and irregularly updated, emphasizing the necessity for a dedicated tool retrieval component. Tool retrieval is nontrivial due to the following challenges: 1) complex user instructions and tool descriptions; 2) misalignment between tool retrieval and tool usage models. To address the above issues, we propose to enhance tool retrieval with iterative feedback from the large language model. Specifically, we prompt the tool usage model, i.e., the LLM, to provide feedback for the tool retriever model in multi-round, which could progressively improve the tool retriever's understanding of instructions and tools and reduce the gap between the two standalone components. We build a unified and comprehensive benchmark to evaluate tool retrieval models. The extensive experiments indicate that our proposed approach achieves advanced performance in both in-domain evaluation and out-of-domain evaluation.

cs.CL

On-chip mechanical exceptional points based on an optomechanical zipper cavity

Exceptional points (EPs) represent a distinct type of spectral singularity in non-Hermitian systems, and intriguing physics concepts have been studied with optical EPs recently. As a system beyond photonics, the mechanical oscillators coupling with many physical systems are expected to be further exploited EPs for mechanical sensing, topology energy transfer, nonreciprocal dynamics etc. In this study, we demonstrated on-chip mechanical EPs with a silicon optomechanical zipper cavity, wherein two near-degenerate mechanical breathing modes are coupled via a single co-localized optical mode. By tailoring the dissipative and coherent couplings between two mechanical oscillators, the spectral splitting with 1/2 order response, a distinctive feature of EP, was observed successfully. Our work provides an integrated platform for investigating the physics related to mechanical EPs on silicon chips and suggests their possible applications for ultrasensitive measurements.

physics.optics

Continual Learning for Task-oriented Dialogue System with Iterative Network Pruning, Expanding and Masking

This ability to learn consecutive tasks without forgetting how to perform previously trained problems is essential for developing an online dialogue system. This paper proposes an effective continual learning for the task-oriented dialogue system with iterative network pruning, expanding and masking (TPEM), which preserves performance on previously encountered tasks while accelerating learning progress on subsequent tasks. Specifically, TPEM (i) leverages network pruning to keep the knowledge for old tasks, (ii) adopts network expanding to create free weights for new tasks, and (iii) introduces task-specific network masking to alleviate the negative impact of fixed weights of old tasks on new tasks. We conduct extensive experiments on seven different tasks from three benchmark datasets and show empirically that TPEM leads to significantly improved results over the strong competitors. For reproducibility, we submit the code and data at: https://github.com/siat-nlp/TPEM

cs.CL

Anticipation in a velocity-based model for pedestrian dynamics

Lane formation in bidirectional pedestrian streams is based on a stimulus-response mechanism and strategies of navigation in a fast-changing environment. Although microscopic models that only guarantee volume exclusion can qualitatively reproduce this phenomenon, they are not sufficient for a quantitative description. To quantitatively describe this phenomenon, a minimal anticipatory collision-free velocity model is introduced. Compared to the original velocity model, the new model reduces the occurrence of gridlocks and reproduces the movement of pedestrians more realistically. For a quantitative description of the phenomenon, the definition of an order parameter is used to describe the formation of lanes at transient states and to show that the proposed model compares relatively well with experimental data. Furthermore, the model is validated by the experimental fundamental diagrams of bidirectional flows.

physics.soc-ph

Prolonged Clogs in Bottleneck Simulations for Pedestrian Dynamics

This article studies clogging phenomena using a velocity-based model for pedestrian dynamics. First, a method to identify prolonged clogs in simulations was introduced. Then bottleneck simulations were implemented with different initial and boundary conditions. The number of prolonged clogs was analyzed to investigate the decisive factors causing this phenomenon. Moreover, the time lapse between two consecutive agents passing the exit, and the trajectories of agents were analyzed. The influence of three types of factors was studied: parameters of the spatial boundaries, algorithmic factors related to the implementation of the model, and the movement model. Parameters of the spatial boundaries include the width and position of the bottleneck exit. Algorithmic factors are the update methods and the size of the time step. Model parameters cover several parameters describing the level of motivation, the strength and range of impact among agents, and the shape of agents. The results show that the occurrence of prolonged clogs is closely linked to parameters of the spatial boundaries and the movement model but has virtually no correlation with algorithmic factors.

physics.soc-ph

Generalized collision-free velocity model for pedestrian dynamics

The collision-free velocity model is a microscopic pedestrian model, which despite its simplicity, reproduces fairly well several self-organization phenomena in pedestrian dynamics. The model consists of two components: a direction sub-model that combines individual desired moving direction and neighbor's influence to imitate the process of navigating in a two-dimensional space, and an intrinsically collision-free speed sub-model which controls the speed of the agents with respect to the distance to their neighbors. In this paper we generalize the collision-free velocity model by introducing the influence of walls and extending the distance calculations to velocity-based ellipses. Besides, we introduce enhancements to the direction sub-module that smooth the direction changes of pedestrians in the simulation; a shortcoming that was not visible in the original model due to the symmetry of the circular shapes. Moreover, the introduced improvements mitigate backward movements, leading to a more realistic distribution of pedestrians especially in bottleneck scenarios. We study by simulation the effects of the pedestrian's shape by comparing the fundamental diagram in narrow and wide corridors. Furthermore, we validate our generalized approach by investigating the flow through bottlenecks with varying exit's widths.

physics.soc-ph

Phonon-laser sensing in a hetero optomechanical crystal cavity

Micro- and nanomechanical resonators have emerged as promising platforms for sensing a broad range of physical properties such as mass, force, torque, magnetic field, and acceleration. The sensing performance relies critically on the motional mass, the mechanical frequency, and the linewidth of the mechanical resonator. Here, we demonstrate a hetero optomechanical crystal (OMC) cavity based on a silicon nanobeam structure. The cavity supports phonon lasing in a fundamental mechanical mode with a frequency of 5.91 GHz, an effective mass of 116 fg, and a mechanical linewidth narrowing from 3.3 MHz to 5.2 kHz, while the optomechanical coupling rate of is as high as 1.9 MHz. With this phonon laser, the on-chip sensing with a resolution of $δ$$λ$/$λ$ = 1.0*10-8 can be attained, which is at least two orders of magnitude larger than that obtained with conventional silicon-based sensors. The use of a silicon-based hetero OMC cavity that harnesses phonon lasing could pave the way towards exciting, high-precision sensors that lend themselves to silicon monolithic integration and offer unprecedented sensitivity for broad physical sensing applications.

physics.optics