SearcharxivSearch

arXiv subjects

Sebastian Schreiber

Publications and source records attributed to Sebastian Schreiber.

8 recordsLinked to original sources

ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMs

Large language models deployed as agents over large tool catalogs face a critical tool-retrieval bottleneck. As embedding-based retrieval approaches rely on compact encoders that may under-capture specialized tool semantics, parametric tool retrieval addresses this by encoding each tool as a virtual token appended to the LLM vocabulary, fine-tuned in two stages (memorization then retrieval SFT) to use the LLM as a retriever, achieving strong performance on standard ToolBench retrieval benchmarks. Yet these benchmarks use verbose, fully-specified queries, and their evaluation applies constrained decoding that restricts outputs to valid token paths, neither reveals whether the model actually understands its tools. We introduce \textbf{ToolSense}, an open-source LLM-powered diagnostic framework that takes any tool catalog as input and automatically generates three benchmarks: a Realistic Retrieval Benchmark (RRB) with queries at three ambiguity tiers, an MCQ probing benchmark, and a QA probing benchmark. Applying ToolSense to ToolBench (~47k tools) and evaluating five parametric model training configurations reveals a knowledge-retrieval dissociation: on RRB queries, several configurations collapse by ~50-64 percentage points compared to fully-specified ToolBench benchmarks, falling below the embedding-model baseline. Additionally, despite strong retrieval performance, some models score near-random on factual probes, suggesting a knowledge-retrieval dissociation. We open-source the ToolSense framework and the ToolBench diagnostic benchmarks at https://github.com/SAP/toolsense.

cs.AI

CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval

Tool retrieval over large API catalogs is a core bottleneck for LLM agents: user queries arrive in colloquial, often underspecified language, while the catalog uses technical API vocabulary that no fixed encoder can bridge on its own. The two dominant training approaches, contrastive encoder fine-tuning and HyDE-style query expansion with a frozen LLM, address this problem from opposite ends and fail in complementary directions: the fine-tuned encoder excels when the query's surface form already matches the catalog but collapses when it does not, while zero-shot HyDE is more robust to underspecified queries yet generates catalog-unaware hypothetical descriptions that degrade retrieval when queries are well-formed. We introduce CoHyDE, an iterative procedure that trains the dense encoder and the LLM rewriter as a single co-evolving system: the encoder is retrained with InfoNCE on catalog-style hypothetical descriptions produced by the rewriter, and the rewriter is preference-aligned via DPO against the encoder's retrieval scores, with both sides warm-started on the tool catalog before the loop begins. On a ~10k tool subset of the ToolBench catalog, three rounds of CoHyDE improve over the strongest single-component baseline by +2.5 pp NDCG@5 on standard queries and +6.3 pp on held-out vague queries, with gains as large as +8 pp on the hardest vague tier. Ablations confirm that co-training is the key ingredient: using either component in isolation fails to match CoHyDE on both well-formed and vague queries, with losses of up to -8 pp on vague queries.

cs.AI

MirrorBench: A Benchmark to Evaluate Conversational User-Proxy Agents for Human-Likeness

Large language models (LLMs) are increasingly used as human simulators, both for evaluating conversational systems and for generating fine-tuning data. However, naive "act-as-a-user" prompting often yields verbose, unrealistic utterances, motivating principled evaluation of *user proxy agents*. We present **MirrorBench**, a reproducible and extensible benchmarking framework that evaluates user proxies solely on their ability to produce human-like user utterances across diverse conversational regimes, explicitly decoupled from downstream task success. **MirrorBench** combines three lexical-diversity metrics (**MATTR**, **Yule's~$K$**, and **HD-D**) with three LLM-judge-based metrics (**GTEval**, **Pairwise Indistinguishability**, and **Rubric-and-Reason**), and contextualizes judge scores using Human-Human and Proxy-Proxy calibration controls. Across four public datasets, **MirrorBench** yields variance-aware comparisons and reveals systematic gaps between user proxies and real human users. The framework is open sourced at https://github.com/SAP/mirrorbench and includes a command-line interface for running and managing user-proxy benchmarking experiments.

cs.AI

Planning Agents on an Ego-Trip: Leveraging Hybrid Ego-Graph Ensembles for Improved Tool Retrieval in Enterprise Task Planning

Effective tool pre-selection via retrieval is essential for AI agents to select from a vast array of tools when identifying and planning actions in the context of complex user queries. Despite its central role in planning, this aspect remains underexplored in the literature. Traditional approaches rely primarily on similarities between user queries and tool descriptions, which significantly limits retrieval accuracy, specifically when handling multi-step user requests. To address these limitations, we propose a Knowledge Graph (KG)-based tool retrieval framework that captures the semantic relationships between tools and their functional dependencies. Our retrieval algorithm leverages ensembles of 1-hop ego tool graphs to model direct and indirect connections between tools, enabling more comprehensive and contextual tool selection for multi-step tasks. We evaluate our approach on a synthetically generated internal dataset across six defined user classes, extending previous work on coherent dialogue synthesis and tool retrieval benchmarks. Results demonstrate that our tool graph-based method achieves 91.85% tool coverage on the micro-average CompleteRecall metric, compared to 89.26% for re-ranked semantic-lexical hybrid retrieval, the strongest non-KG baseline in our experiments. These findings support our hypothesis that the structural information modeled in the graph provides complementary signals to pure similarity matching, particularly for queries requiring sequential tool composition.

cs.AI

Disambiguation-Centric Finetuning Makes Enterprise Tool-Calling LLMs More Realistic and Less Risky

Large language models (LLMs) are increasingly tasked with invoking enterprise APIs, yet they routinely falter when near-duplicate tools vie for the same user intent or when required arguments are left underspecified. We introduce DiaFORGE (Dialogue Framework for Organic Response Generation & Evaluation), a disambiguation-centric, three-stage pipeline that (i) synthesizes persona-driven, multi-turn dialogues in which the assistant must distinguish among highly similar tools, (ii) performs supervised fine-tuning of open-source models with reasoning traces across 3B - 70B parameters, and (iii) evaluates real-world readiness via a dynamic suite that redeploys each model in a live agentic loop and reports end-to-end goal completion alongside conventional static metrics. On our dynamic benchmark DiaBENCH, models trained with DiaFORGE raise tool-invocation success by 27 pp over GPT-4o and by 49 pp over Claude-3.5-Sonnet, both under optimized prompting. To spur further research, we release an open corpus of 5000 production-grade enterprise API specifications paired with rigorously validated, disambiguation-focused dialogues, offering a practical blueprint for building reliable, enterprise-ready tool-calling agents.

cs.AI

Boltzmann relaxation dynamics of strongly interacting spinless fermions on a lattice

Motivated by the recent interest in non-equilibrium phenomena in quantum many-body systems, we study strongly interacting fermions on a lattice by deriving and numerically solving quantum Boltzmann equations that describe their relaxation to thermodynamic equilibrium.The derivation is carried out by inspecting the hierarchy of correlations within the framework of the 1/Z-expansion. Applying the Markov approximation, we obtain the dynamic equations for the distribution functions. Interestingly, we find that in the strong-coupling limit, collisions between particles and holes dominate over particle-particle and hole-hole collisions -- in stark contrast to weakly interacting systems. As a consequence, our numerical simulations show that the relaxation time scales strongly depend on the type of excitations (particles or holes or both) that are initially present.

quant-ph

Convergence of generalized urn models to non-equilibrium attractors

Generalized Polya urn models have been used to model the establishment dynamics of a small founding population consisting of k different genotypes or strategies. As population sizes get large, these population processes are well-approximated by a mean limit ordinary differential equation whose state space is the k simplex. We prove that if this mean limit ODE has an attractor at which the temporal averages of the population growth rate is positive, then there is a positive probability of the population not going extinct (i.e. growing without bound) and its distribution converging to the attractor. Conversely, when the temporal averages of the population growth rate is negative along this attractor, the population distribution does not converge to the attractor. For the stochastic analog of the replicator equations which can exhibit non-equilibrium dynamics, we show that verifying the conditions for convergence and non-convergence reduces to a simple algebraic problem. We also apply these results to selection-mutation dynamics to illustrate convergence to periodic solutions of these population genetics models with positive probability.

math.PR

Pushed beyond the brink: Allee effects, environmental stochasticity, and extinction

A demographic Allee effect occurs when individual fitness, at low densities, increases with population density. Coupled with environmental fluctuations in demographic rates, Allee effects can have subtle effects on population persistence and extinction. To understand the interplay between these deterministic and stochastic forces, we analyze discrete-time single species models allowing for general forms of density-dependent feedbacks and stochastic fluctuations in demographic rates. Our analysis provide criteria for stochastic persistence, asymptotic extinction, and conditional persistence. Stochastic persistence requires that the geometric mean of fitness at low densities is greater than one. When this geometric mean is less than one, asymptotic extinction occurs with a high probability whenever the initial population density is low. If in addition the population only experiences positive density-dependent feedbacks, conditional persistence occurs provided the geometric mean of fitness at high population densities is greater than one. However, if the population experiences both positive and negative density-dependent feedbacks, conditional persistence is only possible if fluctuations in demographic rates are sufficiently small. Applying our results to stochastic models of mate-limitation, we illustrate counter-intuitively that the environmental fluctuations can increase the probability of persistence when populations are initially at low densities, and decrease the likelihood of persistence when populations are initially at high densities. Alternatively, for stochastic models accounting for predator saturation and negative density-dependence, environmental stochasticity can result in asymptotic extinction at intermediate predation rates despite conditional persistence occurring at higher predation rates.

math.PR