SearcharxivSearch

arXiv subjects

Benjamin Elder

Publications and source records attributed to Benjamin Elder.

At least 19 recordsLinked to original sources

Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course

Large language model (LLM)-powered agents can be accurate on average yet unreliable in production, a discrepancy that has been observed but remains largely unaddressed. When given the same task five times, a ReAct agent on the AppWorld benchmark using GPT-4.1 succeeds in all five runs only 53% of the time, even though its per-run pass rate averages 77%. We call this 24-point shortfall the consistency gap, and we argue that addressing it is a precondition for trustworthy AI agent deployment. We present a self-evolving agent framework that reduces this gap by identifying unstable, low-consistency steps in agent trajectories and converting them into episodic memory the agent can draw on in future runs. At its core is a Consistency Analyzer that pinpoints where and why a trajectory is likely to flip across executions, and a Guideline Generator that converts the diagnosis into targeted guidelines, committed to memory and injected into future agent executions on similar tasks. On AppWorld with ReAct/GPT-4.1, our framework raises the fraction of tasks that succeed in all five runs by +16 points on same-task evaluation and +13 points on similar-task generalization.

cs.AI

VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies

Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation. We introduce VAKRA (e\textbf{V}aluating \textbf{A}PI and \textbf{K}nowledge \textbf{R}etrieval \textbf{A}gents), a benchmark of over $8{,}000$ executable APIs across $62$ domains with tasks spanning three settings of increasing difficulty: diverse API interaction styles, multi-hop reasoning over structured APIs, and multi-source reasoning with natural-language tool-use policy constraints. Correctness is verified by re-executing predicted tool calls against live APIs, accommodating multiple valid paths. Using a fixed ReAct harness to isolate model capabilities from agent architecture, we evaluate frontier and open-weight models and find that even the best model achieves only 70.4\% on single-hop endpoint-style tasks and drops to 50--51\% on compositional APIs; performance degrades by over 50\% as reasoning depth increases, and policy-constrained questions expose severe failures (as low as 2.4\% on unanswerable queries). Trace analysis shows failures concentrate at language-mediated reasoning - entity disambiguation, cross-source grounding, rather than tool invocation mechanics. Code is available https://github.com/IBM/VAKRA. Dataset is available https://huggingface.co/datasets/ibm-research/VAKRA

cs.AI

Asymptotic limits of constrained instantons

We revisit the topic of false vacuum decay in field theory. We focus on a toy model of a real massive scalar field with an unstable quartic potential. This model has a false vacuum, and decay out of the false vacuum can be described via the method of constrained instantons, which work by introducing a constraint on the path integral. We identify and develop three different asymptotic limits which enable analytic construction of approximate {constrained} solutions. The first, in which the constrained solution is small compared to the inverse mass of the scalar field, is an application of the perturbative methods of Affleck, although we re-derive the main results and identify several terms which were previously neglected. Second, for very large constrained solutions we adapt the thin-wall approximation of Coleman. However, we find that the large instanton limit does not always exist. In this case we identify another useful limit, in which the Lagrange multiplier used to implement the constraint is large. In this limit, the solution's scaling with the parameters may be found via dimensional analysis and an exact solution is obtained with a single numerical computation.

hep-th

Searching for screened scalar forces with long-baseline atom interferometers

Screened scalars are ubiquitous in many dark-sector models. They give rise to non-trivial fifth forces whilst evading experimental constraints through density-dependent screening mechanisms. We propose equipping a 10\,m-scale long-baseline atom interferometer with an annular planar source mass inside the vacuum chamber to search for such screened fifth forces. Two key challenges arise: distinguishing the static fifth force from backgrounds, and isolating it from the plate's Newtonian gravity. We introduce the `$Q$-flip protocol', which alternates between interferometry sequences to induce controllable time-dependence, aiding signal extraction and de-trending of transient noise. We further develop an \emph{in situ} calibration procedure to characterise the plate's Newtonian gravity and reach shot-noise-limited sensitivity. We show that our proposal could test theoretically motivated parameter space, advancing existing bounds in chameleon and symmetron screened scalar models by $1$ to $1.5$ orders of magnitude. Our proposal is directly applicable to forthcoming experiments, such as AION-10 or VLBAI, and is readily extensible to broader theoretical models and longer baselines.

hep-ph

Constrained instantons in scalar field theories

Instantons, localised saddle points of the action, play an important role in describing non-perturbative aspects of quantum field theories, for example vacuum decay or violation of conservation laws associated with anomalous symmetries. However, there are theories in which no saddle point exists. In this paper, we revisit the idea of constrained instantons, proposed initially by Affleck in 1981, and develop it into a complete method for computing the vacuum decay rate in such cases. We apply this approach to the massive scalar field theory with a negative quartic self-interaction using two different constraints. We solve the field equations numerically and find a two-branch structure, with two distinct solutions for each value of the constraint. By counting the negative modes, we identify one branch of solutions as the constrained instantons and the other as the minima of the action subject to the constraint. We discuss their significance for the computation of the vacuum decay rate.

hep-th

Live API-Bench: 2500+ Live APIs for Testing Multi-Step Tool Calling

Large language models (LLMs) increasingly rely on external tools and APIs to execute complex tasks specified in natural language. Evaluating such tool calling capabilities in realistic enterprise settings is challenging: APIs are often proprietary, heterogeneous, and difficult to share, limiting reproducible benchmarks. To address this, we introduce Live API Bench, a comprehensive benchmark constructed by transforming NL2SQL datasets into interactive API environments. Our pipeline converts SQL queries from BIRD SQL into executable API sequences across three formulations SLOT, SEL, and REST covering minimal general purpose operations, domain specific multi step tasks, and function oriented RESTful interactions, respectively. The benchmark spans 11 databases with over 2,500 invocable tools, paired with human authored queries, ground truth API sequences, and verified final answers. Live API Bench enables systematic evaluation of core challenges in tool use, including error handling, sequential reasoning, parameter generation, response parsing, and robustness across diverse domains. We evaluate 10 LLMs and 4 ReACT agents, observing low task completion rates (7 to 47pct), which improve modestly to 50pct under interactive agent settings, highlighting substantial scope for improving LLM tool calling performance. We release all code and data associated with this paper.

cs.SE

Prospects for detecting new dark physics with the next generation of atomic clocks

Wide classes of new fundamental physics theories cause apparent variations in particle mass ratios in space and time. In theories that violate the weak equivalence principle (EP), those variations are not uniform across all particles and may be detected with atomic and molecular clock frequency comparisons. In this work we explore the potential to detect those variations with near-future clock comparisons. We begin by searching published clock data for variations in the electron-proton mass ratio. We then undertake a statistical analysis to model the noise in a variety of clock pairs that can be built in the near future according to the current state of the art, determining their sensitivity to various fundamental physics signals. Those signals are then connected to constraints on fundamental physics theories that lead directly or indirectly to an effective EP-violating, including those motivated by dark matter, dark energy, the vacuum energy problem, unification or other open questions of fundamental physics. This work results in projections for tight new bounds on fundamental physics that could be achieved with atomic and molecular clocks within the next few years. Our code for this work is packaged into a forecast tool that translates clock characteristics into bounds on fundamental physics.

hep-ph

On the time-dependent density of quadratically coupled dark matter around ordinary matter objects

Wave-like dark matter may feature quadratic couplings to ordinary matter. This carries profound consequences for the phenomenologies of such models. It changes the dark matter density around dense objects made from ordinary matter such as planets and stars, thereby changing the sensitivity of direct detection experiments on Earth as well as implying forces on other ordinary matter objects in the vicinity. In this note we study the time dependence of the dark matter field around spherical objects of ordinary matter. This work indicates the time-scale on which accelerating objects settle into a stationary state and delineates the applicability of stationary solutions for experimental dark matter tests. We also use this to understand (and effectively eliminate) the infinities in energies, forces, and pressures that appear when naively comparing the total energy around objects with different size but the same total number of ordinary matter particles.

hep-ph

Assessment of Prediction Intervals Using Uncertainty Characteristics Curves

Accurate quantification of model uncertainty has long been recognized as a fundamental requirement for trusted AI. In regression tasks, uncertainty is typically quantified using prediction intervals calibrated to an ad-hoc operating point, making evaluation and comparison across different studies relatively difficult. Our work leverages: (1) the concept of operating characteristics curves and (2) the notion of a gain over a null reference, to derive a novel operating point agnostic assessment methodology for prediction intervals. The paper defines the Uncertainty Characteristics Curve and demonstrates its utility in selected scenarios. We argue that the proposed method addresses the current need for comprehensive assessment of prediction intervals and thus represents a valuable addition to the uncertainty quantification toolbox.

cs.LG

Detecting Dark Domain Walls

Light scalar fields, with double well potentials and direct matter couplings, undergo density driven phase transitions, leading to the formation of domain walls. Such theories could explain dark energy, dark matter or source the nanoHz gravitational-wave background. We describe an experiment that could be used to detect such domain walls in a laboratory experiment, solving for the scalar field profile, and showing how the domain wall affects the motion of a test particle. We find that, in currently unconstrained regions of parameter space, the domain walls leave detectable signatures.

gr-qc

Constraining the chameleon-photon coupling with atomic spectroscopy

We compute bounds from atomic spectroscopy on chameleon fields that couple to the photon. Chameleons are a wide class of scalar field models that generically lead to screened fifth forces and a host of novel phenomenologies, particularly when the photon coupling is included. We account for perturbations to the atomic energy levels from both the scalar field "fifth force" and the scalar field's correction to the electric field. We also account for the electromagnetic interaction's contribution to the scalar charge of the proton, which enables a considerably wider class of models to be tested than without this effect. We find bounds that cover different areas of chameleon parameter space. Some regions are redundant with existing experiments, particularly $g - 2$, confirming that those models are ruled out. Other regions were previously unconstrained, and a range of models spanning approximately four orders of magnitude in chameleon coupling parameters are excluded for the first time.

hep-ph

Casimir Tests of Scalar-Tensor Theories

We compute bounds and forecasts on screened modified gravity theories, specialising to the chameleon model in Casimir force experiments. In particular, we investigate the classical interaction between a plate and sphere subject to a screened interaction of the chameleon type. We compare numerical simulations of the field profile and the classical pressure exerted on the sphere to analytical approximations for these non-linear field theories. In particular, we focus on the proximity force approximation (PFA) and show that, within the range of sphere sizes $R$ and plate-sphere distance $D$ simulated numerically, the PFA does not reproduce the numerical results. This differs from the case of linear field theories such as Newtonian gravity and a Yukawa model where the PFA coincides with the exact results. We show that for chameleon theories, the screening factor approximation (SFA) whereby the sphere is modelled as a screened sphere embedded in the external field due to the plates, fares better and can be used in the regime $D\gtrsim R$ to extract constraints and forecasts from existing and forthcoming data. In particular, we forecast that future Casimir experiments would corroborate the closing of the parameter space for the simplest of chameleon models at the dark energy scale.

gr-qc

Mapping the Weak-Field Limit of Scalar-Gauss-Bonnet Gravity

We derive the weak field limit of scalar-Gauss-Bonnet theory and place novel bounds on the parameter space using terrestrial and space-based experiments. In order to analyze the theory in the context of a wide range of experiments, we compute the deviations from Einstein gravity around source masses with planar, cylindrical, and spherical symmetry. We find a correction to the Newtonian potential around spherical and cylindrical sources that can be larger than PPN corrections sufficiently close to the source. We use this to improve on laboratory constraints on the scalar-Gauss-Bonnet coupling parameter $\Lambda$ by two orders of magnitude. Present laboratory and Solar System bounds reported here are superseded by tests deriving from black holes.

gr-qc

Hide and Seek: Screened Scalar Fields in Hydrogen and Muonium

We compute bounds on screened scalar field theories from hydrogen-like systems. New light scalar fields generically have a direct coupling to matter. Such a coupling is strongly constrained by myriad experimental measurements. However, certain theories possess a {\it screening mechanism} that allows the effects of this coupling to weaken dynamically, and to evade many such bounds. We compute the perturbations to the energy levels of hydrogen-like systems due to screened scalar fields. We then use this result in two ways. First, we compute bounds from hydrogen spectroscopy, finding significantly weaker bounds than have been reported before as screening effects were overlooked. Second, we show that muonium is an intrinsically much more sensitive probe of screened scalar fields. For chameleon models, muonium experiments probe a large part of the parameter space that is as yet unexplored by low energy physics and has so far only been tested by high-energy particle physics experiments.

hep-ph

Testing Screened Modified Gravity

Long range scalar fields with a coupling to matter appear to violate known bounds on gravitation in the solar system and the laboratory. This is evaded thanks to screening mechanisms. In this short review, we shall present the various screening mechanisms from an effective field theory point of view. We then investigate how they can and will be tested in the laboratory and on astrophysical and cosmological scales.

gr-qc

Muon $g-2$ and Screened Modified Gravity

We show how light scalar fields can account for the discrepancy between the theoretical and observed values of the anomalous magnetic moment of the (anti)muon. When coupled to both matter and photons, light scalar fields induce a change of the anomalous magnetic moment of charged particles. This arises from two concurrent effects. Classically, light scalars induce a change of the cyclotron frequency, complementing the electromagnetic effects coming from the magnetic and electric fields used experimentally. Light scalars also contribute to the anomalous magnetic moment quantum mechanically at the one-loop level. For unscreened scalar fields coupling with a Yukawa interaction to matter, these contributions are negligible after applying the Cassini bound on deviations from Newtonian gravity. On the other hand, screened scalars such as chameleons or symmetrons can couple strongly to matter in the laboratory and decouple in the Solar System. This allows us to probe branches of their parameter spaces where the recently measured anomalous magnetic moment of the (anti)muon can be accounted for in the chameleon and symmetron cases. This might be a hint that modified gravity is at play in the laboratory.

hep-ph

Uncertainty Characteristics Curves: A Systematic Assessment of Prediction Intervals

Accurate quantification of model uncertainty has long been recognized as a fundamental requirement for trusted AI. In regression tasks, uncertainty is typically quantified using prediction intervals calibrated to a specific operating point, making evaluation and comparison across different studies difficult. Our work leverages: (1) the concept of operating characteristics curves and (2) the notion of a gain over a simple reference, to derive a novel operating point agnostic assessment methodology for prediction intervals. The paper describes the corresponding algorithm, provides a theoretical analysis, and demonstrates its utility in multiple scenarios. We argue that the proposed method addresses the current need for comprehensive assessment of prediction intervals and thus represents a valuable addition to the uncertainty quantification toolbox.

cs.LG

Modified Gravity and Cosmology: An Update by the CANTATA Network

General Relativity and the $\Lambda$CDM framework are currently the standard lore and constitute the concordance paradigm. Nevertheless, long-standing open theoretical issues, as well as possible new observational ones arising from the explosive development of cosmology the last two decades, offer the motivation and lead a large amount of research to be devoted in constructing various extensions and modifications. All extended theories and scenarios are first examined under the light of theoretical consistency, and then are applied to various geometrical backgrounds, such as the cosmological and the spherical symmetric ones. Their predictions at both the background and perturbation levels, and concerning cosmology at early, intermediate and late times, are then confronted with the huge amount of observational data that astrophysics and cosmology are able to offer recently. Theories, scenarios and models that successfully and efficiently pass the above steps are classified as viable and are candidates for the description of Nature. This work is a Review of the recent developments in the fields of gravity and cosmology, presenting the state of the art, high-lighting the open problems, and outlining the directions of future research. Its realization was performed in the framework of the COST European Action ``Cosmology and Astrophysics Network for Theoretical Advances and Training Actions''.

gr-qc