Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 595 records · Page 33Linked to original sources

Causal Behavioral Evaluation of AI Agents at Scale via Automated Behavioral Science

As AI agents are increasingly deployed in complex and new environments, knowing the conditions that influence their behavior becomes an indispensable step for their reliable and safe deployment. Yet causal behavioral evaluation of AI agents remains manual and labor-intensive. We introduce Abs2Sim and AEROBAT, a system of methods that support causal behavioral evaluation of AI agents via automated behavioral science. Given a user-specified target behavior, the methods automatically execute a full pipeline of behavioral science research---generating hypotheses about the behavior, designing and executing controlled experiments, making behavioral assessments, analyzing the results, and writing reports. For 12 target behaviors, the methods generated and tested 73 hypotheses: designing 1,160 controlled experiments and executing 22,954 simulation rounds in total. Moderate-to-strong statistical evidence emerged for 30 hypotheses, revealing potential modulators of AI behavior. In sum, our results demonstrate that automated behavioral science can extend the reach of behavioral evaluation of AI agents.

cs.AI↗

Relative Ehrhart functions: eventual polynomiality, reciprocity law, and shifted duality

Classical Ehrhart theory measures the discrete capacity of a convex rational (or integral) polytope $P$ by counting the number of lattice points in the $t$-th dilate $tP$ of $P$. In this paper, we extend this paradigm by replacing a lattice point with a geometric object $Q$ of dimension at most $\dim P$. We show that the counting function $\mathrm{ehr}(P,Q;t)$ of such valid translations of $Q$ into $tP$ inherits eventual quasi-polynomiality (or eventual polynomiality) with leading term $\mathrm{vol}(P)t^d$, where $d = \dim P$. This result is naturally derived by induction on the dimension, based on the classical quasi-polynomiality (or polynomiality) of Ehrhart functions. Similarly, replacing $P$ with its relative interior $P^\circ$ defines $\mathrm{ehr}^\circ(P,Q;t)$, which is also shown to be an eventual quasi-polynomial (or eventual polynomial). Furthermore, we prove several formulas for relative Ehrhart functions under the condition that $P = kP_0$ and $Q = lQ_0$ for some integers $k$, $l > 0$, and polytopes $P_0$ and $Q_0$ with $\dim Q_0 > 0$ such that $Q_0$ is inscribed in $P_0$. In particular, we prove that under this condition, $\mathrm{ehr}(P,Q;-t) = (-1)^{d}\mathrm{ehr}^\circ (P,Q;t+ρ)$ $(t \gg 0)$ holds for some integer $ρ> 0$ if and only if there exists an integer $ρ$ such that $kρ= 2l$ via the classical Ehrhart--Macdonald reciprocity law. In addition, we also prove that under the same condition, $\mathrm{ehr}(P,Q;t) = \mathrm{ehr}^\circ (P,Q;t+σ)$ $(t \gg 0)$ holds for some integer $σ> 0$ if and only if $\mathrm{ehr}^\circ (P_0;\mathrm{codeg}\,P_0) = 1$ and there exists an integer $σ$ such that $kσ= \mathrm{codeg}\,P_0$.

math.CO↗

Compact Feed-Forward 3D Gaussians via Saliency-Guided Primitive Merging

3D scene reconstruction, modeling, and rendering are highly relevant for numerous tasks, and 3D Gaussian splatting has become a standard choice in this context. Its feed-forward variants provide fast reconstruction from sparse input views but often produce per-pixel primitives, leading to highly redundant and thus inefficient representations. We present a structure-aware merging pipeline that takes per-pixel primitives from any feed-forward method and consolidates them into a compact, content-adaptive Gaussian set while largely retaining visual quality at just $\frac{1}{20}^\text{th}$ of the Gaussians of a per-pixel method. We group spatially coherent Gaussians of similar appearance into variable-size clusters via adaptive superpixel segmentation guided by a saliency map, which allocates fine segments to textured regions and coarse segments to homogeneous areas. We compress each cluster into a compact latent representation through a learned encoder, then match and consolidate representations across views based on geometric overlap and feature similarity via a learned merger. A level-of-detail decoder then produces the final Gaussians at a controllable resolution, enabling a flexible quality-efficiency trade-off at inference. As a post-processing module, the pipeline is backbone-agnostic, leveraging the strengths of existing feed-forward methods. This leads to better and more robust quality than achieved by previous approaches that target a reduction in primitive count, while providing a highly compact representation, that can be rendered efficiently.

cs.CV↗

Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling

On-policy distillation (OPD) has emerged as a promising post-training technique for enhancing LLM reasoning. Under the reverse KL objective, the idealized optimum of OPD aligns the student distribution with that of the teacher. When the teacher consistently outperforms the student, this naturally suggests that OPD should yield broad improvements over the pre-OPD student. However, do such improvements extend across the entire range of test-time sampling budgets? In this work, we revisit this expectation through the lens of test-time scaling by varying the sampling budget $K$ and evaluating performance with pass@$K$. Across multiple settings, we observe two distinct patterns: OPD can improve pass@$K$ at both small and large sampling budgets, but it can also improve small-budget performance while reducing large-budget pass@$K$. We show one condition that guarantees such a reversal and an idealized reverse KL counterexample where it occurs even when the teacher has higher accuracy on every problem. To choose between two candidate teachers at a target sampling budget, we propose the \textit{Teacher Advantage Score at $K$} (TAS@$K$), which can be computed before OPD training to predict which teacher will lead to a larger improvement in pass@$K$. Across three domains and thirteen benchmarks, the ordering predicted by TAS@$K$ agrees with the observed pass@$K$ improvements of the resulting OPD models in 83.6\% of experiments, providing a useful signal for teacher selection at the target pass@$K$.

cs.LG↗

Error-Aware Reverse Auction Mechanism for Large Language Model Routing

Routing each query to a cost-effective large language model (LLM) is critical for balancing quality and cost, yet most routers rely on a centralized task center to predict model performance, creating an information-risk mismatch and a scalability bottleneck as the model pool grows. We formulate LLM routing as a market-based allocation problem among strategic providers and propose a routing paradigm that shifts ex-ante prediction to LLM providers via a reverse auction, where providers submit self-predicted acceptance probabilities and execution costs. To account for noisy provider predictions and center evaluations, we introduce the \textit{\textbf{E}rror-\textbf{A}ware \textbf{R}everse \textbf{A}uction \textbf{M}echanism} (EA-RAM), which explicitly models this Dual Error. We prove that, under a private-evaluation-belief structure, truthful effective-surplus reporting is incentive compatible in the reduced-form score space and individually rational under sellers' subjective beliefs, establish sufficient conditions for center rationality, and derive an explicit social-welfare loss bound. We further identify robustness effects: opposite-signed errors can cancel, vanishing-tail link functions (e.g., logistic) stabilize clear-cut cases via saturation, and extra noise smooths belief maps and reduces their maximal local sensitivity. Simulations and real-world benchmarks show that EA-RAM is robust to Dual Error and achieves a better cost--performance Pareto frontier than centralized baselines, with additional gains from provider-side local information, validating its practical effectiveness.

cs.GT↗

ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval

While Large Language Model (LLM) agents increasingly rely on long-term memory for persistent interactions, the retrieval mechanisms governing this memory are rarely treated as evolvable components. This static approach limits performance on heterogeneous memory queries, which often demand diverse evidence construction strategies. To address this, we introduce \textbf{ERSkill}, a retrieval-centric framework for evolving, skill-guided memory access. ERSkill compiles interaction histories into a structured memory store and represents retrieval behaviors as executable skills composed of fundamental primitives. At inference time, a trained router dynamically matches each query to a suitable retrieval skill to construct tailored evidence for answer generation. ERSkill co-evolves the skill set and the router during training. It employs an experience trie to efficiently record explored retrieval paths, alongside a double-frontier mechanism that separates oracle-side capability expansion from router-validated deployment. Experiments across multiple agent memory benchmarks demonstrate that ERSkill substantially outperforms strong non-evolving and evolving baselines. Notably, it improves the overall average across F1, BLEU-1, and LLM-judge scores by 31.3\% with Qwen3-Next-80B-A3B-Instruct and by 21.4\% with GPT-5.4-nano.

cs.CL↗

Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence

Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLM agents, existing work has mainly followed two lines. One line uses post-training methods, such as supervised fine-tuning and reinforcement learning. Another line adopts an agentic paradigm in which the model calls external spatial tools, such as depth estimation and 3D reconstruction tools, to gather intermediate spatial evidence. We study a complementary and underexplored route: Can a frozen VLM agent improve its spatial reasoning through \textbf{parameter-update-free self-evolution}, without depending on external expert spatial tools at inference time? We present \textbf{Spatial Memory Agent (SMA)}, an \textbf{experience-grounded runtime framework} that converts verified spatial experience into reusable transferable lessons. In a verifiable spatial environment, SMA queries the frozen VLM, obtains a predicted answer and reward, and uses \textbf{verifier-guided reflection} to distill compact transferable lessons from spatial experience. SMA further assigns each lesson a \textbf{Transfer Reliability Score (TRS)}, which is initialized uniformly and calibrated from later retrieval outcomes as visit evidence of future transfer reliability. During \textbf{read-only deployment}, SMA retrieves lessons by semantic filter and similarity-TRS combined ranking, allowing the retrieved memory to guide frozen model inference. Across five representative spatial benchmarks and four base VLMs, SMA achieves the highest macro average in every base-model block and the best accuracy among the evaluated methods in most of the 20 evaluations, establishing a practical parameter-update-free path for spatial self-evolution across the evaluated frozen model scales and environments.

cs.AI↗

AQuA: Recursively Self-Improving Quantitative Trading Research Agents

We study recursive self-improvement at the level of quantitative-investment research: whether an autonomous system can use evidence from earlier experiments to improve the hypotheses and candidates proposed in later iterations. We present AQuA, which comprises two separate language-model-driven research systems: one for symbolic factor discovery and one for trainable model development. Each system records experimental results and uses them to guide subsequent proposals. Each operates in a fixed sandbox, which fixes the data splits, feature and label definitions, and evaluator while allowing the model to act only through constrained factor expressions or configuration diffs. The factor system, a manager-mediated multi-agent pipeline, discovers and combines factors into a signal that reaches a combined validation information coefficient of about $0.190$ on a crypto universe. The model system, a config-driven loop over a hybrid time-series architecture, reaches a per-stock information coefficient of $+0.0843$ on US equities and converts it into a threshold long/short strategy with a held-out Sharpe of up to $+2.50$ at a two-leg cost. The strategy is positive in every year from 2021 to 2025.

cs.CL↗

Critical behavior and critical exponents of rotating QCD matter

We investigate the thermodynamic properties and critical behavior of rotating strongly interacting matter within the two-flavor Nambu--Jona-Lasinio model. The phase structure and critical endpoint (CEP) are determined in the temperature-angular-velocity \((T,ω)\) plane. Exploiting the thermodynamic conjugacy between the angular velocity \(ω\) and rotational polarization \(J\), we characterize the CEP through complementary thermal, coexistence, susceptibility, and critical-isotherm responses. The corresponding effective critical exponents are extracted along distinct thermodynamic trajectories and are found to be consistent with \(α_ω\simeq0\), \(β_ω\simeq1/2\), \(γ_ω\simeq1\), and \(δ_ω\simeq3\). These exponents are numerically compatible with the standard scaling relations and form a mutually consistent Landau mean-field scaling pattern. Our results establish a coherent thermodynamic characterization of the rotation-induced CEP and demonstrate how its singular behavior is encoded in rotational response observables associated with the conjugate pair \((ω,J)\).

hep-ph↗

Defect-controlled twin activation in crystallographically equivalent magnesium micropillars

Tensile twinning plays a central role in accommodating -axis plasticity in Mg. In bulk Mg, it typically shows a relatively deterministic response with a low critical stress, whereas in confined volumes it exhibits broad yield-stress distributions that complicate the prediction of small-scale mechanical behavior. Here, site-specific compression tests are performed on 4 $μ$m-diameter pillars fabricated in a parent Mg crystal and an adjacent {10-12} twin. The two regions share the same [11-20] compression axis but experienced different prior deformation histories, allowing the influence of the residual microstructural state to be examined at fixed crystallographic orientation. Among 27 pillars, most parent-region pillars yield near 300 MPa, whereas pillars from the twin region span approximately 30 to 300 MPa. Interrupted tests combined with cross-sectional EBSD link individual load drops to discrete twin formation and further show that a pillar containing a pre-existing twin yields at approximately 80 MPa through the migration of the existing twin boundary. Molecular dynamics simulations of 30 nm-diameter pillars resolve possible atomistic pathways at the nanoscale. The simulations illustrate how contact geometry and pre-existing twin embryos alter event selection, and how an activated twin advances rapidly, while coherent twin boundary migration proceeds through disconnection motion accompanied by crystallographically required atomic shuffles. The results attribute the experimental scatter to the local availability of embryos and mobile interfaces, such that the first plastic event is governed by the twinning pathway accessible from the local microstructural state rather than by a single characteristic critical stress. Deformation history can therefore strongly modify the distribution of first plastic events even when the loading orientation is fixed.

cond-mat.mtrl-sci↗

Online Convex Optimization with Dueling Feedback

Noisy binary comparison between two candidates is a common interface between human and learning systems, especially in modern large language model (LLM) post-training alignment. We study online convex optimization with dueling (pairwise comparison) feedback, where the learner observes only a binary preference between two queried points. We consider adversarial sequences of convex losses and measure regret with the loss at both queried points, under a comparison link with a known nonzero slope at the origin. We propose a simple reduction that converts dueling feedback into approximate gradients, enabling the use of standard first-order methods. We show that regret guarantees transfer under this reduction, yielding $\mathcal O(T^{3/4})$ static and adaptive regret, and $\mathcal O(T^{3/4}\sqrt{1+P_T/D})$ dynamic regret with unknown comparator path length $P_T$. For strongly convex losses, the static and adaptive bounds improve to $\widetilde{\mathcal O}(T^{2/3})$. For smooth losses, we presents unified dueling ellipsoidal FTRL, and proves $\widetilde{\mathcal O}(T^{2/3})$ static regret, which improves to $\widetilde{\mathcal O}(\sqrt T)$ under additional strong convexity.

cs.LG↗

Analysis note: one-point charge correlator with DELPHI Open Data

We present the first measurement of the one-point charge correlator, the angular flux of electric charge in hadronic final states, using DELPHI Open Data collected at LEP-1 at $\sqrt{s} = 91.2$~GeV during 1994 and 1995. The data, corrected for detector effects, exhibit a clear $\sin(2θ)$ modulation, consistent with the parity-violating hadronic charge flow that the chiral structure of the $Z$ couplings imprints on the final state. The measurement demonstrates the experimental feasibility of the observable and establishes strategies for controlling associated detector effects, thereby motivating a new program to measure charge-flux observables. This note documents the experimental details supporting the companion experimental paper and the joint theory--experiment Letter.

hep-ex↗

Finitistic dimension via modules over the singularity category

We combine results of Rickard and Shaul with methods of intrinsic homological algebra and the theory of purity to prove that the finiteness of the big finitistic dimension of an Artin algebra is an intrinsic property of the singularity category of its opposite algebra, via its category of modules. This `object-free' approach complements work of Dey--Šťovíček and arrives at the same conclusion: the finiteness of the finitistic dimension of Artin algebras is a singular invariant (of the opposite algebras). In fact, our characterisation can be used to prove that finite finitistic dimension descends along certain fully faithful functors between singularity categories. We include an appendix where we explain how our methods can be used to prove non-existence of bounded t-structures on singularity categories.

math.RT↗

SubZero+: Memory-Efficient Adaptive Zeroth-Order LLM Fine-Tuning in Random Subspaces

Zeroth-order (ZO) optimization with SGD in random subspaces enables memory-efficient fine-tuning of large language models without backpropagation. However, high gradient estimation noise fundamentally undermines adaptive optimizers like Adam. We propose SubZero+, which achieves practical adaptive ZO optimization through a carefully designed dual low-dimensionality strategy: (i) multi-query forward-difference gradient estimation in periodically refreshed random subspaces to mitigate noise amplification in moment buffers, and (ii) Adam updates with periodic restarts performed directly in low-dimensional space rather than full-parameter space. In experiments, this dual design retains memory overhead comparable to momentum-free ZO methods while achieving stronger optimization performance than the evaluated ZO baselines. Theoretically, in the exact-directional limit, $K$-query averaging preserves conditional unbiasedness, while the coefficient estimator's covariance and mean-squared error, as well as query-induced second-moment inflation, scale exactly as $1/K$. Extensive experiments across SuperGLUE with models from 1.3B to 32B parameters under both full fine-tuning and LoRA schemes demonstrate consistent improvements over competing ZO methods. SubZero+ significantly narrows the performance gap with first-order optimization while preserving ZO's inference-time memory efficiency.

cs.LG↗

Bypassing the Chiral Obstruction in Two-Dimensional Tensor Networks

Gapped chiral phases pose an intrinsic obstruction to projected entangled pair state (PEPS) representations with finite bond dimension $D$, which generically develop spurious long-range correlations despite the gapped nature of the target state. To address this long-standing problem, we take a qualitatively different route by considering the target chiral system together with a completely decoupled time-reversed auxiliary copy. Constraining the finite-$D$ PEPS to the factorized form would simply reduce the problem to representing each chiral layer independently, leaving the original finite-$D$ obstruction unchanged. Unexpectedly, however, we find that finite-$D$ variational optimization generates an \emph{emergent} residual interlayer entanglement, despite the absence of any explicit interlayer coupling in the Hamiltonian. This emergent entanglement is the key mechanism that bypasses the obstruction, qualitatively changing the PEPS from an effectively gapless representation with long-range correlation tails to one with a finite correlation length. We demonstrate this mechanism for both a free-fermion Chern insulator and an interacting chiral spin liquid. The residual entanglement is systematically suppressed with increasing $D$, approaching the decoupled limit, while the chiral topological information remains accessible through layer-resolved entanglement spectra. Our results thus uncover a previously unrecognized mechanism for bypassing the chiral obstruction through variationally generated auxiliary entanglement.

cond-mat.str-el↗

Token Distribution versus Data Volume: Domain Balancing in Multi-Domain Meeting Summarisation

Jointly fine-tuning an LLM on meeting-summarisation corpora of widely varying size raises a question that prior work leaves confounded: when a domain-balanced training mixture helps, is the gain due to the distribution of tokens across domains, or merely to the volume of data seen? We disentangle these factors by constructing balanced and natural (native-proportional) token mixtures at matched token budgets (2-32M) over five English meeting corpora, fine-tuning Mistral-7B with QLoRA, and evaluating per domain. Balancing redistributes quality, improving the data-scarce minority domains at a low cost to the data-rich ones. The trade favours balancing whenever the minority domains matter: their share under proportional allocation is fixed at 1-2% regardless of budget, so matching balanced quality on those domains requires far more total data. We further find that pruning low-value transcript lines removes ~15% of tokens from the conversational corpora at no measurable cost, and that balancing by tokens is not the same as balancing by examples. Fine-tuning one model per domain is competitive only on the data-rich domains and falls below the zero-shot model on the data-scarce ones. A two-annotator study of 741 judge-labelled facts validates our fact-level evaluation. Together these results give practitioners a basis for deciding when to balance an imbalanced multi-domain mixture, and on what unit.

cs.CL↗

Optimal Weighted $L^2$ Hessian Estimates for Parabolic Equations under General Diffusion Marginals

We study weighted $L^2$ Hessian estimates for parabolic terminal equations under diffusion marginals. The reference diffusion determines the norm, and its generator may have different principal coefficients from the PDE. We characterize the limiting optimal Hessian constant uniformly over shrinking subintervals of a fixed time horizon, with zero terminal data. Under coefficient regularity, uniform ellipticity and bounded reference drift, finiteness characterizes linear well-posedness in the weighted parabolic Sobolev space on the full interval. For initial laws satisfying weighted doubling and coercivity conditions, the constant is the maximum of initial, flat and exponential-mixture contributions. Additional geometric regularity and nonnegative Ricci curvature identify the full mixture contribution, including spatial infinity, through terminal momenta of paths minimizing action plus the initial potential. In the matching case, with reference covariance $a$, PDE matrix $a/2$ and normalized Hessian $a^{1/2}D^2u\,a^{1/2}$, the constant is $2$ with time weight $t^α$, $α\ge1/2$, for every initial law; point starts attain $2\sqrt2$ at $α=0$.

math.PR↗