Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,621 records · Page 90Linked to original sources

LLMs learn different forms of metacognition when trained to predict their own accuracy

Large language models are trained to always produce an answer, regardless of whether they possess the relevant knowledge, which leads them to fabricate facts. Prior work has shown that LLMs' confidence estimates correspond poorly to their actual performance, and that fine-tuning can substantially improve them. However, what models actually learn during such training remains poorly understood. We investigate how LLMs acquire metacognitive monitoring, the ability to know what one knows, by training 10 open-weight LLMs to predict their own accuracy on factual multiple-choice questions before answering them. We find that trained confidence reflects two distinct signals. While on questions close to the training data, it tracks the model's true accuracy, in other domains, it instead tracks output consistency: the concentration of the model's answer distribution. Output consistency tracking emerges early in training and generalizes across datasets, whereas accuracy tracking develops later and remains local to the training distribution. These results suggest that calibration training may not teach models to generally detect errors they commit confidently, and they raise broader questions about the nature of metacognition in artificial systems.

cs.CL↗

Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization

A single random Gaussian probe gives an unbiased estimate of the squared Frobenius norm of a layer's quantization error. The estimator is well-behaved because round-to-nearest error is spectrally flat. Across 1,683 tensors from a 35B MoE and a 9B dense model, effective dimensionality is 0.93 to 0.96 times the i.i.d. noise value of the same shape, and on the MoE the median is unchanged from 2-bit to 8-bit. The probe coefficient of variation is predictable from tensor shape. One probe measures per-tensor sensitivity to within 4 to 7%; twenty probes reach 1.3 to 1.4%.RAM applies the propagated form of this estimator to budget-targeted mixed-precision quantization with no calibration data. Gaussian probes carrying the network's own input statistics score every tensor at six bit-widths. A knapsack solver allocates bits under an exact byte budget, with guardrails against catastrophic 2-bit assignments. One probe pass serves any budget. Isolated and propagated scores rank tensors independently on Qwen3.5-35B-A3B (Spearman -0.01), yet the propagated probe rank-correlates 0.81 to 0.83 with the GPTQ layer objective from real activations, while the isolated estimator is uncorrelated with it. That objective is the wrong allocation target: at matched bytes on Qwen3.8-27B, a block-output probe beats a vendor IQ3_M mix and an oracle that allocates from the real-activation objective. On Qwen3-8B the propagated probe ties HAWQ-V2 at matched bytes. Across seven architectures from 8B to 122B, with probe timing up to a 400B model in nine minutes on one workstation, RAM reaches 3.5 to 13.6% lower median WikiText-2 perplexity than size-comparable uniform 4-bit builds on the tested MoE models. (Black Sheep Ai baa.ai)

cs.CL↗

On the Relevance of Incorporating Decision Dependence in Distributional Ambiguity

Most existing studies on distributionally robust optimization (DRO) with a decision-dependent ambiguity set focus on the computational and theoretical challenges posed by this class of problems. In this paper, we adopt a combined modeling and computational perspective to understand the trade-offs between modeling fidelity, solution quality, and computational effort, particularly in comparison with DRO models using decision-independent ambiguity sets. Motivated by representative applications in joint pricing-stocking newsvendor problems with price-dependent demand and facility location problems with location-dependent demand, we consider a two-stage stochastic mixed-integer program with (non)convex continuous recourse. Assuming a finite sample space, we model the decision-dependent distributional ambiguity with a polyhedral ambiguity set and reformulate the problem as a nonconvex mixed-integer nonlinear program. To efficiently solve the reformulations, we propose decomposition-based cutting-plane algorithms. Our experiments on benchmark instances indicate that a decision-dependent ambiguity set substantially reduces the postdecision disappointment for the newsvendor and out-of-sample cost for the facility location problems, up to 65% and 7% on average, respectively; thereby mitigating the optimizer's curse relative to a decision-independent DRO. However, this improvement is achieved at the expense of increased computational effort, with the absolute runtime ratio falling between 2 and 25 for medium- to large-sized newsvendor instances and between 2 and 12 for facility location instances.

math.OC↗

Greenpixie's AI Token Methodology: Assessing the Energy, Water and CO2-eq Impact of AI Tokens for Open and Closed Weight Models

We describe a methodology for estimating the per-token energy cost of cloud-hosted large language model (LLM) inference, separating between input (prefill) and output (decode) tokens. Graphics processing unit (GPU) energy usage is measured during inference benchmarking with open-weights models on a wide range of text-based tasks. The remaining server energy contribution from non-GPU hardware is estimated from the inference wall time. Bayesian linear regression is used to model the relationship between energy per token and LLM size, request traffic, and hardware deployment configuration. Proprietary frontier LLMs of unknown size and deployment are binned into size buckets based on naming conventions and performance priors, and the space of possible LLM configurations is sampled with Monte-Carlo methods to give a representative average energy per token and uncertainty. We also describe how these energy measurements can be used to estimate the carbon-dioxide equivalent ($\mathrm{CO_2\text{-}eq}$) emissions, both usage and embodied, and water consumed per token of AI inference. This methodology provides actionable data that enables reductions in cost, electricity usage, $\mathrm{CO_2\text{-}eq}$ emitted and water consumed in cloud and Software as a Service (SaaS).

cs.CY↗

FINGR: Learning Dexterous Hand Control for Real-World Rubik's Cube Solving

Manipulating a Rubik's Cube with a single dexterous hand is a challenging test of sustained, contact-rich control: the hand must execute successive layer turns while keeping the cube secure. Each turn requires some fingers to support the cube while others push a moving layer, release contact, and reset for the next move. To learn this coordination, we introduce FINGR (Future-supervised Interaction Network with Geometric Representations), a policy that combines finger-relative geometry with future interaction prediction. A shared point encoder expresses the cube relative to each fingertip and aggregates its points without depending on cubie indexing. Learned future tokens share the observation encoder and receive supervision for contact-force changes, layer-turn progress, and finger joint displacement at multiple time scales. The resulting representation conditions a flow policy that directly generates finger actions. On a real dexterous hand, our policy achieves 99.0% success over 300 turn attempts, compared with 79.7% for the base flow policy. Integrated with grasping and table-assisted regrasping, the policy solves all ten scrambled $2\times2\times2$ cubes in a mean complete-system time of approximately 137 seconds. The project website is available at https://www.lyt0112.com/projects/FINGR

cs.RO↗

Selfless nonnuclear crossed products

We prove that strictly ergodic actions of countable groups with Ozawa's property PHP on the Cantor set with dynamical comparison have completely selfless reduced crossed products. We deduce that these crossed products have stable rank one, strict comparison, a unique quasitracial state, and real rank zero. This covers actions of C*-simple acylindrically hyperbolic and linear groups with an amenable subgroup whose restricted action is strictly ergodic and, in particular, actions in the class $\mathrm{A}^*(F_d,X)$ introduced by Bell, Geffen, and Kerr.

math.OA↗

Uncovering Non-Normality in Information Flow: Network Structure and Dynamics of Social Media Cascades

Information cascades on social media are conventionally conceptualized as directed, feedforward branching processes. However, real-world diffusion pathways frequently deviate from pure hierarchical trees due to localized clustering, reciprocal commentary, and multi-wave temporal surges. In this work, we quantify the directional asymmetry and hierarchical structure of empirical information cascades on X (formerly Twitter) using spectral non-normality via Henrici's departure from normality. Analyzing approximately 58,000 cascade networks across diverse topics (including politics, entertainment, natural disasters, etc.), we investigate (1) how non-normality relates to temporal dynamics such as endogenous-like versus exogenous-like patterns and burstiness, (2) whether non-normality is correlated with the peak concentration or overall size of a cascade, (3) whether the overall non-normality of a cascade's network structure can be predicted from its early stages. We find that non-normality strongly aligns with peak concentration (peak/N) rather than overall cascade size, characterizing cascades governed by rapid, asymmetric forwarding. Furthermore, while early-stage structural forecasting (<= 30% of nodes observed) exhibits expected baseline uncertainty (51%-72% accuracy at a +/- 20% error tolerance), predictability consolidates rapidly during intermediate growth, exceeding 80% across all dynamic clusters once 50%-60% of the network is observed. By identifying the topological and dynamic correlates of cascade structures, this study advances our understanding of information flow and establishes a quantifiable benchmark for forecasting directional diffusion architectures.

cs.SI↗

Finite-Size Effects on Symmetry Restoration in a Rotating Scalar System with Spatially Dependent Thermal Self-Energy

We investigate $\mathbb{Z}_2$ symmetry restoration in a real $λϕ^4$ theory confined to a finite cylindrical region undergoing rigid rotation. The finite transverse extent is imposed through Dirichlet boundary conditions, resulting in a discrete Fourier-Bessel spectrum and an explicit radial dependence of the fluctuation propagator. Using the background-field method, we derive the one-loop effective potential while retaining the spatial structure induced by the finite geometry. At finite temperature, rotation modifies the thermal mode energies through angular-momentum-dependent shifts, and the coincident-point propagator generates a position-dependent thermal self-energy that is incorporated through ring resummation. We use the resulting spatially resolved effective potential to determine the symmetry-restoration temperature $T_c(Ω,R)$ and examine its dependence on the angular velocity and transverse size. The results show that the finite transverse size of the system modifies the symmetry-restoration temperature, leading to a nontrivial dependence of $T_c(Ω,R)$ on $R$. An additional scaling behavior emerges when the normalized critical temperature is expressed in terms of the boundary velocity $ΩR$: systems with different transverse sizes exhibit the same relative modification of $T_c$ when their boundary velocities are equal. This scaling persists after the inclusion of the position-dependent thermal self-energy through ring resummation, indicating that the $ΩR$ dependence of the rotational response is maintained in the interacting finite system.

hep-ph↗

Discovery of a 62-min long-period transient with near-orthogonal rotator geometry

We report the serendipitous discovery of J155543.7$-$563102 (hereafter LPT J1555$-$5631) with a period of 62.2 min in MeerKAT and ASKAP survey data. The source emitted bright (10-30 mJy), steep-spectrum ($α\sim-1.5$), highly polarized radio pulses during a 15-hr active window in May 2022, with no other quiescent or pulsed emission detected before or after over $\sim5$ years. The mean light curve exhibits two alternating pulses per cycle. A main pulse (MP) and an interpulse (IP) separated by $Δϕ=165.4^\circ$ with similar amplitudes, but different polarization properties: the wider MP (3.9%) is partially linearly and circularly polarized and shows a constant polarization position angle, while the narrow IP (1.1%) is $\sim$100% circularly polarized. We classify LPT J1555$-$5631 as a near-orthogonal rotator, for which emission from both magnetic poles is visible during one rotation and the magnetic and rotation axes are almost perpendicular. The 14.6$^\circ$ deviation from an antipodal separation and the nearly constant polarization position angle suggest that the emission does not arise from the standard dipolar polar-cap geometry of ordinary pulsars, but instead originates within a broader, more complex magnetosphere. The observed properties of LPT J1555$-$5631 can be accommodated by extended emission regions, and favor coherent emission mechanisms such as the electron cyclotron maser instability, although they do not uniquely determine the progenitor or emission physics. If these inferences apply more broadly to the LPT population, they could account for why their beam widths are decoupled from rotation periods, the prevalence of flat polarization position angles, and may favor an elevated incidence of orthogonal rotators.

astro-ph.HE↗

GUITAR: Structured Failure Diagnosis of GUI Agents via State Transitions

Understanding where and why Graphical User Interface (GUI) agents fail is essential for building more reliable systems, yet current evaluation relies on step accuracy, a metric that treats each screen independently and overlooks the underlying structure of GUI environments. This leads to two critical blind spots: (1) functionally equivalent screens are evaluated in isolation, obscuring systematic failure patterns across shared screens; and (2) the long-tailed GUI distribution renders failures on rare but critical screens invisible under standard metrics. To address these issues, we propose \textbf{GUITAR}, a state-centric diagnostic framework that performs structured failure analysis over both states and transitions, using a State Transition Graph (STG) by mapping visually diverse screens to shared functional states. Across 8 agents and 6 tasks from AndroidControl and Mind2Web, GUITAR reveals that 60.4\% of failures occur in 20\% of states, localizing errors to a small set of bottlenecks. Bottleneck-targeted guidance improves SR by 2.8\% and retains a 1.88\% average gain across 7 agents under three-fold trajectory-held-out evaluation with fully automatic STGs. These findings demonstrate the diagnostic and actionable value of structure-aware evaluation within the evaluated mobile and web tasks. Code is available at https://github.com/sqzhang-lazy/GUITAR

cs.AI↗

Basis Functions for Time-Dependent Kohn-Sham Inversion

Floquet theory provides insight into the inversion of time-dependent Kohn-Sham density functional theory. Specifically, mathematical derivations show that the fundamental frequencies of a time-dependent wavefunction solution are the leading-order harmonics for the TD-KS state. Numerical tests of the resulting ansatz in 1D and 3D for atomic and molecular cases demonstrate its utility. In particular, low $L_{2}$ errors in the time-dependent density and longitudinal current were found, even though currents were not an explicit optimization objective. In all, the proposed inversion ansatz provides exchange-correlation potentials from time-dependent wavefunctions, is highly interpretable, and may significantly help in the development of nonadiabatic density functionals.

physics.chem-ph↗

Natural Image Autoencoder-Based fMRI Representations for Trait and State Prediction

Foundation models pre-trained on large-scale fMRI datasets have shown strong downstream performance, but at substantial data and computation cost. To investigate how much fMRI-specific pre-training is actually needed for such performance, we introduce FReD, which derives fMRI representations from a frozen Deep Compression AutoEncoder (DCAE) pre-trained exclusively on natural images and pairs them with a task specific readout. For trait prediction, FReD summarizes frame-wise representations by their temporal mean and log-standard deviation and applies linear probing, with late fusion across two normalization schemes. For state prediction, it represents each frame as a single token and models temporal dependencies with a shallow Transformer. Across four resting-state datasets spanning six trait-prediction targets, linear probes on frozen DCAE features generally outperform those on fMRI foundation model representations and remain competitive with fully fine-tuned fMRI foundation models. On three task-fMRI state-prediction tasks, a temporal readout on DCAE features performs comparably to the strongest foundation models evaluated. A Gaussian injection analysis further shows that localized signal changes are recovered more accurately from the frozen DCAE features than from the evaluated foundation-model representations. Together, these results show that strong performance on current fMRI benchmarks is possible without fMRI-specific representation pre-training, making frozen natural-image features as a useful baseline for assessing its added value.

cs.CV↗

GradLev: Token-Parallel Test-Time Training Via Costate Prediction

Test-time training (TTT) allows a model to improve its predictions at inference time by updating weights after every observed token. However, sequential gra- dient writes make parallel training difficult. We observe that, given layer inputs and activation gradients (costates), online gradient descent admits exact parallel scans for both forward evaluation and reverse backpropagation. GradLev lever- ages this duality: a causal auxiliary network predicts costates across all tokens in parallel; associative scans compute the adapted weights and forward activations and propagate gradients backward; and the resulting gradient targets supervise the predictor via a consistency loss. Exact consistency guarantees exact recovery of the sequential online learner. At deployment, the auxiliary predictor is discarded, and the model updates natively via token-by-token forward and backward passes.

cs.LG↗

Unified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous Manipulation

Dexterous manipulation requires tactile feedback. However, robot tactile demonstrations are difficult to scale,because dexterous-hand teleoperation provides limited tactile feedback to the operator. In contrast, human demonstrations offer a substantially more scalable source of diverse tactile interactions. Motivated by a simple premise: hands can change, but the underlying physics of interaction does not. We leverage human tactile data to improve dexterous manipulation policies. Specifically, we first build a tactile motion-capture system that synchronously records images, tactile signals, and hand motions. Using this system, we construct the UVTA dataset spanning five contact-rich tasks, with 1,000 human demonstrations covering diverse interaction patterns and 150 robot demonstrations per task. To transfer the underlying physics of human interaction to robot control, we propose a Unified Visual-Tactile-Action Model that maps both embodiments into aligned tactile and action representations and jointly predicts future action and tactile trajectories. The joint objective enables human demonstrations to supervise contact-aware representation learning, while only robot actions are executed during deployment. In real-robot evaluations across five tasks, our method achieves an average success rate of 70%, outperforming the strongest visual-tactile baseline, which achieves 29%, and an architecture ablation, which achieves 42%. Performance improves consistently with additional human demonstrations and exhibits no saturation at 1,000 demonstrations per task, validating the effectiveness of scalable human tactile data for dexterous manipulation. Project page is available at https://uni-vta.github.io/.

cs.RO↗

LLMs are not stochastic parrots: Evidence for meaning-mediated abstraction from conlang-like tasks

The strong version of the stochastic parrot argument claims that, although large language models (LLMs) may exceed rote regurgitation, they cannot move beyond statistical pattern matching into abstraction or reasoning, remaining ontologically near the lower bound of pattern reuse despite producing alluringly fluent text. We test this hypothesis using conlang-like tasks. Several LLMs are given only natural-language descriptions of fictional languages that subvert prominent superficial patterns in training data by combining statistically uncommon and unattested features. Crucially, no example outputs are given. We argue that if the models exhibit rule-following behaviour, they cannot be relying solely on superficial statistical patterns; such patterns often work against the correct output. Instead, successful performance requires representations of the constraints specified in the prompt. Across three complementary task families, models systematically move in the meaning-predicted direction: they distinguish prompt exposure from instructed use, alter semantic relationships in response to novel constraints, and sometimes produce exact matches to complex translation answer keys. Although performance varies across the spectrum of models used, these results provide evidence for meaning-mediated abstraction in LLMs and refute the strong stochastic parrot hypothesis. Our work shows that, under appropriate architectural and contextual constraints, statistical learning can produce meaning-mediated abstractions, although generation remains strongly constrained by superficial plausibility. We discuss implications for model development and for understanding how increasingly abstract representations may emerge from plausible-text-generation objectives.

cs.CL↗

Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

Language-model judges compare agent upgrades with their predecessors, but a fixed judge can make version-dependent mistakes. We analyze 35 public coding-agent submissions (20 prespecified version pairs on 250 SWE-bench Verified issues), two customer-service agents (155 tau-bench tasks), and 1,106 expert-labeled AgentRewardBench trajectories. An upstream outage left three judges for the primary SWE-bench analysis (8,743 aligned cells); the fourth is descriptive. All three coding-agent judges and all four tau-bench judges reject task-conditioned error invariance after multiplicity adjustment. On SWE-bench, 32 of 60 judge-by-pair units have a detectable differential comparison component; eight judge-only intervals declare upgrades that execution-based intervals cannot establish, despite rank correlations of 0.71-0.79. In tau-bench, one judge reverses a nine-point reference-reward gap by penalizing a procedural habit the reward ignores. False acceptance of failed coding patches rises with agent capability conditional on task and execution outcome, while a task-solvability prediction reverses sign across domains. A separately fixed post-submission OpenHands follow-up on the same 250 issues (eight configurations, 1,981 three-judge cells) reproduces the capability/false-acceptance association (mean Spearman +0.944, exact p=0.000099) and decreasing Youden contrast (mean -0.937, p=0.000397); this is observational, not a new-task replication. Transporting old-version calibration raises SWE-bench comparison error from 3.8 to 19.5 points, with 24.6% undefined bootstrap ratios. A paired audit saves only 5% in interval width at 80 labeled tasks. A randomized self-report test is negative (three adjusted p-values=1.0). These results favor paired audits of current outputs over judge-only release decisions or transported calibration; independent human patch review remains pending.

cs.LG↗

RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineers

Coding agents are beginning to move beyond purely digital tasks to tackle physical-world challenges, particularly in robotics. Existing robotics benchmarks, however, primarily focus on the performance of individual artifacts, such as policies or controllers, offering limited coverage of coding agents' broader engineering capabilities. Real-world robotics extends beyond control: agents must build, integrate, diagnose, and improve heterogeneous artifacts under resource constraints and reason from multimodal feedback. To evaluate these broader capabilities, we introduce RLE-Bench, a benchmark of robot-learning tasks spanning four representative robotics development workflows: interactive control, policy learning, perception and estimation, and mechanical design. We use diverse task-specific metrics to evaluate the artifacts submitted by the coding agents, from the success rate the agents achieved to the policy agents trained, the harness agent built, and the mechanical structures the agent designed. We aggregate these metrics into an overall RLE Index and report workflow-specific capability profiles, enabling systematic comparison of coding agents' capabilities across multiple capability dimensions. Beyond performance ranks, we also conduct in-depth case studies examining agent behavior on representative tasks, highlighting both current capabilities and limitations, and pointing to the opportunities robotics tasks have to offer for future agent training.

cs.RO↗

Coherence-Aware Distributional Evaluation of Open-Ended Text Generation

Existing open-ended generation metrics measure likelihood, lexical diversity, or distributional similarity in generic representation space, yet can miss fundamental dimensions of quality. A prominent blind spot is global coherence: a generated passage may be locally fluent while remaining globally contradictory, causally inconsistent, or topically disconnected. We identify representation as a central bottleneck in detecting these failures and introduce CHORD (Coherence-aware Hidden-state Open-generation Reference Distance), a coherence-sensitive distributional metric. CHORD encodes generated and human-written corpora in the hidden-state space of a frozen LLM using a coherence-eliciting prompt, and compares the resulting distributions using RBF-MMD. To test coherence sensitivity and selectivity, we construct a counterfactual evaluation suite pairing graded coherence-degrading perturbations with meaning-preserving controls. CHORD selectively detects relation, discourse, structural, and mixture failures that perplexity, entropy, MAUVE, FBD, and MMD-based baselines either miss or cannot separate from benign rewriting. Factorial ablations show that representation is the primary source of coherence sensitivity, while RBF-MMD improves sample efficiency. Larger backbones capture finer-grained distinctions, but coherence prompting improves selectivity only when the backbone can follow the prompt. On unconditional generation and prefix continuation, CHORD yields model rankings that strongly align with human judgments of whether outputs make sense and appear human-written. Together, these results establish representation design as central to reliable distributional evaluation. Code: https://github.com/MAPS-research/CHORD. Experiments: https://github.com/MAPS-research/CHORD-Experiment.

cs.CL↗