Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Flip-packability: uniform characterisations of tame graph classes

A class of graphs is monadically dependent if one cannot encode all graphs in coloured graphs from the class using a fixed first-order formula, and monadically stable if one cannot even encode arbitrarily long linear orders. Bonnet et al. (ICALP 2025) characterised monadic dependence by flip-separability: for every vertex weighting, boundedly many flips - complementations of the adjacency relation within a vertex subset - make every ball of radius $r$ carry at most an $\varepsilon$-fraction of the weight, so that every set carrying an $\varepsilon$-fraction has two elements pulled apart. We introduce flip-packability: boundedly many graphs, each obtained from the input by boundedly many flips and all determined by the weighting before any set is presented, such that every set carrying an $\varepsilon$-fraction of the weight has $m$ elements pairwise far apart in one of them. The number of flips producing each graph depends on the radius alone; only the number of graphs depends on $\varepsilon$ and $m$. We prove that a class of graphs is flip-packable if and only if it is monadically stable, and $2$-flip-packable, that is, flip-packable with $m=2$, if and only if it is monadically dependent. The passage from two scattered elements to $m$ is thus exactly what separates the two notions. For monadically stable classes we show that the flipped graphs can be computed from the weighting in cubic time. Varying the three parameters of the definition - the sparsifying operation, the radius, and the number $m$ of elements scattered - produces eight known characterisations of sparse and dense graph classes from the same template. In each case $m$ separates a depth-like notion from its width-like relaxation: treedepth from treewidth, shrubdepth from cliquewidth, and monadic stability from monadic dependence.

cs.LO↗

Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents

Executable environments are critical for post-training agents on tasks that require tool use and multi-step interaction, but constructing executable tasks together with their environments remains difficult to scale. Skills provide reusable domain knowledge, operational procedures, and tool-use instructions, but a substantial gap remains between the information contained in a skill and a concrete, challenging task with a complete executable environment. To address this gap, we introduce Skill2Env, a capability-oriented framework that starts from a skill and uses agent capability demands to guide task and environment synthesis. Skill2Env represents these demands through reusable difficulty patterns and instantiates them into task blueprints that specify objectives, challenges, environment facts, information boundaries, and acceptance criteria. These blueprints guide the joint construction of task instructions, execution substrates, workspaces, and rubric-based evaluators around source skills. We further propose Iterative Task Hardening, which uses solver execution evidence to identify insufficiently challenging task designs, strengthen or extend their difficulty-pattern instantiations, and revise the corresponding blueprints and environments. Using 1.5K high-scoring trajectories generated from Skill2Env environments for supervised fine-tuning, we observe consistent improvements across a broad range of agent benchmarks, demonstrating the effectiveness of capability-oriented environment synthesis for agent post-training.

cs.AI↗

Diffusion Reward Models

Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate or to a distribution from a fixed parametric family. This is at odds with human preference, which is inherently multimodal: the same response can be reasonably judged in many ways, and no single family covers all of them. To better fit this structure, we introduce DRM, a Diffusion Reward Model that recasts reward modeling as conditional density estimation over $p(\mathbf{r}\mid x,y)$. Conditioned on a frozen LLM encoder, a lightweight Diffusion Transformer denoises Gaussian noise into a reward vector, placing no parametric assumption on the output distribution and naturally representing its multimodal structure. A single architecture handles both multi-attribute regression and pairwise preference data, and at inference $N$ samples form an empirical reward distribution that can be aggregated into a scalar, a variance, or quantiles. Across five benchmarks, DRM matches or surpasses baselines under matched data and backbone, stays competitive with much larger discriminative, distributional, and generative RMs despite its modest training scale, and recovers multimodal reward structure where conventional heads collapse to a point. Uncertainty-aware rejection and lower-confidence-bound (LCB) aggregation further demonstrate that DRM can exploit distributional information beyond a scalar reward to improve reward-model decisions. Downstream RLHF experiments additionally show that using DRM as the training-time reward leads to improved policy performance, directly validating the practical benefit of diffusion-based reward modeling for RLHF training.

cs.LG↗

Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents

Outcome-only reinforcement learning gives software engineering (SWE) agents a terminal success signal but little direct guidance about intermediate decisions. We introduce Counterfactual Rollout Replay (CRR), a training-time procedure that uses forkable executable environments to obtain step-level return contrasts. CRR selects a small set of decision points, restores each state, samples an alternative action, and rolls the branch forward under the policy. It retains the realised training trajectory and replaces the advantage at selected steps with the difference between its terminal return and the sampled counterfactual return. The method needs no human process labels or learned process reward model; free refers to those supervision costs, not replay compute. With a 14B policy, CRR improves pass@1 on SWE-bench Verified, SWE-bench Live, and SWE-rebench, and combines with process-reward and trajectory-search methods. On SWE-bench Verified, an equal-wall-clock comparison on the same hardware yields 41.7% versus 36.7% for extended outcome-only GRPO, a 5.0-point gain with fork overhead included. These results apply to environments with affordable, reliable state restoration; stochastic continuations and expensive or imperfect replay remain limitations.

cs.SE↗

LLMs learn different forms of metacognition when trained to predict their own accuracy

Large language models are trained to always produce an answer, regardless of whether they possess the relevant knowledge, which leads them to fabricate facts. Prior work has shown that LLMs' confidence estimates correspond poorly to their actual performance, and that fine-tuning can substantially improve them. However, what models actually learn during such training remains poorly understood. We investigate how LLMs acquire metacognitive monitoring, the ability to know what one knows, by training 10 open-weight LLMs to predict their own accuracy on factual multiple-choice questions before answering them. We find that trained confidence reflects two distinct signals. While on questions close to the training data, it tracks the model's true accuracy, in other domains, it instead tracks output consistency: the concentration of the model's answer distribution. Output consistency tracking emerges early in training and generalizes across datasets, whereas accuracy tracking develops later and remains local to the training distribution. These results suggest that calibration training may not teach models to generally detect errors they commit confidently, and they raise broader questions about the nature of metacognition in artificial systems.

cs.CL↗

Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization

A single random Gaussian probe gives an unbiased estimate of the squared Frobenius norm of a layer's quantization error. The estimator is well-behaved because round-to-nearest error is spectrally flat. Across 1,683 tensors from a 35B MoE and a 9B dense model, effective dimensionality is 0.93 to 0.96 times the i.i.d. noise value of the same shape, and on the MoE the median is unchanged from 2-bit to 8-bit. The probe coefficient of variation is predictable from tensor shape. One probe measures per-tensor sensitivity to within 4 to 7%; twenty probes reach 1.3 to 1.4%.RAM applies the propagated form of this estimator to budget-targeted mixed-precision quantization with no calibration data. Gaussian probes carrying the network's own input statistics score every tensor at six bit-widths. A knapsack solver allocates bits under an exact byte budget, with guardrails against catastrophic 2-bit assignments. One probe pass serves any budget. Isolated and propagated scores rank tensors independently on Qwen3.5-35B-A3B (Spearman -0.01), yet the propagated probe rank-correlates 0.81 to 0.83 with the GPTQ layer objective from real activations, while the isolated estimator is uncorrelated with it. That objective is the wrong allocation target: at matched bytes on Qwen3.8-27B, a block-output probe beats a vendor IQ3_M mix and an oracle that allocates from the real-activation objective. On Qwen3-8B the propagated probe ties HAWQ-V2 at matched bytes. Across seven architectures from 8B to 122B, with probe timing up to a 400B model in nine minutes on one workstation, RAM reaches 3.5 to 13.6% lower median WikiText-2 perplexity than size-comparable uniform 4-bit builds on the tested MoE models. (Black Sheep Ai baa.ai)

cs.CL↗

On the Relevance of Incorporating Decision Dependence in Distributional Ambiguity

Most existing studies on distributionally robust optimization (DRO) with a decision-dependent ambiguity set focus on the computational and theoretical challenges posed by this class of problems. In this paper, we adopt a combined modeling and computational perspective to understand the trade-offs between modeling fidelity, solution quality, and computational effort, particularly in comparison with DRO models using decision-independent ambiguity sets. Motivated by representative applications in joint pricing-stocking newsvendor problems with price-dependent demand and facility location problems with location-dependent demand, we consider a two-stage stochastic mixed-integer program with (non)convex continuous recourse. Assuming a finite sample space, we model the decision-dependent distributional ambiguity with a polyhedral ambiguity set and reformulate the problem as a nonconvex mixed-integer nonlinear program. To efficiently solve the reformulations, we propose decomposition-based cutting-plane algorithms. Our experiments on benchmark instances indicate that a decision-dependent ambiguity set substantially reduces the postdecision disappointment for the newsvendor and out-of-sample cost for the facility location problems, up to 65% and 7% on average, respectively; thereby mitigating the optimizer's curse relative to a decision-independent DRO. However, this improvement is achieved at the expense of increased computational effort, with the absolute runtime ratio falling between 2 and 25 for medium- to large-sized newsvendor instances and between 2 and 12 for facility location instances.

math.OC↗

Greenpixie's AI Token Methodology: Assessing the Energy, Water and CO2-eq Impact of AI Tokens for Open and Closed Weight Models

We describe a methodology for estimating the per-token energy cost of cloud-hosted large language model (LLM) inference, separating between input (prefill) and output (decode) tokens. Graphics processing unit (GPU) energy usage is measured during inference benchmarking with open-weights models on a wide range of text-based tasks. The remaining server energy contribution from non-GPU hardware is estimated from the inference wall time. Bayesian linear regression is used to model the relationship between energy per token and LLM size, request traffic, and hardware deployment configuration. Proprietary frontier LLMs of unknown size and deployment are binned into size buckets based on naming conventions and performance priors, and the space of possible LLM configurations is sampled with Monte-Carlo methods to give a representative average energy per token and uncertainty. We also describe how these energy measurements can be used to estimate the carbon-dioxide equivalent ($\mathrm{CO_2\text{-}eq}$) emissions, both usage and embodied, and water consumed per token of AI inference. This methodology provides actionable data that enables reductions in cost, electricity usage, $\mathrm{CO_2\text{-}eq}$ emitted and water consumed in cloud and Software as a Service (SaaS).

cs.CY↗

FINGR: Learning Dexterous Hand Control for Real-World Rubik's Cube Solving

Manipulating a Rubik's Cube with a single dexterous hand is a challenging test of sustained, contact-rich control: the hand must execute successive layer turns while keeping the cube secure. Each turn requires some fingers to support the cube while others push a moving layer, release contact, and reset for the next move. To learn this coordination, we introduce FINGR (Future-supervised Interaction Network with Geometric Representations), a policy that combines finger-relative geometry with future interaction prediction. A shared point encoder expresses the cube relative to each fingertip and aggregates its points without depending on cubie indexing. Learned future tokens share the observation encoder and receive supervision for contact-force changes, layer-turn progress, and finger joint displacement at multiple time scales. The resulting representation conditions a flow policy that directly generates finger actions. On a real dexterous hand, our policy achieves 99.0% success over 300 turn attempts, compared with 79.7% for the base flow policy. Integrated with grasping and table-assisted regrasping, the policy solves all ten scrambled $2\times2\times2$ cubes in a mean complete-system time of approximately 137 seconds. The project website is available at https://www.lyt0112.com/projects/FINGR

cs.RO↗

Uncovering Non-Normality in Information Flow: Network Structure and Dynamics of Social Media Cascades

Information cascades on social media are conventionally conceptualized as directed, feedforward branching processes. However, real-world diffusion pathways frequently deviate from pure hierarchical trees due to localized clustering, reciprocal commentary, and multi-wave temporal surges. In this work, we quantify the directional asymmetry and hierarchical structure of empirical information cascades on X (formerly Twitter) using spectral non-normality via Henrici's departure from normality. Analyzing approximately 58,000 cascade networks across diverse topics (including politics, entertainment, natural disasters, etc.), we investigate (1) how non-normality relates to temporal dynamics such as endogenous-like versus exogenous-like patterns and burstiness, (2) whether non-normality is correlated with the peak concentration or overall size of a cascade, (3) whether the overall non-normality of a cascade's network structure can be predicted from its early stages. We find that non-normality strongly aligns with peak concentration (peak/N) rather than overall cascade size, characterizing cascades governed by rapid, asymmetric forwarding. Furthermore, while early-stage structural forecasting (<= 30% of nodes observed) exhibits expected baseline uncertainty (51%-72% accuracy at a +/- 20% error tolerance), predictability consolidates rapidly during intermediate growth, exceeding 80% across all dynamic clusters once 50%-60% of the network is observed. By identifying the topological and dynamic correlates of cascade structures, this study advances our understanding of information flow and establishes a quantifiable benchmark for forecasting directional diffusion architectures.

cs.SI↗

Discovery of a 62-min long-period transient with near-orthogonal rotator geometry

We report the serendipitous discovery of J155543.7$-$563102 (hereafter LPT J1555$-$5631) with a period of 62.2 min in MeerKAT and ASKAP survey data. The source emitted bright (10-30 mJy), steep-spectrum ($α\sim-1.5$), highly polarized radio pulses during a 15-hr active window in May 2022, with no other quiescent or pulsed emission detected before or after over $\sim5$ years. The mean light curve exhibits two alternating pulses per cycle. A main pulse (MP) and an interpulse (IP) separated by $Δϕ=165.4^\circ$ with similar amplitudes, but different polarization properties: the wider MP (3.9%) is partially linearly and circularly polarized and shows a constant polarization position angle, while the narrow IP (1.1%) is $\sim$100% circularly polarized. We classify LPT J1555$-$5631 as a near-orthogonal rotator, for which emission from both magnetic poles is visible during one rotation and the magnetic and rotation axes are almost perpendicular. The 14.6$^\circ$ deviation from an antipodal separation and the nearly constant polarization position angle suggest that the emission does not arise from the standard dipolar polar-cap geometry of ordinary pulsars, but instead originates within a broader, more complex magnetosphere. The observed properties of LPT J1555$-$5631 can be accommodated by extended emission regions, and favor coherent emission mechanisms such as the electron cyclotron maser instability, although they do not uniquely determine the progenitor or emission physics. If these inferences apply more broadly to the LPT population, they could account for why their beam widths are decoupled from rotation periods, the prevalence of flat polarization position angles, and may favor an elevated incidence of orthogonal rotators.

astro-ph.HE↗

GUITAR: Structured Failure Diagnosis of GUI Agents via State Transitions

Understanding where and why Graphical User Interface (GUI) agents fail is essential for building more reliable systems, yet current evaluation relies on step accuracy, a metric that treats each screen independently and overlooks the underlying structure of GUI environments. This leads to two critical blind spots: (1) functionally equivalent screens are evaluated in isolation, obscuring systematic failure patterns across shared screens; and (2) the long-tailed GUI distribution renders failures on rare but critical screens invisible under standard metrics. To address these issues, we propose \textbf{GUITAR}, a state-centric diagnostic framework that performs structured failure analysis over both states and transitions, using a State Transition Graph (STG) by mapping visually diverse screens to shared functional states. Across 8 agents and 6 tasks from AndroidControl and Mind2Web, GUITAR reveals that 60.4\% of failures occur in 20\% of states, localizing errors to a small set of bottlenecks. Bottleneck-targeted guidance improves SR by 2.8\% and retains a 1.88\% average gain across 7 agents under three-fold trajectory-held-out evaluation with fully automatic STGs. These findings demonstrate the diagnostic and actionable value of structure-aware evaluation within the evaluated mobile and web tasks. Code is available at https://github.com/sqzhang-lazy/GUITAR

cs.AI↗

Basis Functions for Time-Dependent Kohn-Sham Inversion

Floquet theory provides insight into the inversion of time-dependent Kohn-Sham density functional theory. Specifically, mathematical derivations show that the fundamental frequencies of a time-dependent wavefunction solution are the leading-order harmonics for the TD-KS state. Numerical tests of the resulting ansatz in 1D and 3D for atomic and molecular cases demonstrate its utility. In particular, low $L_{2}$ errors in the time-dependent density and longitudinal current were found, even though currents were not an explicit optimization objective. In all, the proposed inversion ansatz provides exchange-correlation potentials from time-dependent wavefunctions, is highly interpretable, and may significantly help in the development of nonadiabatic density functionals.

physics.chem-ph↗

Natural Image Autoencoder-Based fMRI Representations for Trait and State Prediction

Foundation models pre-trained on large-scale fMRI datasets have shown strong downstream performance, but at substantial data and computation cost. To investigate how much fMRI-specific pre-training is actually needed for such performance, we introduce FReD, which derives fMRI representations from a frozen Deep Compression AutoEncoder (DCAE) pre-trained exclusively on natural images and pairs them with a task specific readout. For trait prediction, FReD summarizes frame-wise representations by their temporal mean and log-standard deviation and applies linear probing, with late fusion across two normalization schemes. For state prediction, it represents each frame as a single token and models temporal dependencies with a shallow Transformer. Across four resting-state datasets spanning six trait-prediction targets, linear probes on frozen DCAE features generally outperform those on fMRI foundation model representations and remain competitive with fully fine-tuned fMRI foundation models. On three task-fMRI state-prediction tasks, a temporal readout on DCAE features performs comparably to the strongest foundation models evaluated. A Gaussian injection analysis further shows that localized signal changes are recovered more accurately from the frozen DCAE features than from the evaluated foundation-model representations. Together, these results show that strong performance on current fMRI benchmarks is possible without fMRI-specific representation pre-training, making frozen natural-image features as a useful baseline for assessing its added value.

cs.CV↗

GradLev: Token-Parallel Test-Time Training Via Costate Prediction

Test-time training (TTT) allows a model to improve its predictions at inference time by updating weights after every observed token. However, sequential gra- dient writes make parallel training difficult. We observe that, given layer inputs and activation gradients (costates), online gradient descent admits exact parallel scans for both forward evaluation and reverse backpropagation. GradLev lever- ages this duality: a causal auxiliary network predicts costates across all tokens in parallel; associative scans compute the adapted weights and forward activations and propagate gradients backward; and the resulting gradient targets supervise the predictor via a consistency loss. Exact consistency guarantees exact recovery of the sequential online learner. At deployment, the auxiliary predictor is discarded, and the model updates natively via token-by-token forward and backward passes.

cs.LG↗

Unified Visual-Tactile-Action Modeling from Human Demonstrations for Dexterous Manipulation

Dexterous manipulation requires tactile feedback. However, robot tactile demonstrations are difficult to scale,because dexterous-hand teleoperation provides limited tactile feedback to the operator. In contrast, human demonstrations offer a substantially more scalable source of diverse tactile interactions. Motivated by a simple premise: hands can change, but the underlying physics of interaction does not. We leverage human tactile data to improve dexterous manipulation policies. Specifically, we first build a tactile motion-capture system that synchronously records images, tactile signals, and hand motions. Using this system, we construct the UVTA dataset spanning five contact-rich tasks, with 1,000 human demonstrations covering diverse interaction patterns and 150 robot demonstrations per task. To transfer the underlying physics of human interaction to robot control, we propose a Unified Visual-Tactile-Action Model that maps both embodiments into aligned tactile and action representations and jointly predicts future action and tactile trajectories. The joint objective enables human demonstrations to supervise contact-aware representation learning, while only robot actions are executed during deployment. In real-robot evaluations across five tasks, our method achieves an average success rate of 70%, outperforming the strongest visual-tactile baseline, which achieves 29%, and an architecture ablation, which achieves 42%. Performance improves consistently with additional human demonstrations and exhibits no saturation at 1,000 demonstrations per task, validating the effectiveness of scalable human tactile data for dexterous manipulation. Project page is available at https://uni-vta.github.io/.

cs.RO↗

LLMs are not stochastic parrots: Evidence for meaning-mediated abstraction from conlang-like tasks

The strong version of the stochastic parrot argument claims that, although large language models (LLMs) may exceed rote regurgitation, they cannot move beyond statistical pattern matching into abstraction or reasoning, remaining ontologically near the lower bound of pattern reuse despite producing alluringly fluent text. We test this hypothesis using conlang-like tasks. Several LLMs are given only natural-language descriptions of fictional languages that subvert prominent superficial patterns in training data by combining statistically uncommon and unattested features. Crucially, no example outputs are given. We argue that if the models exhibit rule-following behaviour, they cannot be relying solely on superficial statistical patterns; such patterns often work against the correct output. Instead, successful performance requires representations of the constraints specified in the prompt. Across three complementary task families, models systematically move in the meaning-predicted direction: they distinguish prompt exposure from instructed use, alter semantic relationships in response to novel constraints, and sometimes produce exact matches to complex translation answer keys. Although performance varies across the spectrum of models used, these results provide evidence for meaning-mediated abstraction in LLMs and refute the strong stochastic parrot hypothesis. Our work shows that, under appropriate architectural and contextual constraints, statistical learning can produce meaning-mediated abstractions, although generation remains strongly constrained by superficial plausibility. We discuss implications for model development and for understanding how increasingly abstract representations may emerge from plausible-text-generation objectives.

cs.CL↗

Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation

Language-model judges compare agent upgrades with their predecessors, but a fixed judge can make version-dependent mistakes. We analyze 35 public coding-agent submissions (20 prespecified version pairs on 250 SWE-bench Verified issues), two customer-service agents (155 tau-bench tasks), and 1,106 expert-labeled AgentRewardBench trajectories. An upstream outage left three judges for the primary SWE-bench analysis (8,743 aligned cells); the fourth is descriptive. All three coding-agent judges and all four tau-bench judges reject task-conditioned error invariance after multiplicity adjustment. On SWE-bench, 32 of 60 judge-by-pair units have a detectable differential comparison component; eight judge-only intervals declare upgrades that execution-based intervals cannot establish, despite rank correlations of 0.71-0.79. In tau-bench, one judge reverses a nine-point reference-reward gap by penalizing a procedural habit the reward ignores. False acceptance of failed coding patches rises with agent capability conditional on task and execution outcome, while a task-solvability prediction reverses sign across domains. A separately fixed post-submission OpenHands follow-up on the same 250 issues (eight configurations, 1,981 three-judge cells) reproduces the capability/false-acceptance association (mean Spearman +0.944, exact p=0.000099) and decreasing Youden contrast (mean -0.937, p=0.000397); this is observational, not a new-task replication. Transporting old-version calibration raises SWE-bench comparison error from 3.8 to 19.5 points, with 24.6% undefined bootstrap ratios. A paired audit saves only 5% in interval width at 80 labeled tasks. A randomized self-report test is negative (three adjusted p-values=1.0). These results favor paired audits of current outputs over judge-only release decisions or transported calibration; independent human patch review remains pending.

cs.LG↗