Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,603 records · Page 89Linked to original sources

Profit Reallocation Mechanisms in Tree-based Data Trading

Markets for data promise to unlock its economic value, yet in practice they remain far less active than expected---one reason is that those who supply data are not rewarded for the value it creates downstream. A defining feature of data is its \emph{replicability}: a buyer can refine purchased data into a new product and resell it to \emph{many} downstream buyers, so a single source seeds a branching cascade of resales that naturally forms a \emph{tree}. Because an upstream seller captures none of this downstream value, its incentive to trade is weakened. Existing works propose \emph{profit reallocation}---returning part of downstream revenue to upstream contributors---as a natural remedy. But whether profit reallocation works on the tree-structured markets that replicable data actually induces has remained open. To bridge this gap, we develop a principled framework for profit reallocation on tree-structured data markets. We introduce a sequential trading game on a tree and a general class of budget-feasible profit reallocation mechanisms (PRMs) over it. We derive efficient algorithms to compute the induced equilibria---a polynomial-time exact algorithm for discrete valuations and a fully polynomial-time approximation scheme (FPTAS) for continuous ones---via a subtree decomposition technique that tames the potential coupling across a seller's children. We then prove that, under mild assumptions, \emph{any} budget-feasible PRM weakly expands the trades that occur in equilibrium, and any budget-balanced PRM additionally weakly improves social welfare, relative to the baseline that reallocates nothing. Experiments on synthetic markets confirm that these benefits are substantial, persist even when the assumptions fail, and grow with the depth and branching of the tree.

cs.GT↗

A Disk-Shaped Magnetoelastic Torque Sensor for Robotic Joints Using Permanent Magnetization

Direct torque sensing is a growing need in the robotics community to enable precise control and interactions where torque estimation from motor current is not sufficient. This paper presents a novel disk-shaped magnetoelastic torque sensor with a compact axial envelope of about 1 cm, suitable for integration in robotic joints. A four-magnetometer architecture is used to measure the field modulated by the stress affecting a narrow magnetized region while rejecting the effects of parasitic cross forces. A custom-designed magnetic shield enhances the torque sensitivity while reducing the external stray fields by a factor 7x. The device measures the torque with an accuracy of 1.34 %FS relative to the 50 Nm full scale (FS). The paper details the development of the sensor through the mechanical design, the magnetization procedure, and the experimental validation. The results demonstrate the potential of the proposed sensor for robotic applications.

cs.RO↗

PGL-3D: Towards Progressive Geometric Learning for 3D Visual Query Localization

3D Visual Query Localization (3DVQL) retrieves the latest contiguous occurrence of a queried object in an RGB--point-cloud sequence and predicts a 9-DoF cuboid for every response frame. The query is captured independently of the search sequence, so its annotated pose may differ from how the object appears in the search frames. The benchmark baseline predicts cuboids after feature modeling, leaving their geometry unused for subsequent feature refinement. We investigate whether complete intermediate cuboids can improve query and proposal representations before final decoding. We introduce Progressive Geometric Learning for 3DVQL (PGL-3D), a predict--select--refine--re-predict framework that uses intermediate cuboids to guide the aggregation of search evidence and update query and proposal representations. A shared head first predicts a complete cuboid for every proposal. Query--Tube--Memory (QTM) then selects reference observations by combining proposal association, cuboid quality, frame response, and target absence, since association confidence alone establishes neither target presence nor geometric accuracy. The center, size, and orientation of each selected cuboid define soft pooling weights over query-conditioned proposal features. The pooled memory updates the query and proposal representations, and the head re-predicts from the updated features. A training-only objective, ST-D9O, supervises cuboid geometry at every stage by adding boundary, signed-distance, and soft-overlap terms to parameter regression. PGL-3D achieves a mean stAP of $0.270 \pm 0.004$ on 3DVQL, compared with $0.044$ reported for LaF. Ablations support the benefits of geometry-guided feature updates, while stage-wise analyses show improved cuboid accuracy. Replacing the geometry objective in our PROT3D reproduction with ST-D9O improves mAO on GSOT3D from $21.63\%$ to $25.78\%$. Our code and models will be released.

cs.CV↗

Control Data Scheduling over Shared Communication Channels: A Sparse and Collision-Free Mechanism

In systems where controllers operate remotely and communicate with actuators over shared communication channels, it is crucial to efficiently schedule the control data transmission to reduce bandwidth usage and actuator effort. In this article, we investigate a novel scheduling mechanism that jointly coordinates control actions across different controllers and different time steps. We propose an algorithm based on the alternating direction method of multipliers (ADMM) to solve the resulting optimization problem. While ADMM is often treated as a black-box solver, the proposed algorithm offers a clear physical interpretation, ensures convergence to a stationary point, and is computationally efficient. Simulation results validate our theoretical results and demonstrate the effectiveness of our proposed algorithm.

eess.SY↗

SymbolicLM: Training Language Models as Symbolic Regressors

Large Language Models (LLMs) have shown promising capabilities in scientific reasoning, yet scientific discovery ultimately requires deriving precise laws directly from observational data, known as Symbolic Regression (SR). This poses a challenge for LLMs due to the gap between probabilistic text generation and the exact structural requirements of SR. Existing approaches rely on complex external scaffolds, which are computationally expensive and separate symbolic reasoning from the model itself. To address this limitation, we propose to directly equip LLMs with symbolic regression capabilities through dedicated numerical-symbolic and physical supervision. We introduce PhysSymbArena, a large-scale benchmark containing over 160,000 equations and 1.8B tokens of numerical-symbolic data with physical descriptions, enabling systematic training and evaluation. Based on PhysSymbArena, we develop SymbolicLM, which enhances the symbolic regression ability of LLMs through mathematical and physical supervision. During inference, we further introduce SymbolicSGA, a refinement framework that leverages quantitative feedback to iteratively improve generated equations. Experiments on multiple symbolic regression benchmarks show that SymbolicLM substantially improves structural recovery while maintaining competitive numerical fitting performance. These results demonstrate that symbolic regression can be explicitly learned as an intrinsic capability of LLMs.

cs.CE↗

Boolean Cumulants and Exact Reduced Descriptions of Renewal-Driven Systems

Reduced descriptions of unresolved fluctuations are commonly based on Gaussian processes, although many realistic forcings have finite correlation times and a renewal structure. For memoryless step (Kubo--Anderson) noise, we show that the reduced dynamics admit two exact and complementary descriptions: a kernel representation governed by the Boolean cumulants of the jump distribution, and a frozen-noise representation adapted to stationary probability densities. For exponentially distributed waiting times, the totally time-ordered $G$-cumulants coincide with the Boolean cumulants of the jump law, and the memory kernel is resummed exactly as the Boolean generator $η$ evaluated on a resolvent operator. Second-order closure is exact if and only if the jump law is symmetric Bernoulli; otherwise, the leading closure error is controlled by $(b_4/b_2)λ^2$. The Boolean hierarchy, however, acts on the kernel, not on the stationary measure. Stationary densities are approximated by replacing the jump law with its $N$-point Gauss quadrature: each surrogate preserves all multi-time correlations up to order $2N-1$ for any waiting-time law, its kernel is a Padé resummation of the Boolean one, and it is itself an exactly solvable renewal problem. For the linear system with linear multiplicative interaction (LIMI/CAM), the frozen-noise representation reduces the dynamics to a random affine recursion and yields exact support boundaries, singularity exponents and Kesten tail indices, confirmed numerically. Rates, moments and closure errors are governed by the Boolean hierarchy, whereas stationary densities are determined by the geometry of the jump distribution. For memoryless renewal noise, the combinatorial structure controlling finite-correlation reductions is Boolean.

cond-mat.stat-mech↗

Climbing the Hill: Prompt Injection Red-Teaming Against Frontier Models with Curriculum Reinforcement Learning

Prompt injection is a leading security risk for LLMs and LLM-based applications such as agents. State-of-the-art red-teaming methods for prompt injection leverage reinforcement learning (RL) to train an attacker LLM to generate effective injected prompts. However, when targeting frontier LLMs such as GPT-6-Luna, a major challenge is the cold-start problem: every attack attempt by the attacker LLM fails and thus receives zero reward, providing no signal for learning. In this work, we propose a curriculum learning-based method to address the cold-start problem. In particular, we propose to train the attacker LLM against a sequence of increasingly robust target LLMs, with each stage warm-starting from the attacker LLM obtained in the previous one. However, simply training against a weak target (e.g., GPT-4o-mini) may not sufficiently prepare the attacker LLM to obtain useful learning signals against a frontier LLM (e.g., GPT-5.6-Terra). Instead, we find that the design of the curriculum is critical: after each stage, the attacker LLM needs to partially succeed against the next target LLM such that it can learn from successful attempts to attack the new target. Our extensive evaluation shows that our method can effectively red-team frontier LLMs, achieving an attack success rate (ASR@10) of 93.8\% and 45.0\% against GPT-5.6-Luna and GPT-5.6-Terra on AgentDyn, whereas state-of-the-art RL methods such as RL-Hammer and PISmith achieve 0\% ASR under the same setting. Moreover, we find that the attacker LLM transfers across targets, e.g., an attacker LLM trained to defeat one strong LLM (GPT-5.6-Terra) also succeeds against six other frontier LLMs (e.g., GPT-6-Luna) it was never trained on. Our code is available at https://github.com/albert-y1n/PIForge.

cs.LG↗

Quantifying Behavioral Tails in Black-Box Language Models

We introduce RareTrap, a framework for estimating the probability of severe behaviors in black box large language models (LLMs). A key challenge for probability estimation is defining a tractable distribution over the input space. To accomplish that, RareTrap uses a surrogate LLM and constructs a geometry-aware mapping from a lower-dimensional latent reference space into its token-embedding space to induce an explicit and reproducible distribution over input prompts. A response-level performance function is utilized on the response to quantify behavior severity. This enables sequential rare event simulation that concentrates evaluations on progressively more severe behaviors while preserving probability under the induced prompt distribution, which would otherwise be prohibitive to measure. Across 10 open-weight and two frontier models (GPT-5.4 and Claude Sonnet 4.6), we find that RareTrap successfully induces severe resource consumption behaviors and computes their probability with as few as 200 evaluations. RareTrap provides model developers a principled approach for evaluating language models under a common distribution, and prioritizing alignment effort to improve safety and mitigate risks.

cs.LG↗

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Connectionist temporal classification (CTC) naturally supports offline and streaming speech recognition with utterance-level supervision, but conventional implementations materialize frame-by-vocabulary activations in memory, making CTC training with native LLM vocabularies prohibitively memory-intensive. A key observation is that every valid CTC alignment uses only target tokens and blank, and their union across a batch typically forms a small subset of the full vocabulary. We introduce Pruned CTC, which restricts alignment computation to this subset while retaining full-vocabulary normalization. We prove that this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and gradients. Head-and-loss activation memory no longer scales linearly with vocabulary size. We further apply finite-beam alignment pruning. Building on Pruned CTC, we develop LLM-CTC, which adapts pretrained LLMs for non-autoregressive ASR while retaining causal attention and native vocabularies, and extend it to bounded-history streaming, avoiding chunk-level speech--text alignments. Experiments show that, with Zipformer-M encoder and 180K vocabulary, Pruned CTC reduces full-step memory by 5.1$\times$ with only 17% step-time overhead. Across three corpora, it matches standard CTC accuracy. On GigaSpeech, across six Qwen3 model sizes from 0.6B to 32B, LLM-CTC remains within 7% relative WER of LLM-CE with 7 to 10$\times$ faster recognition; when fine-tuning Qwen3-ASR for bounded-history streaming, LLM-CTC remains within 3% relative WER of matched offline models on the test set. Together, these results establish Pruned CTC as a scalable sequence objective for native-vocabulary LLM ASR across offline and streaming settings.

eess.AS↗

Who Blocks Whom? Probabilistic Pass-Blocking Assignments for Evaluating Blockers and Pass Rushers in American Football

Historically, statistical analysis of offensive lineman has been hindered by the lack of easily measurable quantities. More recently, with the introduction of player tracking data new methodological advances are now possible. Using high-dimensional spatio-temporal data, we adapt the defensive-matchup hidden Markov model of \cite{franks2015characterizing} from basketball to football pass protection, producing frame-by-frame probabilistic assignments of each pass blocker to the rushers. We show how this probabilistic assignment is a usable modeling artifact that augments existing player-evaluation frameworks. We directly quantify the attention a rusher commands, upgrade adjusted plus-minus \citep{Macdonald+2012} from all-or-nothing stints to partial, continuous blocking credit in continuous time, yield block-shedding survival metrics, and measure the space a rusher generates for his teammates. Fit to the first eight weeks of the 2021 NFL season, the resulting metrics recover widely-recognized elite rushers and pass protectors and align with independent charting.

stat.AP↗

Auditing Agent Actions through Query-Conditioned Attribution

LLM agents increasingly take consequential actions through interactions with users, policies, and external tools. Auditing these agents requires automated attribution of realized actions to their historical basis. However, existing attribution formulations do not provide question-specific traces for diverse auditing objectives. Additionally, when access to the acting model is limited (e.g., in API-only deployments), applicable methods commonly rely on costly input perturbations or external LLM analysis of complete trajectories. We therefore formulate query-conditioned agent action attribution, a new task that takes a natural-language auditing query as input and recovers the source and ordered intermediate evidence for the query-specified aspect of an action. We instantiate this task with $A^3Bench$, a benchmark comprising 1,396 auditing queries across policy basis, parameter provenance, failure propagation, and unsafe-behavior tracing. To enable efficient, query-specific attribution, we use small open-weight models as attribution proposers that combine query-conditioned gradient saliency with query-semantic relevance to rank history units. Our proposer consistently achieves stronger source and evidence rankings at lower inference cost than open-weight baselines, improving source MRR by up to 40.9\% and evidence MAP by 42.1\% with only two forward passes and one backward pass. Controlled evaluations confirm that our proposer improves attribution specificity by adapting its rankings to fine-grained changes in the auditing query. Building on a proposer ensemble, our end-to-end system surpasses the strongest frontier-model baseline in source accuracy (64.5\% vs.\ 60.4\%) while reducing empirical deployment latency by 29.9\% relative to the fastest frontier API baseline. Code and data will be released after the initial review period following final validation and cleanup.

cs.AI↗

Implicit-Explicit time integration scheme with Physics-based preconditioning for two-fluid tokamak boundary simulations

In this work, a globally stiffly accurate Implicit-Explicit (IMEX) Runge-Kutta scheme is developed and implemented in the GBS code [Ricci et al., Plasma Phys. Control. Fusion, 2012], for two-fluid plasma turbulence simulations. The stiffest phenomena, governed by shear Alfvén waves and parallel diffusion, are treated implicitly, while the remaining non-stiff terms are advanced explicitly. This splitting enables time steps well beyond the Courant-Friedrichs-Lewy limit, without incurring the full computational cost of a globally implicit formulation. To efficiently solve the implicit subsystem at each time step, a three-dimensional physics-based preconditioner, inspired by techniques developed in the magnetohydrodynamic (MHD) context, is used. The resulting framework is verified through the method of manufactured solutions and exhibits both algorithmic and parallel scalability. Significant advantages in numerical stability and computational efficiency are demonstrated with respect to an adaptive explicit Runge-Kutta scheme.

physics.plasm-ph↗

Flip-packability: uniform characterisations of tame graph classes

A class of graphs is monadically dependent if one cannot encode all graphs in coloured graphs from the class using a fixed first-order formula, and monadically stable if one cannot even encode arbitrarily long linear orders. Bonnet et al. (ICALP 2025) characterised monadic dependence by flip-separability: for every vertex weighting, boundedly many flips - complementations of the adjacency relation within a vertex subset - make every ball of radius $r$ carry at most an $\varepsilon$-fraction of the weight, so that every set carrying an $\varepsilon$-fraction has two elements pulled apart. We introduce flip-packability: boundedly many graphs, each obtained from the input by boundedly many flips and all determined by the weighting before any set is presented, such that every set carrying an $\varepsilon$-fraction of the weight has $m$ elements pairwise far apart in one of them. The number of flips producing each graph depends on the radius alone; only the number of graphs depends on $\varepsilon$ and $m$. We prove that a class of graphs is flip-packable if and only if it is monadically stable, and $2$-flip-packable, that is, flip-packable with $m=2$, if and only if it is monadically dependent. The passage from two scattered elements to $m$ is thus exactly what separates the two notions. For monadically stable classes we show that the flipped graphs can be computed from the weighting in cubic time. Varying the three parameters of the definition - the sparsifying operation, the radius, and the number $m$ of elements scattered - produces eight known characterisations of sparse and dense graph classes from the same template. In each case $m$ separates a depth-like notion from its width-like relaxation: treedepth from treewidth, shrubdepth from cliquewidth, and monadic stability from monadic dependence.

cs.LO↗

Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents

Executable environments are critical for post-training agents on tasks that require tool use and multi-step interaction, but constructing executable tasks together with their environments remains difficult to scale. Skills provide reusable domain knowledge, operational procedures, and tool-use instructions, but a substantial gap remains between the information contained in a skill and a concrete, challenging task with a complete executable environment. To address this gap, we introduce Skill2Env, a capability-oriented framework that starts from a skill and uses agent capability demands to guide task and environment synthesis. Skill2Env represents these demands through reusable difficulty patterns and instantiates them into task blueprints that specify objectives, challenges, environment facts, information boundaries, and acceptance criteria. These blueprints guide the joint construction of task instructions, execution substrates, workspaces, and rubric-based evaluators around source skills. We further propose Iterative Task Hardening, which uses solver execution evidence to identify insufficiently challenging task designs, strengthen or extend their difficulty-pattern instantiations, and revise the corresponding blueprints and environments. Using 1.5K high-scoring trajectories generated from Skill2Env environments for supervised fine-tuning, we observe consistent improvements across a broad range of agent benchmarks, demonstrating the effectiveness of capability-oriented environment synthesis for agent post-training.

cs.AI↗

Solving Vertex Integrity Faster than $2^n$

Vertex Integrity is a classical graph parameter measuring the vulnerability of a network to vertex failures. We settle two fundamental questions concerning its exact exponential complexity. First, we prove that, unless the Exponential Time Hypothesis fails, Vertex Integrity has no $2^{o(n)}n^{O(1)}$-time algorithm. Secondly, we give a deterministic exact algorithm running in $O(1.9602^n)$ time and space, improving the straightforward $O(2^n)$ bound.

cs.DS↗

Diffusion Reward Models

Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate or to a distribution from a fixed parametric family. This is at odds with human preference, which is inherently multimodal: the same response can be reasonably judged in many ways, and no single family covers all of them. To better fit this structure, we introduce DRM, a Diffusion Reward Model that recasts reward modeling as conditional density estimation over $p(\mathbf{r}\mid x,y)$. Conditioned on a frozen LLM encoder, a lightweight Diffusion Transformer denoises Gaussian noise into a reward vector, placing no parametric assumption on the output distribution and naturally representing its multimodal structure. A single architecture handles both multi-attribute regression and pairwise preference data, and at inference $N$ samples form an empirical reward distribution that can be aggregated into a scalar, a variance, or quantiles. Across five benchmarks, DRM matches or surpasses baselines under matched data and backbone, stays competitive with much larger discriminative, distributional, and generative RMs despite its modest training scale, and recovers multimodal reward structure where conventional heads collapse to a point. Uncertainty-aware rejection and lower-confidence-bound (LCB) aggregation further demonstrate that DRM can exploit distributional information beyond a scalar reward to improve reward-model decisions. Downstream RLHF experiments additionally show that using DRM as the training-time reward leads to improved policy performance, directly validating the practical benefit of diffusion-based reward modeling for RLHF training.

cs.LG↗

JET: Justification Evaluation in Transformer

JET uses pretrained language and vision-language models to select among a finite set of answers without additional training. It evaluates candidate likelihoods directly and shares computation across candidates. Experiments on desktop CPUs and consumer GPUs assess decision accuracy and execution cost. Qwen3.6-35B-A3B achieves 87.48% accuracy on the full MMLU test set and 3.69 requests per second on a separately timed MMLU subset. The accuracy-throughput comparison covers model, hardware, and reasoning choices, with Jev as an external reference. Controlled execution experiments show 2.18-2.23-fold speedups from prefix reuse and cache management, and a 30.8% reduction in process time from input preparation optimizations, with unchanged outputs. Optional reasoning has a task-dependent accuracy-throughput trade-off. These results support local decision inference from existing models.

cs.LG↗

Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents

Outcome-only reinforcement learning gives software engineering (SWE) agents a terminal success signal but little direct guidance about intermediate decisions. We introduce Counterfactual Rollout Replay (CRR), a training-time procedure that uses forkable executable environments to obtain step-level return contrasts. CRR selects a small set of decision points, restores each state, samples an alternative action, and rolls the branch forward under the policy. It retains the realised training trajectory and replaces the advantage at selected steps with the difference between its terminal return and the sampled counterfactual return. The method needs no human process labels or learned process reward model; free refers to those supervision costs, not replay compute. With a 14B policy, CRR improves pass@1 on SWE-bench Verified, SWE-bench Live, and SWE-rebench, and combines with process-reward and trajectory-search methods. On SWE-bench Verified, an equal-wall-clock comparison on the same hardware yields 41.7% versus 36.7% for extended outcome-only GRPO, a 5.0-point gain with fork overhead included. These results apply to environments with affordable, reliable state restoration; stochastic continuations and expensive or imperfect replay remain limitations.

cs.SE↗