Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

pytest-gpu-proof: Enabling Cloud-CPU Continuous Integration for GPU Code with Local GPU Attestation

GPU acceleration is now routine across robotics, but cloud-hosted GPU continuous integration (CI) runners are expensive, resulting in severe under-testing of GPU-accelerated code. We present pytest-gpu-proof, an open-source pytest plugin offering a practical middle ground. Tests can be run on a local machine, signed with a receipt of exactly what ran and what it produced, and integrated into standard CPU CI workflows (e.g., GitHub Actions). The tool is open source and on PyPI, and we are actively integrating it across our lab's software stack.

cs.DC↗

Affine Pricing Models from Group Quantization and Holonomy

The analytic tractability of affine pricing models is usually expressed through two complementary formulations: a coordinate-space pricing operator and an exponential-affine transform representation governed by generalized Riccati equations. We develop \emph{Affine Holonomy Group Quantization} (AHGQ) as a geometric framework in which these two formulations arise from the same underlying structure. The construction separates the affine pricing symbol into a homogeneous quadratic sector and a complementary affine sector. The first generates a finite-dimensional symplectic transport and a centrally extended Lie group, while the second is represented by a multiplicative holonomy carried by a thin-path groupoid. Their combination determines an affine Poincaré--Cartan form. Its characteristic dynamics reduce in momentum variables to the generalized Riccati system and its scalar amplitude, whereas the coordinate representation recovers the standard affine pricing operator. Representative Gaussian and square-root models illustrate the construction. The contribution is structural: AHGQ gives a common geometric origin to the coordinate and transform representations of continuous-path, time-homogeneous affine pricing models.

q-fin.MF↗

Magnetostrictive properties in Fe$_{4-x}$Co$_{x}$N films: Insight from experiments and first-principles calculations

The ferromagnetic nitride Fe$_{4}$N has attracted attention for application in spintronic devices due to its high spin polarization and large magnetostriction. We present a combined experimental and theoretical study on the magnetostriction of Fe$_{4-x}$Co$_{x}$N films across a wide composition range. The Fe$_{4-x}$Co$_{x}$N films were grown on SrTiO$_{3}$(001) substrates using molecular beam epitaxy, and the magnetostriction constants along the [100] direction ($λ$$_{100}$) and [111] direction ($λ$$_{111}$) were precisely evaluated using an optical cantilever method. The experimental results reveal that $λ$$_{111}$ remains positive in the whole composition range and shows a maximum value of +82 ppm around x = 0.9. The variation in $λ$$_{111}$ with x is much smaller than that in $λ$$_{100}$, for which giant tunability and sign reversal are observed. First-principles calculations show reasonable agreement with the experimental $λ$$_{111}$ for x ${\ge}$ 1.6, but give negative values at lower x, and an exceptionally large negative $λ$$_{111}$ is obtained at x = 0.8, where the Fermi level coincides with a pronounced minority-spin peak in the density of states. The calculated $λ$$_{111}$ depends strongly on the smearing parameter, indicating that the rhombohedral magnetostriction is highly sensitive to the treatment of atomic disorder. The saturation magnetostriction constant ($λ$$_{s}$) derived from $λ$$_{100}$ and $λ$$_{111}$ is also compared with the $λ$$_{s}$ measured for the (001)-oriented polycrystalline Fe$_{4-x}$Co$_{x}$N films, and a possible scenario for deviation between them is discussed. Our findings clarify the basic features of magnetostriction in the Fe$_{4-x}$Co$_{x}$N system, providing essential magnetoelastic parameters for designing nitride-based spintronic devices.

cond-mat.mtrl-sci↗

Direction-Scale Decomposition in Action Representation: Rethinking What to Tokenize for Vision-Language-Action Models

Action representation plays a central role in discrete-token vision-language-action (VLA) learning but remains underexamined. Under conventional pose-increment representations, action tokens are sensitive to execution speed and dataset-specific normalization, potentially obscuring geometric structure shared across demonstrations and datasets. We introduce Direction-Scale Decomposition (DSD), an action representation that decomposes translation and rotation increments into direction and scale components before tokenization. DSD isolates motion direction while retaining magnitudes in separate scale channels. We evaluate DSD with uniform binning (BIN) and BEAST, a B-spline-based tokenizer, in simulation and real-world manipulation under both single-dataset and mixed-dataset training. On LIBERO, DSD improves average success rates with both tokenizers. On SimplerEnv, DSD-BIN outperforms BIN by 10.3 percentage points in overall success rate under mixed-dataset training. Real-robot experiments further show gains both with and without robotics pretraining. These results support DSD as an effective action representation for discrete-token VLA models and suggest its potential to mitigate performance degradation when training on large and diverse dataset mixtures. Our project page with additional resources is available at https://vla-dsd.github.io/

cs.CV↗

Refutable Exclusion Restrictions in Competing Risks with Categorical Covariates

Competing-risks data do not identify latent marginal duration distributions or their dependence without additional restrictions. This paper asks whether restrictions introduced to restore identification themselves restrict the observable law. For a two-risk Archimedean model with categorical exclusion restrictions, we derive a necessary-and-sufficient observable characterization. A discrete single-crossing argument identifies the scalar copula parameter from cell-specific overall survival probabilities, while cause indicators recover the remaining allocation and generate additional specification restrictions. We construct an identification-robust, self-normalized quadratic statistic, invert it to obtain confidence sets, and use empty inverted sets as a conservative specification test. Simulations show that the cause indicator can be decisive: under the weakest contrast considered it converts a frequently uninformative survival-based confidence set into an informative joint set without loss of coverage, whereas excessive categorical contrast can eliminate causes from individual cells and make cause-specific recovery inadmissible. The results turn latent exclusion restrictions into refutable restrictions without estimating covariate derivatives.

stat.ME↗

Bondal--Orlov reconstruction for tame stacks with trivial generic stabilizer

We generalize the Bondal--Orlov Reconstruction Theorem to smooth, proper, tame algebraic stacks with generically trivial stabilizer, whose canonical bundles are nowhere torsion: They do not become trivial after pullback along any finite morphism from a projective curve. Our main tools are coherent Tannaka duality as proved by Lurie and by Hall and Rydh, and Grothendieck duality for proper tame stacks as developed by Hall and Priver. We additionally use the notion of nowhere torsion line bundle on a scheme to prove a version of the Bondal--Orlov Reconstruction Theorem for finite type, separated, Gorenstein schemes which are not necessarily proper, which generalizes work of Ballard, Favero, Ito, and Matsui.

math.AG↗

Image Fidelity is Not Field Fidelity: Joint Thermodynamic Reconstruction and Error Localization in Neural Tomography

Neural fields for scientific tomography are optimized from 2D images, but the actual quantity of interest is often a latent 3D physical field. Because the forward map is many-to-one, low 2D image error need not certify a correct 3D field. Moreover, the latent field is not directly supervised during training, and its error cannot be evaluated against truth at deployment. We develop CoroNeRF to jointly optimize 3D electron density and temperature fields directly from multiview, multiline intensities through a differentiable atomic-emission renderer. Using solar coronal tomography as a controlled testbed, we evaluate physical-field recovery and test whether cross-seed instability provides a ground-truth-free-at-inference indicator of local physical-field error. We underscore the following two observations. (i) Image fidelity is not field fidelity: spectral ablations show that limited-channel reconstructions can fit their available observations well while recovering substantially worse fields, whereas evaluation on a common richer probe exposes the discrepancy. (ii) Cross-seed instability ranks local physical-field error across tested matched-model conditions, supported by sparsification and physical signal-strength controls. Seed-deviation projections provide complementary directional validation, but shared forward-model mismatch can still produce incorrect cross-seed consensus. These results characterize joint thermodynamic recovery and the usefulness and limits of seed-based error localization in a controlled, single-scene solar tomography testbed.

cs.LG↗

Computations of Cohomology of Arithmetic Groups, Part 1

The key part of the current paper is the computation of boundary and Eisenstein cohomology of $GL_4({\mathbb Z})$ with coefficient in any highest weight representations. The method we develop let us compute in an alternative way the cohomology of $SL_3({\mathbb Z})$ and of $GL_3({\mathbb Z})$ with coefficients in any highest weight representation. This is done in a simpler, faster and in a more structured way compared to \cite{BHHM}. We state a duality for the boundary cohomology of $GL_m({\mathbb Z})$ of the type of Serre's duality, where the dualizing sheaf is a power of the determinant representation. We refine this duality to a duality on the level of the spectral sequence for the boundary cohomology $E_\infty^{p,q}$. We state it as a conjecture. However, all the computations, 30 different families of representations, satisfy this conjecture. We compute the Eisenstein cohomology of $GL_4({\mathbb Z})$ with coefficients in the symmetric powers and their twist by the determinant representation. For several other representations, we compute the Eisenstein cohomology, based a few conjectures. Based on those conjectured, one can compute the Eisenstein cohomology in most of the cases. They will be included in the next version of the paper.

math.NT↗

When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse

Long-running LLM applications repeatedly send growing context, making prefix caching critical for reducing prefill cost. Yet prefix-cache behavior under agentic workloads remains poorly understood. We study production traces from two companies and evaluate 14 eviction algorithms across HBM-constrained and large memory-pool settings. Despite a large gap to Belady, sophisticated policies designed for traditional caches provide little benefit over LRU. The reason is structural: prefix reuse is dominated by the regular pacing of active sessions, making recency unusually predictive. Prefix caching nevertheless introduces new challenges, including heavy-tailed session footprints and highly variable miss costs as attention computation grows with sequence length. We introduce the compute-savings ratio and two offline oracles to quantify these effects. Our results show that effective prefix-cache management should retain recency as its foundation while selectively adding quick demotion for one-hit prefixes, compute-aware partial eviction for expensive misses, and capacity-dependent eviction granularity. We will release the traces and simulator to support future research.

cs.DC↗

Panel Conditioning in Fixed-Effects Models: Identification and Bias Propagation

Panel conditioning, the causal effect of prior survey participation on responses, can vary with tenure. Under an additive model of cell means in period, entry cohort, and tenure, we characterize which features of the conditioning path a staggered panel identifies on its observed support, and how the unidentified component affects common panel estimators. The identified set of the path is an affine translate of the tenure projection of the cell design's kernel, and a linear functional of the path is identified exactly when it annihilates that projection. It always contains an affine direction and, when the entry cohorts share a stride, periodic directions, which exhaust it under a connectivity condition on observed increments; second differences at that stride are then identified, and ordinary ones generally are not when the stride exceeds one. Under a recruitment condition, an interrupted schedule such as the four-eight-four rotation of the Current Population Survey (CPS) distinguishes a constant increment per interview from one per calendar month, which no equally spaced schedule can. We give support conditions for recovery under a plateau, entry-wave negative controls, or bounded cohort drift. A second set of results links identification to regression: two-way fixed effects absorb every unidentified direction, so the remaining conditioning bias is normalization-invariant and itself identified, and a two-way regression with tenure indicators corrects it under a residual-rank condition. When event time is aligned with tenure, conditioning shifts event-study coefficients by a known linear functional of the path, producing pre-trends without anticipation; bounds on identified curvature bound those shifts. Simulations verify the identities, a 19-wave Japanese panel illustrates the support calculations, and published CPS month-in-sample indices give a descriptive, not identifying, example.

stat.ME↗

Least-false Cox coefficients under affine follow-up contamination: exact continuous- and grouped-time benchmarks

Covariates summarized over a subject's completed follow-up are sometimes entered into Cox regression as though observed at baseline. This practice incorporates future event or censoring information and changes both the estimand and its sampling behavior. We analyze an affine class in which a genuine baseline covariate is contaminated by realized follow-up time. Treating partial likelihood as an observed-data estimation criterion, we derive the population score and characterize its unique least-false coefficient. Exact continuous-time benchmarks show scale reduction and saturation under strong contamination, while administrative censoring destroys the reduction and may produce overshoot. Grouping exit times with Breslow ties changes the geometry: the coefficient has a single hump and eventually returns to zero even though the induced association diverges. We also derive an observed-data influence function and show why model-based variance can be either too small or too large. A subject-level sandwich consistently estimates uncertainty around the least-false target under the stated conditions, but it does not correct the target itself.

stat.ME↗

Improved Revenue Guarantees for Selling Separately and Bundling

We study how much revenue a seller can lose by restricting attention to selling separately or grand bundling, in the setting of a single additive buyer with independent item values. Although revenue-optimal mechanisms can require lotteries and infinite menus, Babaioff, Immorlica, Lucier, and Weinberg showed that the better of these two simple formats always achieves a constant fraction of optimal revenue. We prove that $\mathrm{OPT} \le 3.52 \max\{\mathrm{SREV}, \mathrm{BREV}\}$, where $\mathrm{SREV}$ and $\mathrm{BREV}$ are the optimal revenues from selling separately and grand bundling, respectively. This improves the previous best-known approximation factor of $5.2$ due to Ma and Simchi-Levi and narrows the gap to the known lower bound of $2$.

cs.GT↗

A Task-Based Framework for Evaluating Raman Spectral Quality Measures

Raman spectral preprocessing and enhancement are often evaluated by comparing output spectra with a reference. Interpreting these comparisons requires evidence that spectral quality measures reflect downstream task performance. We present a controlled-perturbation framework for testing this relationship. Five perturbation types (baseline distortion, independent noise, correlated noise, a global wavenumber shift, and nonlinear axis warping) generate paired changes in a spectral measure (metric harm) and in downstream performance (task harm). An alignment gap (AG) quantifies how much the relationship between metric harm and task harm changes with perturbation type. Ordering concordance (OC) measures how often a metric correctly ranks two conditions by their task harm. The framework evaluates thirteen outputs (MSE, RMSE, MAE, NMSE, spectral angle, Pearson correlation, Wasserstein distance, a structure-to-noise ratio, peak precision, recall, F1, artifact ratio, and missing ratio). Three public datasets provide bacterial classification, sugar-mixture quantification, and mineral identification tasks. PCA with logistic regression, partial least squares regression, and cosine library matching supply the task outcomes. Classifiers and calibrations are fitted either to unperturbed training spectra or to each perturbed training condition, then evaluated on the same perturbed test spectra. Mineral queries are compared with an unchanged or correspondingly perturbed library. The resulting comparisons identify task-specific strengths and limitations, including cases where better ordering does not accompany a smaller AG. Removing axis perturbations and comparing spectra on a common physical grid test how these findings depend on the evaluation design. The framework provides a reproducible procedure for assessing existing measures and testing new candidates against downstream task performance.

physics.chem-ph↗

Concentration of bounded sparse chaoses and sparse Khatri-Rao embeddings

We establish moment and mixed-tail inequalities for fixed-order decoupled homogeneous chaoses generated by independent, centered, sparse bounded random variables. Our bounds apply to arbitrary real rectangular coefficient tensors and describe the fluctuation scales through weighted slice and partition norms, with a Bennett-type logarithmic improvement in the largest-entry term. As an application, we derive guarantees for sparse Khatri--Rao embeddings that explicitly account for sparsity and input geometry.

math.PR↗

Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents

We introduce Forecast-Dojo, a replayable environment for benchmarking and training LLM forecasting agents. It combines resolved prediction-market questions with dated news, allowing agents to research an event and revisit their predictions at successive historical dates. The same tasks and tools support repeated evaluation, collection of training interactions, and feedback from recorded outcomes without waiting for new events to resolve. Forecast-Dojo contains 1,568 Polymarket events, split by time into training and evaluation periods, and 18.8M dated news articles. In an evaluation of 12 models, research tools lower Brier score for all 12. Forecasts also improve as events unfold, with the largest gains at steps where more newly dated evidence is recorded. Every model still trails historical market forecasts in both Brier score and accuracy. A belief notebook carried between dates lowers research cost but does not consistently improve forecast quality. Beyond evaluation, Forecast-Dojo provides interaction trajectories and outcome feedback for agent learning, with supervised fine-tuning as a proof of concept.

cs.AI↗

Learning New Words from Unlabeled Test Data in Automatic Speech Recognition

New words are invented every day. A human listener can learn a new word by hearing it clearly once and inferring its usage from sentence context. This paper proposes granting ASR a similar ability to learn the contextual representations and spellings of new words from unlabeled test data at test time. A frozen CTC acoustic model provides spellings, a frozen language model provides contextual evidence for out-of-vocabulary (OOV) word detection, and an adaptation module expands the vocabulary by learning the lexical token representations with distributions over CTC-generated candidates. The spelling model of each token is optimized by minimizing a Kullback-Leibler divergence (KLD) objective. We demonstrate that the CTC-weighted language model log likelihood ratio can be interpreted as the KLD between the unknown correct ASR and the unsupervised learned ASR, and that, using a Pinsker bound, the square root of KLD can be interpreted as an upper bound on the total variation distance between the true and estimated spelling of the unknown word. Experiments show relative OOV character-error-rate reductions of up to 14.97% on LibriSpeech and 6.67% on dysarthric Speech Accessibility Project data for recurring OOV words, relative to the corresponding rescoring system.

eess.AS↗

Online Sim-to-Real Adaptation via Closed-Loop System Modeling

Sim-to-real transfer has made substantial progress, but can still produce controllers that remain stable and functional on hardware while suffering from degraded tracking accuracy due to residual dynamics mismatch. Correcting these errors typically requires identifying the underlying system dynamics, adapting the control policy, or returning to simulation for additional training and finetuning, all of which can require substantial data and computation. We propose OSRAM (Online Sim-to-Real Adaptation via Closed-Loop System Modeling), a framework that instead adapts the reference commands provided to an existing controller. OSRAM treats the deployed robot and its policy as a unified closed-loop dynamical system and learns its task-level command-response behavior directly from tracking observations. A closed-loop dynamics model is meta-trained across randomized dynamics in simulation and rapidly finetuned after deployment using limited real-world interaction. The adapted model is then used to optimize future reference commands while leaving the underlying control policy unchanged. We evaluate OSRAM on bipedal velocity tracking and loco-manipulation in simulation and on hardware. Results show that closed-loop modeling improves prediction and tracking accuracy under unseen dynamics, while online reference adaptation reduces residual sim-to-real tracking errors across different control objectives and hardware configurations. These results demonstrate that adapting the behavior of the robot-policy closed loop provides a practical alternative to finetuning the policy or identifying the full physical dynamics for sim-to-real transfer. More information can be found at http://generalroboticslab.com/OSRAM.

cs.RO↗

Broadening Uncertainty Estimation for Audio Question Answering Across Methods, Formats, and Inputs

Audio-language models can produce confident answers unsupported by the audio, motivating uncertainty estimates that identify unreliable responses. We compare probability-based, sampling-based, self-verification, evidential, and contrastive measures across four open-weight models and five audio QA benchmarks. In multiple-choice evaluation, first-token measures are strongest overall, with top-1 probability achieving a mean AUROC of .740, compared with .708 for ten-sample discrete semantic entropy, while requiring no additional model calls. Across four benchmarks, shifting from multiple-choice to open-ended evaluation lowers mean accuracy from 57.6% to 36.6%, yet uncertainty remains predictive of errors: semantic entropy, maximum token entropy, and semantic agreement achieve mean AUROCs of .697, .694, and .693, respectively. To test whether uncertainty reflects the evidence available to answer the question, we perform input ablations that remove either the audio or the question. Across top-1 confidence, entropy, and sampling-based measures, removing audio reduces error-detection AUROC by .101 on average, compared with .010 when removing the question. Together, these results establish efficient uncertainty baselines and show that uncertainty in audio-language models depends substantially more on available audio evidence than on question text.

cs.SD↗