Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 973 records · Page 54Linked to original sources

Bouncing-ball modes in ergodic billiards

We show that for a horizontally stretched Bunimovich billiard, for almost every stretching parameter, there exists a sequence of eigenfunctions concentrating in the rectangular part. Thanks to the work of Markarian--Oliffson Kamphorst--Pinto de Carvalho, we know that for an open set of parameters such billiards are ergodic. The proof is a result of interaction with ChatGPT 6: we were curious if it could establish the existence of bouncing ball modes for Bunimovich or Sinai billiards. That was not successful until we suggested varying the wings. That produced an essentially complete argument for the existence of quasimodes. We then suggested using Hassell's parameter-variation argument which led to the result about eigenfunctions. The argument was then significantly simplified and clarified by the authors.

math.SP↗

Towards Breaking the Learning System Wall Using Multimodal Tutoring Transcriptions

Past research using log data has faced the "learning system wall," whereby few methods exist for generalizing models of student learning across platforms. Increasingly, online learning is captured by richer forms of data, including dialog and video, with new affordances. An example of this is remote tutoring programs, where human tutors support students who use learning systems while video conferencing. Toward better platform-general modeling of learning, we introduce an AI-driven multimodal transcription system that processes screen-recording videos into unified screenplay-style transcripts containing audio dialogue and annotated learning log actions. We describe a planned method for temporally aligning AI-generated multimodal transcripts with MATHia learning logs and for identifying and classifying student learning processes to align with MATHia logs. Lastly, we highlight challenges and potential solutions in capturing learning processes in one system, offering initial steps towards generalizing log data across diverse systems.

cs.HC↗

AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes

Adding or removing a sounding object requires coordinated changes to visual content and sound while preserving the surrounding scene. Yet paired supervision for localized non-speech audiovisual editing remains limited, as visual and acoustic edits must target the same object and isolate its sound from overlapping sources. To address this gap, we introduce \textit{AVIOBench}, a dataset comprising 37.9 hours of paired audiovisual examples spanning 1{,}878 target-object names. AVIOBench links the visual presence and acoustic contribution of each target object through a shared identity and visual mask. Our automated pipeline uses visual grounding and cross-modal consistency to select target-sound removal candidates, then jointly refines the audiovisual pairs to improve perceptual quality and cross-modal consistency. Building on this dataset, we propose \textit{AVIO}, which adapts a pretrained text-to audiovisual generation model through source-conditioned feature modulation to jointly learn object addition and removal. A reference-frame curriculum gradually reduces reference conditioning during training, enabling one model to perform instruction-only editing with optional visual guidance. Quantitative and qualitative evaluations demonstrate effective audiovisual object removal and addition, with optional reference guidance providing appearance and placement control for addition.

cs.AI↗

Efficient and Scalable Physics-Guided Fully Convolutional Spatiotemporal Learning for 3D Microstructure Evolution Prediction

Accurate prediction of three-dimensional (3D) microstructure evolution remains computationally demanding because high-fidelity phase-field simulations require repeated numerical integration over large volumetric domains and long temporal horizons. This study develops an efficient and scalable physics-guided fully convolutional spatiotemporal framework for direct multi-frame prediction of complete 3D microstructure sequences. The model combines shared 3D spatial encoding and decoding with a factorized latent translator that integrates temporal, local 3D spatial, and channel interactions. A discrete Cahn--Hilliard (CH) residual is incorporated during training to regularize the learned evolution toward the governing dynamics without altering the inference pathway. The framework is evaluated on high-resolution 3D spinodal-decomposition trajectories under nominal, long-horizon, and reduced-temporal-context forecasting. Under full temporal context, the model accurately reproduces volumetric evolution, with average 3D structural similarity remaining above 0.97 over the nominal prediction horizon. Physics guidance becomes increasingly beneficial as temporal information is reduced, improving predictive robustness and preservation of interface-level morphology. The framework also achieves more than a 30-fold wall-clock speedup relative to the reference spectral phase-field solver, while physics guidance introduces no additional inference cost. These results establish direct multi-frame, physics-guided fully convolutional learning as a high-throughput surrogate strategy for dense 3D phase-field dynamics and repeated microstructure forecasting.

cs.LG↗

BRIDGE: Bilevel Retrieval-Credit-Aware Agentic Reinforcement Learning

Agentic reinforcement learning (ARL) with verifiable rewards improves the ability of large language models (LLMs) to tackle knowledge-intensive tasks by learning to interleave search and reasoning. However, most existing ARL methods optimize only LLM-generated tokens and treat retrieved evidence as environment observations. This creates an information-credit gap: failures caused by missing or misleading evidence are attributed to the LLM policy rather than to the retriever, which motivates training the LLM and the retriever jointly. In this paper, we show that retrieval and LLM policy learning are order-sensitive: adapting the retriever before optimizing the policy yields a larger reward gain than the reverse order. To preserve this hierarchy while allowing both components to co-adapt, we formulate retrieval-augmented agentic RL as a bilevel optimization problem. To solve it efficiently, we introduce BRIDGE, a memory-efficient first-order bilevel method motivated by a loss-landscape analysis of the RL and retrieval objectives. Across seven open-domain QA benchmarks, BRIDGE achieves the highest average accuracy with both 3B and 7B backbones, improving the multi-hop average over the strongest baseline by 9.6 and 3.4 EM points, respectively. It also achieves the best averaged answer accuracy and reasoning quality across medical QA benchmarks.

cs.AI↗

Compressed Sensing with Quantized Tensor Trains (QTTs)

The storage and recovery of a vector of length \(2^d\) can become prohibitively expensive as \(d\) grows, a manifestation of the curse of dimensionality. For certain structured functions and multiscale quantities, the resulting discretized vectors admit compact quantized tensor train (QTT) representations after binary tensorization: a length-\(2^d\) vector is reshaped into a \(d\)-way binary tensor with low tensor-train (TT) rank. Motivated by this construction, we study compressed sensing of tensorized signals that combine bounded TT rank with sparse low-order interaction structure, using randomly sampled Walsh--Hadamard measurements. For this structured model class, we establish a uniform restricted isometry property and derive a measurement condition guaranteeing uniqueness of noiseless recovery. We also propose a structure-aware initializer for the alternating linear scheme (ALS), obtained by projecting the adjoint backprojection onto the low-order interaction space before applying TT-SVD. We prove a uniform initialization error bound whose measurement requirement, for fixed structural parameters, grows polynomially with the tensor order \(d\) rather than with the ambient dimension \(2^d\). Numerical experiments support the predicted RIP and initialization scaling and demonstrate more reliable ALS recovery.

math.NA↗

Geographic Disparities in Hospice Quality and Family Caregiver Experience: The Roles of Ownership, Social Vulnerability, and Workforce Capacity

Hospice quality should be interpreted in relation to both provider organization and the local conditions under which care is delivered. This study develops a provider-county performance assessment framework by linking national Centers for Medicare & Medicaid Services (CMS) hospice data and Consumer Assessment of Healthcare Providers and Systems (CAHPS) Hospice Survey outcomes with county measures of rurality, social vulnerability, health burden, and workforce and health-resource context. The adjusted analysis includes 2,928 providers in 1,078 counties and combines geographic mapping, blockwise regression, six secondary CAHPS outcomes, and eight sensitivity analyses. Adding county context increased adjusted R-squared from 0.077 to 0.202. After full adjustment, for-profit hospices had overall caregiver ratings 3.832 percentage points lower than nonprofit hospices. This negative association appeared across all six secondary CAHPS domains and remained significant in every sensitivity specification. Higher county social vulnerability was also associated with poorer caregiver experience, although its magnitude depended partly on the specification of community health burden. These findings show that county context materially improves hospice performance assessment but does not eliminate the ownership difference. The framework supports context-aware monitoring, peer comparison, and targeted quality improvement.

stat.AP↗

On Extensions of the Unanimous Vote Problem

The Unanimous Vote problem is to determine a fixed order in which to flip each of $n$ biased coins, where each coin can be flipped only once, such that the expected number of flips until seeing both a head and a tail (or flipping all coins) is minimized. Duman Keles et al. (arXiv:2510.16678 [cs.DS]) gave an $\mathcal{O}(n \log n)$-time algorithm for this problem. Extensions of the Unanimous Vote problem are a rich source of stochastic optimization problems. We focus on three: (1) a variant in which each coin can be flipped arbitrarily many times (a solution is thus an infinite sequence of coin choices), (2) a generalization with $d$-sided dice, that can each be rolled once, where dice must be rolled until two different outcomes are observed (or all dice have been rolled), and (3) a different generalization with $d$-sided dice, where dice must be rolled until all $d$ outcomes have been observed. For (1), we show that there is an optimal sequence which follows a simple greedy rule; the same rule only gives a 1-additive approximation for the original problem (arXiv:2510.16678 [cs.DS]). The rule also yields a correspondence between a particular optimal sequence and a related mechanical word, which we exploit to characterize the conditions under which this optimal sequence is periodic. We establish tight multiplicative and additive adaptivity gaps for this variant. For (2), we show that two different generalizations of the greedy rule from (arXiv:2510.16678 [cs.DS]) can be combined to obtain a PTAS. For (3), we give an $\mathcal{O}(\log d)$-approximation algorithm by reducing the problem to Submodular Ranking (arXiv:1007.2503 [cs.DS]); the same reduction technique can be used to yield approximation algorithms for other stochastic probing problems. Finally, we pose a number of related open questions.

cs.DS↗

From $3$-spin to Negative $r$-spin: Limits of Cohomological Field Theories and Tautological Relations

We prove that the negative $r$-spin tautological relations of Chidambaram--Garcia-Failde--Giacchetto are a consequence of Pixton's $3$-spin relations, for every $r\ge2$, and deduce that they hold in the tautological Chow ring of $\overline{\mathcal{M}}_{g,n}$. In particular, our result for $r=2$ implies that Kazarian--Norbury $K$-relations hold in Chow. The proof factors through two essential parts. The first part shows that the tautological relations arising from Chiodo's class imply the negative $r$-spin relations. This requires an intricate study of limiting behavior of the Givental--Teleman graph contributions. In the second part, we show that the Chiodo relations are a consequence of Pixton's $3$-spin relations as an application of a generalization of Janda's work on tautological relations from CohFTs.

math.AG↗

Expediting AC Contingency Analysis using a Basecase Machine Learning Model

Grid planning and operation under increasingly variable operating conditions require fast and accurate AC power flow (AC-PF) contingency analysis for numerous line outages. While ML models can accelerate these computations, existing approaches often require either training separate models for different contingencies or a single model using data from multiple topologies. Both approaches incur substantial offline data generation and training costs. To address this gap, this work proposes a fixed-point framework that reuses a single ML model trained exclusively on basecase topology data to predict post-contingency AC-PF states under any non-critical single-line outage. We analyze the proposed method's convergence under the DC model approximation and derive its convergence rate in terms of network parameters. As a side result, we show that power transfer distribution factors (PTDFs) for non-critical lines have magnitudes less than one. We further characterize the ML training region to accommodate PF specifications encountered during fixed-point iterations and derive bounds on ML prediction errors for post-contingency state estimates. Numerical tests on a 6,717-bus synthetic Texas system demonstrate convergence across all non-critical single-line outages, with prediction errors remaining close to those of basecase. The proposed framework offers a favorable tradeoff between the speed of DC solvers and Newton-Raphson accuracy.

eess.SY↗

Affine spin-pencil factorization of generalized driven anisotropic Rabi--Stark Hamiltonians in an inhomogeneous orthosymplectic superalgebra

We construct a supersymmetric factorization for generalized driven anisotropic Rabi--Stark Hamiltonians using an affine factor $\hat{A}=\hat{C}+\hat{a}\hat{D}$ in an inhomogeneous orthosymplectic superalgebra, with two real invertible spin matrices $\hat{C}$ and $\hat{D}$. The six real pencil parameters constrain nine Hamiltonian frequencies and one energy offset up to an overall frequency scale. The normal product $\hat{A}^{\dagger}\hat{A}$ has an exact zero-energy ground doublet and, with its antinormal partner $\hat{A}\hat{A}^{\dagger}$, gives a Witten index of two. For distinct pencil roots with finite baseline slopes, coherent gauges define a Bargmann hierarchy and endpoint conditions for exceptional finite-polynomial states. The spin grading organizes four limiting cases, comprising two driven-oscillator families with every baseline exact and the constrained anisotropic-Rabi and Rabi--Stark sheets, which meet at the degenerate-qubit isotropic Rabi model. An aligned interior example adds direct and longitudinal drives, Stark deformation, and intensity-dependent coupling while retaining the factorization.

quant-ph↗

SAIVE: Selecting AI Valuable Entities

Data lakes store large amounts of telemetry, with logs from network sensors, hosts, and applications containing possibly hundreds of fields for every event. Large enterprises are then left with data lakes that cannot be analyzed efficiently with AI. Aggregate analysis looks at persistent shifts in behavior over time. Many of the fields and columns in data lakes are not useful as they do not contain information that is sufficiently diverse or concentrated to support AI analysis. SAIVE is a simple method for examining a few rows in a large table and applies a histogram of histograms filtering criterion to select the fields that for AI analysis is more likely to yield useful results. This paper provides a principled foundation for the SAIVE heuristics by assuming of a Zipf-Mandelbrot power-law distribution of the underlying data. Constraining the Zipf-Mandelbrot exponent alpha to a reasonable range provides a a practical, cheap, expert-free filter for selecting AI valuable entities in large data sets.

cs.DB↗

Local Search for Fair Max-Min Diversification

Given $n$ points in a metric space, Max-Min diversification asks for a subset of $k$ points maximizing the minimum pairwise distance between the selected points. This is arguably the most fundamental notion of diversity with applications across a wide range of domains. We consider this problem under partition constraints, previously studied as Fair Max-Min Diversification (FMMD). Here, each point has a color in $[m]$, and a feasible solution must contain exactly $k_i$ points of color $i$, where $k_1,\ldots,k_m$ are prescribed parameters satisfying $\sum_i k_i=k$. We give the first constant factor approximation for the problem using local search, that runs in time $f(m)\cdot \operatorname{poly}(n)$, in which all constraints are satisfied exactly. All previously known algorithms either provided an $\widetilde Θ(m)$ approximation factor, had running times exponential in the solution size $k$, or satisfied the fairness constraints only approximately or in expectation. We further generalize our result to the problem where each point may belong to an arbitrary subset of colors. Given lower and upper bounds $\ell_i$ and $u_i$ for every color $i$, the goal is to find $k$ points whose color counts satisfy all these bounds while maximizing their diversity.

cs.DS↗

ECLIPSE-X: Plate-Scale Calibration as the Binding Constraint on Relativistic Astrometry with Spacecraft Navigation

Precision stellar astrometry near the Sun can constrain the parametrized post-Newtonian (PPN) parameter gamma through gravitational light deflection. This study develops ECLIPSE-X, a localized multi-frame estimation architecture in which one persistent relativistic parameter is estimated simultaneously with frame-dependent spacecraft line-of-sight and roll states, separated by their differing spatial signatures and by the persistence of gamma across frames. The aim is not a competitive determination of gamma, but an assessment of whether such a joint solution and its formal covariance stay statistically valid under realistic astrometric model error. A synthetic field of 250 stars spanning 1.22-8 apparent solar radii is observed over 40 frames at 5-s cadence; weighted least squares estimates gamma with three navigation states per frame, and Schur-complement marginalization gives the formal uncertainty in gamma after navigation coupling. The nominal solution recovers gamma = 1.0001198167 with sigma_gamma = 3.075104e-4, and a 1000-realization ensemble confirms calibration. Perturbations unknown to the estimator identify plate-scale mismatch as the dominant failure mechanism: a 300-ppm scale error alone inflates the empirical-to-formal uncertainty ratio to 7126.6 with zero 1-sigma and 2-sigma coverage, against 2.18 for catalog-coordinate mismatch and 1.12 for a frame-coherent radial perturbation, while every realization converges numerically: convergence does not guarantee statistical validity. Augmenting the state with a plate-scale parameter restores consistency, at a fixed 21.2% cost in the precision of gamma set by the correlation between the two persistent parameters. A calibration sweep shows the unaugmented estimator is covariance-consistent only below sigma_p = 2.7884e-8, so the augmented state is a requirement, not a refinement, for near-Sun relativistic astrometric estimators.

astro-ph.IM↗

Large-scale factor analysis shows machine intelligence is only partially interpretable

A common assumption in language model development is that cognitive abilities are organized around a general, domain-free intelligence factor, like fluid intelligence in humans. This assumption is rarely tested directly, and prior attempts have done so only at a much smaller scale. We take a latent variable approach to intelligence in language models, similar to how psychometricians study psychological constructs. Performance in every specific problem set is influenced by a domain-specific and a domain-agnostic latent factor. Using factor analysis as a dimension-reduction technique, we analyzed 13,251 published evaluation scores covering 1,618 language models across 456 different text-only benchmarks. Due to the super-sparse nature of the dataset, we triangulate our analysis across different data densifiers and imputation methods. A robust pattern across different modes of bias is that 1. A general intelligence factor accounts for 70.8% of variance in model performance at our most generous estimate, and far less than that in most of our solutions, 2. Content-similar benchmarks do not necessarily cluster together, and 3. The $g$ factor is not dominated by any common theme, and there is a lack of evidence that it is well-proxied by standard "intelligence" benchmarks. Our findings go against current endeavors of defining, identifying, and targeting general intelligence as a tangible construct in language model development. This leaves the strategy of targeting a single conceptual ability without support, since the first-order abilities it would have to reach are often partially idiosyncratic and not identifiable in practice.

cs.CL↗

Asymptotic Rigidity and Boundary Structure of Yang's Numerical Invariants for the Bidisk Submodules $[z^k-w^\ell]$

Let \[ M_{k,\ell}=[z^k-w^\ell]\subset H^2(\mathbb D^2), \qquad k,\ell\in\mathbb N,\quad k\ne\ell . \] We study the large-index asymptotics of Yang's numerical invariants and the boundary structure of their generating function. Starting from the exact staircase formula obtained in our preceding work, we prove that \[ Σ_j(M_{k,\ell}) = \frac{C_{k,\ell}}{j} + O(j^{-2}), \] where \[ C_{k,\ell} = \int_0^\infty \frac{x^2} {(x+1/k)^2(x+1/\ell)^2}\,dx . \] The first strict descent together with $C_{k,\ell}$ recovers the unordered pair $\{k,\ell\}$, yielding an asymptotic rigidity principle. At the next order we obtain a periodic correction \[ Σ_j(M_{k,\ell}) = \frac{C_{k,\ell}}{j} + \frac{Ψ_{k,\ell}(j)}{j^2} + O(j^{-3}), \] whose least period is \(\operatorname{lcm}(k,\ell)\), and we determine the leading amplitudes of the strict drops. For Yang's generating function \[ \mathcal P_{k,\ell}(t) = \sum_{j=0}^{\infty}Σ_j(M_{k,\ell})t^j, \] we prove that it has radius of convergence one and a logarithmic singularity at $t=1$; hence it is never a polynomial for $k\ne\ell$, giving a negative answer to Yang's polynomiality question within this quasi-homogeneous family. The periodic higher-order corrections generate root-of-unity polylogarithmic boundary modes. The leading logarithmic coefficient together with the second-order boundary support determines the unordered pair $\{k,\ell\}$, while successive renormalized boundary limits recover the Fourier coefficients of every finite-order periodic asymptotic term.

math.FA↗

Angle Distributions for Intersecting Random Segments in Star-Shaped Planar Domains

Let $Ω\subset\mathbb{R}^2$ be a bounded planar set that is star-shaped with respect to the origin, and let $A,B,C,D$ be independent random points uniformly distributed on $Ω$. We consider the random segments $S_{AB}$ and $S_{CD}$ and study the distribution of the smaller angle $Θ\in[0,π/2]$ formed by them, conditional on the event that they intersect. Using the radial function of $Ω$, together with a parametrization of each segment in terms of its supporting line and the positions of its endpoints along that line, we derive an integral representation for the conditional distribution \[ \Pr\{Θ\leqθ| S_{AB}\cap S_{CD}\neq\varnothing\}. \] The resulting expression makes explicit how the geometry of the boundary of $Ω$ determines the angular distribution. The probability of intersection appears naturally as the normalizing constant and is related to the probabilistic version of Sylvester's four-point problem.

math.PR↗

LIBERO-MAX: Do Robot Policies Adapt When the World Changes?

Robots must often continue a task after a target moves, the viewpoint shifts, or an obstacle appears, even though their earlier observations and committed actions reflect the previous scene. Many simulation robustness benchmarks fix external conditions at reset, leaving this temporal challenge underexamined. We introduce LIBERO-MAX, a benchmark of 8,000 paired cases spanning eight types of changes to geometry, observations, appearance, clutter, and paths. Each pair compares task execution with and without a mid-task event, holding the task, initial state, policy seed, and pre-event action sequence fixed. This controlled comparison distinguishes event-associated regressions from failures already present without the change. Across fourteen current VLA, hybrid, and world-action policies, events reduce success by 11.0-25.7 percentage points. Event profiles reveal shared vulnerabilities to geometry and observation changes, while policy-family rankings interleave. Camera controls show that robustness reflects both competence under the changed conditions and the trajectory from which they are encountered; varying query cadence does not eliminate the gap. Together, the paired protocol and temporal diagnostics establish LIBERO-MAX as a reproducible testbed for diagnosing failures under mid-execution changes and measuring progress toward robot policies that remain effective as the world changes.

cs.RO↗