Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,801 records · Page 100Linked to original sources

Neural Audio Codec for Robust Audio Deepfake Detection

Audio deepfake detectors are typically evaluated on uncompressed audio, although real-world audio often undergoes low-bitrate coding. In this work, we investigate how audio coding affects deepfake detection across codecs, bitrates, and detectors, finding higher errors at lower rates. A mixed-pair protocol isolates codec-induced changes in bona fide and spoof audio, revealing asymmetric, codec-dependent failures: low-rate DAC and EnCodec mainly degrade bona fide detection, whereas X-Codec shows a stronger spoof-side limitation. Motivated by these, we propose a forensic-preserving neural audio codec (FP-NAC), which fine-tunes a pretrained codec using a detector-guided objective while preserving its native hard quantization path and bitrate. On ASVspoof 2019 LA, FP-NAC reduces EER by up to 49.8~pp compared with the original DAC at 0.5~kbps while maintaining comparable reconstruction quality. Although supervised by only one detector, FP-NAC improves performance across multiple detectors, highlighting forensic transparency as a codec design objective alongside perceptual quality. Our codes are available at https://github.com/kjungwoo03/FP-NAC.

cs.SD↗

Making Waves: A Membrane-Coupled Delta Array for Manipulating Objects Below the Actuator Spacing

Distributed manipulator systems manipulate objects through the coordinated motion of many actuators. However, an object must be supported by several actuators at once, so the centre-to-centre actuator spacing imposes a hard lower bound on manipulable object size. We remove this bound by coupling the end-effectors of an 8 x 8 array of three degrees-of-freedom delta robots with a stretchable fabric, turning 64 discrete contacts into a continuous surface capable of manipulating objects smaller than the actuator spacing. Viewing the array as a displacement field over that surface, we investigate local quasi-static and cyclic fields as manipulation primitives. These primitives can be applied globally across the array or locally confined around each tracked object to independently manipulate several objects in parallel. We then train a policy acting on low-order discrete cosine transform coefficients: at equal action dimension, commanding a nineteen-delta neighbourhood halves the placement error of commanding the whole array. The policy transfers to hardware without adaptation at 76% success. The platform manipulates objects from 15mm-90mm, a six-fold range spanning both sides of the actuator spacing, 43.3mm, on a single surface.

cs.RO↗

Measurement of the triple-differential cross section of Z+jet production in proton-proton collisions at $\sqrt{s}$ = 13 TeV

A measurement is presented of the triple-differential production cross section of Z/$γ^*$($\to$ $μμ$)+jet events using proton-proton collision data recorded at a center-of-mass energy of 13 TeV by the CMS experiment at the CERN LHC. The data, collected in the years 2016$-$2018, correspond to an integrated luminosity of 138 fb$^{-1}$. The cross section is measured as a function of the transverse momentum of the muon pair, the boost in rapidity of the center-of-mass system of the muon pair and the leading jet with respect to the laboratory frame, and half of their absolute rapidity separation. This choice of observables ensures sensitivity to the scattering angle in the center-of-mass frame and the fractional momenta of the interacting partons. After unfolding for detector effects in all three dimensions simultaneously, the results are compared with predictions at next-to-next-to-leading order in perturbative quantum chromodynamics corrected for electroweak and nonperturbative effects. This measurement provides valuable information towards characterizing the partonic structure of the proton and determining the strong coupling.

hep-ex↗

CYNAR: a trajectory-based estimator of entropy production

Inferring the entropy production rate $σ$ from recorded trajectories is a key challenge in stochastic thermodynamics. In experiments, the forces governing the dynamics are usually unknown and often only some of the system's degrees of freedom can be observed. Building on a recently introduced method [I. Di Terlizzi, Phys. Rev. Lett. 135, 237101 (2025)], we develop and validate CYNAR (computation yields nonequilibrium analysis and reconstruction), a pipeline that estimates $σ$ from trajectories. It exploits a kinetic decomposition of $σ$ into a traffic term ${\cal T}$, inferred from the short-time curvature of correlation functions, and an inflow rate ${\cal G}$, computed from the steady-state score, i.e., the gradient of the log-density. We show that the same decomposition, evaluated on any subset of observed coordinates, always bounds $σ$ from below, for any diffusion tensor and without knowledge of the hidden degrees of freedom. We then compare CYNAR with a more standard approach that estimates $σ$ from steady-state probability currents, on four systems of increasing complexity. CYNAR proves to be more reliable, with no detectable bias close to equilibrium, weak dependence on spatial discretization, and marked robustness to measurement noise. It is also readily applicable to high-dimensional systems, where the traffic alone provides a lower bound close to $σ$, and remains informative under partial observation, including cases where the flux-based estimate vanishes identically. CYNAR is provided as an open-source Python package.

cond-mat.stat-mech↗

On the Quotient of a Pseudo-null Module

Motivated by questions arising in noncommutative Iwasawa theory, let $R$ be a (not necessarily commutative) ring, $T \in R$ a regular central element, and $M$ a pseudo-null $R$-module. We investigate necessary and sufficient conditions under which the quotient $M/TM$ is pseudo-null as an $R/TR$-module. We first give necessary and sufficient conditions in terms of associated prime ideals when $R$ is commutative. Then we apply this to the case when $R$ is a Krull domain and obtain a precise relationship between the characteristic ideals of $M/TM$ and the $T$-torsion submodule $M[T]$. We give necessary and sufficient conditions in terms of $\operatorname{Ext}$ groups when $R$ is a noncommutative ring and then compare them with the criterion when $R$ is commutative. Lastly, we give an application to noncommutative Iwasawa theory. Roughly speaking, if `big' dual fine Selmer group is pseudo-null, then `most' specialized dual fine Selmer group is pseudo-null.

math.NT↗

Testing Induced-Subgraph Freeness in Outerplanar Graphs under the Random-Neighbor Oracle

We prove that, for every fixed nonempty graph $H$, induced-$H$-freeness is testable with $\varepsilon^{-O_H(1)}$ queries on outerplanar graphs with no maximum-degree bound in the $\textit{random-neighbor model}$, where each query at a vertex returns a uniformly random neighbor. Thus, the query complexity is polynomial in $1/\varepsilon$ and independent of the number $n$ of vertices. Previously, the best bound known for this problem was the $\operatorname{poly}(\log n)$-query guarantee that follows from the general outerplanar-graph tester of Babu, Khoury, and Newman (2016) in the stronger $\textit{adjacency-list model}$, which provides exact degree queries and indexed access to neighbors. Our tester has $\textit{two-sided error}$, which is necessary in general: induced-$P_3$-freeness has no one-sided constant-query tester in the random-neighbor model, even on outerplanar graphs of maximum degree two.

cs.DS↗

Diffusable Latents from Structure-Agnostic Distillation

Distilling pretrained foundation models into an autoencoder bottleneck improves latent diffusability, enabling diffusion models to converge faster and reach higher sample quality. Standard distillation aligns the latent at each position to a co-located teacher feature, tying the latent layout to the teacher's. We show this constraint is unnecessary: aligning a single pooled image-level descriptor to the teacher's performs as well as or slightly better than dense position-wise distillation. We compare first-order and relational pooled objectives across latent shapes and teacher modalities. First-order matching extends naturally to 1D token-sequence latents and across modalities, where distilling a text encoder into an image autoencoder still improves diffusability; a relational objective based only on each image's nearest neighbours improves it as well. Code and blog post are available at https://github.com/AdrienRR/structure-agnostic-distillation and https://kyutai.org/blog/2026-09-28-structure-agnostic-distillation/.

cs.CV↗

Graph Residual Conjugate Diffusion: SNR-Equalized Heat Flow for Graph Signals

Diffusion models generate data by reversing a forward corruption process that typically approaches a simple Gaussian prior. Recent work has extended this framework to signals supported on fixed graphs, e.g., road-network traffic and sensor-network measurements. Many graph signals have nonuniform spectral energy, whereas isotropic corruption adds the same conditional noise variance to every graph-frequency mode. Driving all modes to near-zero terminal signal-to-noise ratio (SNR) requires strong corruption, which increases the noise range that must be covered under a fixed sampling budget. We introduce Graph Residual Conjugate Diffusion (GRCD), which replaces the shared clock of graph heat diffusion with a mode-dependent clock that gives every graph-Fourier mode the same conditional SNR. GRCD fits a zero-mean graph-spectral Gaussian reference on the training split and stops at a finite terminal SNR at which the propagated reference still carries the fitted spectral variances. The Gaussian component has an exact modewise propagator in the probability-flow ODE, so sampling advances it analytically and integrates only the learned residual score numerically. We evaluate GRCD on five settings (METR-LA traffic, Molene weather, and three stochastic block models) against seven comparators under a matched protocol: Graph-Aware Diffusion (GAD), EDM (graph backbone), two adaptations of Whitened Score Diffusion (WSD), and three preconditioning controls. At four function evaluations (NFEs), GRCD lowers averaged maximum mean discrepancy (aMMD) by 22 to 36 times over the best comparator on all five settings, reaching 0.054 on METR-LA, where it clears an aMMD 0.1 target with 87% less sampling wall-clock time than the cheapest comparator that reaches it. Fitting the terminal reference reduces aMMD by 2.7 to 7.3 times at finite terminal SNR, while the factors shrink to 1.00 to 1.01 near zero.

cs.LG↗

A Grassmann formula for the spectral order of matrices

For Hermitian matrices $A$ and $B$, we prove that $(A\join B)\oplus(A\meet B)$ is unitarily equivalent to $A\oplus B$, where the join and meet are taken in Olson's spectral order. Taking traces answers a question of Bourin and Lee: for positive semidefinite matrices, $\Tr(A\join B)=\Tr(A+B)$ if and only if $A\meet B=0$. Iterating the direct-sum identity gives a trace formula for a finite family of positive matrices. The trace of the spectral supremum equals the trace of the sum precisely when the ranges form an algebraic direct sum. For each fixed $1<p<\infty$, the spectral supremum and the sum have equal Schatten $p$-norms if and only if the ranges are pairwise orthogonal. Finally we establish an elegant formula for the Frobenius inner product and the spectral order.

math.FA↗

BAM! Bayesian Anything Model: a foundation model for generative computational imaging

Generative models are transforming Bayesian computational imaging, yet the field still lacks physics-aware foundation models. Current practice falls into two camps. Large foundation image models are deployed as plug-and-play priors with zero-shot approximate likelihood guidance, which introduces significant bias and computational cost. Physics-aware generative models avoid this bias, but each is tied to a specific dataset, task and instrument. We introduce BAM (Bayesian Anything Model), a lightweight foundation model for few-step, physics-aware posterior sampling that generalises robustly to unseen data and tasks, zero-shot or with minimal finetuning. BAM upgrades the operator-conditioned Reconstruct Anything Model (RAM) backbone (Terris et al.) into a conditional flow map, so instrument physics is specified at inference time rather than fixed during training. BAM has just 36M parameters and is pre-trained jointly on large image corpora and libraries of forward operators. A single network then draws posterior samples in a few steps, with no likelihood approximation and no guidance weights to tune. Across linear inverse problems on FFHQ, AFHQ, LSUN, DIV2K and the Kohler camera-shake benchmark, BAM outperforms in just 3 steps both specialised models and leading zero-shot methods in sample quality, at a fraction of their computational cost. BAM gives the community an accessible entry point to generative computational imaging, lowers the economic and environmental cost of training imaging models, and opens a new path for research on physics-aware Bayesian computational imaging. Official page: https://bayesian-anything-model.github.io/

cs.CV↗

The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends

Self-attention gives LLMs fine-grained, query-dependent access to context, but dense token interactions incur quadratic prefill cost and a key--value cache growing with context length. Research thus spans explicit-memory compression, sparse access, recurrent state construction, structured state dynamics, and heterogeneous mechanism composition. This survey analyzes these developments as model-internal contextual memory. We introduce a five-dimensional lens---Memory Representation, Memory Update, Access, Readout, and Integration---describing what is represented, how it changes, what is query-eligible, how it is read, and how readouts form outputs. This lens compares overlapping research lines without imposing one computational model. We reconstruct mechanism-level developments and architectural adoption using 59 release-level records from 14 major model lineages and 11 high-performing open-weight endpoints. First, explicit-memory and recurrent-state methods retain distinct interfaces but increasingly control overlapping memory functions. Second, heterogeneous architectures increasingly coordinate across network depth: layer-wise composition distributes complementary memory processing across representational stages, while cross-layer reuse carries selected memory and routing artifacts forward. Depth thus becomes a dimension along which contextual memory is constructed and managed. Third, these developments motivate a stateful multidimensional memory-routing hypothesis: persistent memory is organized across temporal scope, network depth, substrate type, and representation granularity, while coordinated Sparse Write and Sparse Read determine what is maintained and what contributes to each query. Overall, efficient sequence architecture design increasingly concerns the organization, lifecycle, and selective use of contextual memory rather than an isolated Attention operator.

cs.CL↗

Typographic Attack Against VLM-based AI-generated Image Detection

Vision-language models (VLMs) are increasingly used for AI-generated image (AIGI) detection, providing natural-language explanations for authenticity judgments. However, their ability to interpret text within images may also expose these judgments to misleading semantic cues. We systematically evaluate typographic attack strategies across detection-oriented, open-weight, and commercial VLMs, considering both real-to-fake and fake-to-real attacks. Our results show that reasoning modes generally exhibit greater vulnerability than direct modes and that attack effectiveness exhibits pronounced directional asymmetry. Moreover, larger models tend to exhibit higher clean detection accuracy but also higher attack success rates. We further examine attack robustness under image and text transformations and investigate whether overlays indicating the correct class can aid error correction. Together, these analyses characterize how typographic attacks influence authenticity judgments and expose limitations of current VLM-based AIGI detection systems.

cs.CV↗

Full asymptotics for the parabolic Anderson model with Pareto potential in the heaviest-tailed regime

The parabolic Anderson model is the Cauchy problem for the heat equation on the integer lattice with a random potential $ξ$. We consider the case where $\{ξ(z): z\in \mathbb{Z}^d\}$ are independent and identically distributed Pareto random variables with parameter $α$, and assume that the solution is initially localised at the origin. We establish the full asymptotic behaviour of the total mass of the solution as time tends to infinity in the heaviest-tailed regime $α\in(d,2d)$. In particular, we find qualitatively different behaviour in dimension $d=1$ and in dimensions $d\ge 2$.

math.PR↗

ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning

Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This requires connecting 4D interaction understanding, which explains how past actions changed the scene, with spatially grounded planning, which determines how and where to act toward a goal and anticipates the resulting scene changes. We introduce ChronoGraph, a functional 4D scene graph that links actions on affordance parts to semantic and geometric state changes. By representing observed and anticipated transitions in the same form, it provides a shared basis for understanding and planning. We construct ChronoGraphBench through an automatic data engine that converts human-interaction videos and simulated robot trajectories into graph-annotated questions for training and evaluating Vision-Language Models (VLMs) on both tasks. Using these annotations, we train ChronoGraphVLM by adapting pretrained VLMs in two stages. Graph-as-Chain-of-Thought supervised fine-tuning teaches the models to reconstruct observed transitions and predict future ones as graph traces before answering. Subsequent joint 4D graph reinforcement learning directly rewards graph properties and answer correctness. Experiments across model scales show improvements over the corresponding pretrained baselines and zero-shot transfer to VLM4D. Real-world demonstrations further show that graph-based planning and affordance grounding support mobile manipulation through existing robot skills without additional fine-tuning.

cs.AI↗

Connected Dominating Set on Semi-Ladder-Free Graphs

We study \textsc{Connected Dominating Set} on graphs whose closed-neighborhood set systems are $d$-semi-ladder-free. This structural condition strictly generalizes the biclique-free setting and provides a natural regime for connectivity-constrained domination. We obtain both a fixed-parameter algorithm and an approximate kernelization framework for the problem on this class. Our algorithmic result is based on a new compact representation theorem for inclusion-wise minimal set covers in $d$-semi-ladder-free set systems. Although the number of minimal set covers of size at most $k$ may be as large as $n^{Ω(k)}$, we show that all such set covers can nevertheless be encoded by a family of at most $k^{kd+1}$ tuples, and that this family can be enumerated in time $\Oh(k^{kd+2}\cdot nm)$. Combining this representation with a \textsc{Group Steiner Tree} subroutine, we obtain an algorithm for \textsc{Connected Set Cover}, which in turn yields an algorithm for \textsc{Connected Dominating Set} running in time $k^{kd+2}\cdot 2^k \cdot n^{\Oh(1)}$ and polynomial space. For the preprocessing result, we introduce grouped domination cores and dominator cores, and prove polynomial upper bounds on their sizes in $d$-semi-ladder-free graphs. Using these structures, we obtain, for every fixed $d$ and $\varepsilon>0$, a polynomial-time $(1+\varepsilon)$-lossy compression for \textsc{Connected Dominating Set} to an equivalent reduced instance of size $k^{\Oh(d^2/\varepsilon)}$. The reduced instance is a \textsc{Connected Dominating Set} instance on a $(d+2)$-semi-ladder-free graph.

cs.DS↗

Distribution of Age of Information in the Erlang Loss System

In this paper, we study the exact distributions of the age of information (AoI) and peak AoI (PAoI) in a bufferless setting in which time-stamped updates, or processed tasks, generated by one or several information sources according to a Poisson process for each source, are submitted to a shared pool of $c$ homogeneous servers with exponentially distributed service times, i.e., the so-called M/M/c/c or the Erlang loss system. We consider three update management policies which come into play when a new update arrives to find all the servers busy: the arriving update is blocked (non-preemptive, NP) as in the Erlang loss system, or it preempts a randomly chosen update of its own source in service (preempt at random, PR), or the stalest such update (preempt the stalest, PS). All three policies preserve the same birth-death structure of the server occupancy process associated with the underlying Erlang loss system. However, their AoI and PAoI distributions can be very different. The approach we take is the absorbing Markov chain (AMC) method, in which a single AoI cycle, rather than the entire sample path of the system, is modeled by an absorbing Markov chain. Via the AMC method, we derive the exact distributions of AoI and PAoI in matrix-exponential form. The method extends to multiple sources sharing the same pool of servers, with the AoI of a given source being affected by the remaining sources only through their aggregate update rate. Numerical examples illustrate the implications of our findings, including distribution-based server provisioning under age violation constraints.

cs.IT↗

QMA(2) with Limited Shared Entanglement

A $\mathsf{QMA}(2)$ protocol involves two provers submitting unentangled witnesses to a polynomial-time quantum verifier. In The Power of Unentanglement (ToC, 2009), Aaronson et al. proposed $\mathsf{QMA}(2;h)$, a variant of $\mathsf{QMA}(2)$ in which the two provers may share $h$ EPR pairs. Our main result shows that the power of $\mathsf{QMA}(2)$ remains unchanged for up to logarithmically many shared EPR pairs: $\mathsf{QMA}(2;h)=\mathsf{QMA}(2)$ for $h=O(\log n)$, where $n$ is the input length. The result follows from a simulation using four unentangled witnesses, combined with the Harrow-Montanaro equality $\mathsf{QMA}(4)=\mathsf{QMA}(2)$ (FOCS, 2010). We also prove monotonicity in the EPR budget: $\mathsf{QMA}(2;h)\subseteq\mathsf{QMA}(2;H)$ for $h\le H$, preserving completeness and soundness. Combined with input padding, this shows that establishing $\mathsf{QMA}(2;n^\varepsilon)=\mathsf{QMA}(2)$ for any fixed $\varepsilon>0$ would imply equality for every polynomially bounded budget, resolving the open problem raised by Aaronson et al. Finally, we extend these results to a variant of the model in which the provers may use local operations and classical communication (LOCC) during witness preparation. For logarithmic-size witnesses and inverse-polynomial gaps, both models remain equivalent to their unentangled counterpart when $h=O(\log n)$. We show that extending this equivalence to any superlogarithmic EPR budget in the LOCC model would imply $\mathsf{NP}\subseteq\mathsf{BQP}$.

quant-ph↗

Separable decompositions of 2xn states with operator Schmidt rank three

We prove that every bipartite state $ρ$ of a qubit coupled to an $n$-level system whose operator Schmidt rank equals three can be written as a mixture of $\text{rank}(ρ)$ pure product states, and as a mixture of at most $n+1$ mixed product states in general. The second result is tight in the sense that there exist states that cannot be decomposed with a smaller number of terms. Our proof gives constructive methods to obtain the decompositions. It involves eigendecompositions of suitable unitary dilations of an operator explicitly determined by the state. Our framework recovers previously known results for the length of decompositions into pure product states and is also able to deal with decompositions into mixed states. For pure product states, our decomposition is different from existing ones.

quant-ph↗

Refine your search to explore more results.