Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,603 records · Page 89Linked to original sources

A Grassmann formula for the spectral order of matrices

For Hermitian matrices $A$ and $B$, we prove that $(A\join B)\oplus(A\meet B)$ is unitarily equivalent to $A\oplus B$, where the join and meet are taken in Olson's spectral order. Taking traces answers a question of Bourin and Lee: for positive semidefinite matrices, $\Tr(A\join B)=\Tr(A+B)$ if and only if $A\meet B=0$. Iterating the direct-sum identity gives a trace formula for a finite family of positive matrices. The trace of the spectral supremum equals the trace of the sum precisely when the ranges form an algebraic direct sum. For each fixed $1<p<\infty$, the spectral supremum and the sum have equal Schatten $p$-norms if and only if the ranges are pairwise orthogonal. Finally we establish an elegant formula for the Frobenius inner product and the spectral order.

math.FA↗

BAM! Bayesian Anything Model: a foundation model for generative computational imaging

Generative models are transforming Bayesian computational imaging, yet the field still lacks physics-aware foundation models. Current practice falls into two camps. Large foundation image models are deployed as plug-and-play priors with zero-shot approximate likelihood guidance, which introduces significant bias and computational cost. Physics-aware generative models avoid this bias, but each is tied to a specific dataset, task and instrument. We introduce BAM (Bayesian Anything Model), a lightweight foundation model for few-step, physics-aware posterior sampling that generalises robustly to unseen data and tasks, zero-shot or with minimal finetuning. BAM upgrades the operator-conditioned Reconstruct Anything Model (RAM) backbone (Terris et al.) into a conditional flow map, so instrument physics is specified at inference time rather than fixed during training. BAM has just 36M parameters and is pre-trained jointly on large image corpora and libraries of forward operators. A single network then draws posterior samples in a few steps, with no likelihood approximation and no guidance weights to tune. Across linear inverse problems on FFHQ, AFHQ, LSUN, DIV2K and the Kohler camera-shake benchmark, BAM outperforms in just 3 steps both specialised models and leading zero-shot methods in sample quality, at a fraction of their computational cost. BAM gives the community an accessible entry point to generative computational imaging, lowers the economic and environmental cost of training imaging models, and opens a new path for research on physics-aware Bayesian computational imaging. Official page: https://bayesian-anything-model.github.io/

cs.CV↗

The Evolution of Attention in Large Language Models: Mechanisms, Trade-offs, and Emerging Trends

Self-attention gives LLMs fine-grained, query-dependent access to context, but dense token interactions incur quadratic prefill cost and a key--value cache growing with context length. Research thus spans explicit-memory compression, sparse access, recurrent state construction, structured state dynamics, and heterogeneous mechanism composition. This survey analyzes these developments as model-internal contextual memory. We introduce a five-dimensional lens---Memory Representation, Memory Update, Access, Readout, and Integration---describing what is represented, how it changes, what is query-eligible, how it is read, and how readouts form outputs. This lens compares overlapping research lines without imposing one computational model. We reconstruct mechanism-level developments and architectural adoption using 59 release-level records from 14 major model lineages and 11 high-performing open-weight endpoints. First, explicit-memory and recurrent-state methods retain distinct interfaces but increasingly control overlapping memory functions. Second, heterogeneous architectures increasingly coordinate across network depth: layer-wise composition distributes complementary memory processing across representational stages, while cross-layer reuse carries selected memory and routing artifacts forward. Depth thus becomes a dimension along which contextual memory is constructed and managed. Third, these developments motivate a stateful multidimensional memory-routing hypothesis: persistent memory is organized across temporal scope, network depth, substrate type, and representation granularity, while coordinated Sparse Write and Sparse Read determine what is maintained and what contributes to each query. Overall, efficient sequence architecture design increasingly concerns the organization, lifecycle, and selective use of contextual memory rather than an isolated Attention operator.

cs.CL↗

Typographic Attack Against VLM-based AI-generated Image Detection

Vision-language models (VLMs) are increasingly used for AI-generated image (AIGI) detection, providing natural-language explanations for authenticity judgments. However, their ability to interpret text within images may also expose these judgments to misleading semantic cues. We systematically evaluate typographic attack strategies across detection-oriented, open-weight, and commercial VLMs, considering both real-to-fake and fake-to-real attacks. Our results show that reasoning modes generally exhibit greater vulnerability than direct modes and that attack effectiveness exhibits pronounced directional asymmetry. Moreover, larger models tend to exhibit higher clean detection accuracy but also higher attack success rates. We further examine attack robustness under image and text transformations and investigate whether overlays indicating the correct class can aid error correction. Together, these analyses characterize how typographic attacks influence authenticity judgments and expose limitations of current VLM-based AIGI detection systems.

cs.CV↗

Full asymptotics for the parabolic Anderson model with Pareto potential in the heaviest-tailed regime

The parabolic Anderson model is the Cauchy problem for the heat equation on the integer lattice with a random potential $ξ$. We consider the case where $\{ξ(z): z\in \mathbb{Z}^d\}$ are independent and identically distributed Pareto random variables with parameter $α$, and assume that the solution is initially localised at the origin. We establish the full asymptotic behaviour of the total mass of the solution as time tends to infinity in the heaviest-tailed regime $α\in(d,2d)$. In particular, we find qualitatively different behaviour in dimension $d=1$ and in dimensions $d\ge 2$.

math.PR↗

ChronoGraph: Functional 4D Scene Graphs with Vision-Language Models for Interaction Understanding and Grounded Planning

Embodied agents must determine where to act, anticipate the resulting scene changes, and interpret observed outcomes to guide subsequent actions. This requires connecting 4D interaction understanding, which explains how past actions changed the scene, with spatially grounded planning, which determines how and where to act toward a goal and anticipates the resulting scene changes. We introduce ChronoGraph, a functional 4D scene graph that links actions on affordance parts to semantic and geometric state changes. By representing observed and anticipated transitions in the same form, it provides a shared basis for understanding and planning. We construct ChronoGraphBench through an automatic data engine that converts human-interaction videos and simulated robot trajectories into graph-annotated questions for training and evaluating Vision-Language Models (VLMs) on both tasks. Using these annotations, we train ChronoGraphVLM by adapting pretrained VLMs in two stages. Graph-as-Chain-of-Thought supervised fine-tuning teaches the models to reconstruct observed transitions and predict future ones as graph traces before answering. Subsequent joint 4D graph reinforcement learning directly rewards graph properties and answer correctness. Experiments across model scales show improvements over the corresponding pretrained baselines and zero-shot transfer to VLM4D. Real-world demonstrations further show that graph-based planning and affordance grounding support mobile manipulation through existing robot skills without additional fine-tuning.

cs.AI↗

Connected Dominating Set on Semi-Ladder-Free Graphs

We study \textsc{Connected Dominating Set} on graphs whose closed-neighborhood set systems are $d$-semi-ladder-free. This structural condition strictly generalizes the biclique-free setting and provides a natural regime for connectivity-constrained domination. We obtain both a fixed-parameter algorithm and an approximate kernelization framework for the problem on this class. Our algorithmic result is based on a new compact representation theorem for inclusion-wise minimal set covers in $d$-semi-ladder-free set systems. Although the number of minimal set covers of size at most $k$ may be as large as $n^{Ω(k)}$, we show that all such set covers can nevertheless be encoded by a family of at most $k^{kd+1}$ tuples, and that this family can be enumerated in time $\Oh(k^{kd+2}\cdot nm)$. Combining this representation with a \textsc{Group Steiner Tree} subroutine, we obtain an algorithm for \textsc{Connected Set Cover}, which in turn yields an algorithm for \textsc{Connected Dominating Set} running in time $k^{kd+2}\cdot 2^k \cdot n^{\Oh(1)}$ and polynomial space. For the preprocessing result, we introduce grouped domination cores and dominator cores, and prove polynomial upper bounds on their sizes in $d$-semi-ladder-free graphs. Using these structures, we obtain, for every fixed $d$ and $\varepsilon>0$, a polynomial-time $(1+\varepsilon)$-lossy compression for \textsc{Connected Dominating Set} to an equivalent reduced instance of size $k^{\Oh(d^2/\varepsilon)}$. The reduced instance is a \textsc{Connected Dominating Set} instance on a $(d+2)$-semi-ladder-free graph.

cs.DS↗

Distribution of Age of Information in the Erlang Loss System

In this paper, we study the exact distributions of the age of information (AoI) and peak AoI (PAoI) in a bufferless setting in which time-stamped updates, or processed tasks, generated by one or several information sources according to a Poisson process for each source, are submitted to a shared pool of $c$ homogeneous servers with exponentially distributed service times, i.e., the so-called M/M/c/c or the Erlang loss system. We consider three update management policies which come into play when a new update arrives to find all the servers busy: the arriving update is blocked (non-preemptive, NP) as in the Erlang loss system, or it preempts a randomly chosen update of its own source in service (preempt at random, PR), or the stalest such update (preempt the stalest, PS). All three policies preserve the same birth-death structure of the server occupancy process associated with the underlying Erlang loss system. However, their AoI and PAoI distributions can be very different. The approach we take is the absorbing Markov chain (AMC) method, in which a single AoI cycle, rather than the entire sample path of the system, is modeled by an absorbing Markov chain. Via the AMC method, we derive the exact distributions of AoI and PAoI in matrix-exponential form. The method extends to multiple sources sharing the same pool of servers, with the AoI of a given source being affected by the remaining sources only through their aggregate update rate. Numerical examples illustrate the implications of our findings, including distribution-based server provisioning under age violation constraints.

cs.IT↗

QMA(2) with Limited Shared Entanglement

A $\mathsf{QMA}(2)$ protocol involves two provers submitting unentangled witnesses to a polynomial-time quantum verifier. In The Power of Unentanglement (ToC, 2009), Aaronson et al. proposed $\mathsf{QMA}(2;h)$, a variant of $\mathsf{QMA}(2)$ in which the two provers may share $h$ EPR pairs. Our main result shows that the power of $\mathsf{QMA}(2)$ remains unchanged for up to logarithmically many shared EPR pairs: $\mathsf{QMA}(2;h)=\mathsf{QMA}(2)$ for $h=O(\log n)$, where $n$ is the input length. The result follows from a simulation using four unentangled witnesses, combined with the Harrow-Montanaro equality $\mathsf{QMA}(4)=\mathsf{QMA}(2)$ (FOCS, 2010). We also prove monotonicity in the EPR budget: $\mathsf{QMA}(2;h)\subseteq\mathsf{QMA}(2;H)$ for $h\le H$, preserving completeness and soundness. Combined with input padding, this shows that establishing $\mathsf{QMA}(2;n^\varepsilon)=\mathsf{QMA}(2)$ for any fixed $\varepsilon>0$ would imply equality for every polynomially bounded budget, resolving the open problem raised by Aaronson et al. Finally, we extend these results to a variant of the model in which the provers may use local operations and classical communication (LOCC) during witness preparation. For logarithmic-size witnesses and inverse-polynomial gaps, both models remain equivalent to their unentangled counterpart when $h=O(\log n)$. We show that extending this equivalence to any superlogarithmic EPR budget in the LOCC model would imply $\mathsf{NP}\subseteq\mathsf{BQP}$.

quant-ph↗

Separable decompositions of 2xn states with operator Schmidt rank three

We prove that every bipartite state $ρ$ of a qubit coupled to an $n$-level system whose operator Schmidt rank equals three can be written as a mixture of $\text{rank}(ρ)$ pure product states, and as a mixture of at most $n+1$ mixed product states in general. The second result is tight in the sense that there exist states that cannot be decomposed with a smaller number of terms. Our proof gives constructive methods to obtain the decompositions. It involves eigendecompositions of suitable unitary dilations of an operator explicitly determined by the state. Our framework recovers previously known results for the length of decompositions into pure product states and is also able to deal with decompositions into mixed states. For pure product states, our decomposition is different from existing ones.

quant-ph↗

From Local Whole-Body VLA Behaviors to Scene-Scale Aerial Manipulation

Vision-language-action (VLA) models enable task-conditioned interaction, but extending them to scene-scale aerial manipulation remains challenging due to costly whole-body demonstrations, latency-induced action-state misalignment, and cross-site behavior composition. We present a unified framework for synthetic policy training and scene-scale execution on articulated uncrewed aerial manipulators (UAMs). A scene-reconfigurable pipeline synthesizes task-conditioned, kinodynamically feasible trajectories and synchronized multiview observations for VLA training without physical-platform demonstrations. Measured-progress-aligned realization (MPAR) aligns asynchronously returned action chunks with measured execution progress and realizes them as continuous, dynamically feasible trajectories. A relational Scene Graph grounds language goals to object instances and feasible interaction regions, while topology-guided transfer connects local behaviors across sites. Local VLA skills achieve 39/60 successes (65.0%) in simulation under oracle target and feasible-handoff conditions. Under 500-ms added latency, with and without a transient command-update stall, MPAR reduces median takeover phase error by 0.212 s over nominal-time alignment. The complete system completes 21/50 simulated multi-site missions (42.0%) and is further validated on a physical articulated UAM.

cs.RO↗

An electromechanically coupled multiphase-field model with generalized kinetic relations for ferroelectrics

Classical phase-field models typically use the Allen--Cahn evolution law to model interface motion as a consequence of energy-gradient descent. This implies a linear kinetic relation between the velocity of an interface (e.g., a domain wall in a ferroelectric) and its driving force. When nonlinear kinetic relations are to be modeled, as commonly found in ferroelectric ceramics, an alternative evolution law and hence an alternative model is needed. Here, we propose an electromechanically (stress-driven) coupled multiphase-field framework with a general kinetic formulation, which integrates a prescribed kinetic relation directly into the evolution law. We show that this model correctly evolves domain walls with the assigned nonlinear kinetics, e.g., of mixed exponential-power law type, following the so-called Merz--Stadler law. The multiphase formulation further enables distinct kinetic relations and interfacial energies to be assigned to different order-parameter pairs, which reflects the distinguishable properties of the different types of ferroelectric domain walls. The proposed framework is broadly applicable for simulating kinetic relations in materials with applications beyond ferroelectrics.

cond-mat.mtrl-sci↗

Poisson spot and wave scattering of scalar fields by conformal anomaly black holes: probing the near-horizon geometry

We investigate the scattering and diffraction of scalar plane waves by static conformal anomaly black holes. We determine the truncation requirements for the partial-wave series (PWS) method at finite distances and compute the full waveforms, which display a clear diffraction pattern: on the negative $z$-axis the field reduces to the incident plane wave, while on the positive $z$-axis a distorted plane wave coexists with a scattered spherical wave, and a bright Poisson spot appears at $θ=0$ surrounded by concentric diffraction rings. We obtain the on-axis intensity of the Poisson spot as a function of the radial coordinate. The intensity is most sensitive to the near-horizon geometry: close to the horizon it differs markedly from that of the RN black hole, whereas in the far field the two curves converge. We trace this to the metric functions, where different values of $f'(r_+)$ produce different phase increments that compress and shift the interference fringes, while the conformal anomaly correction falls off as $O(\tildeαM^2/r^4)$. By analyzing the potential barrier we study the absorption cross section, and use the series reduction method to accelerate the PWS convergence for the differential scattering cross section. The low- and high-frequency absorption cross sections are strongly correlated with the $\ell=0$ and $\ell$-dependent parts of the potential barrier, respectively. The differential cross section and glory scattering arise mainly from scattering in a finite region excluding a small neighborhood of the outer horizon. Charge and $\tildeα$ have opposite effects on the width of the glory peak but the same effect on its height.

gr-qc↗

NodeGround: A Node Classification Benchmark in the Graph Foundation Model Era

Can a pretrained graph model replace training and tuning a separate predictor for each dataset? Answering this requires evaluating prediction quality alongside computational cost. We present NodeGround, a node classification benchmark that puts graph foundation models (GFMs) and dataset-specific supervised learning under a common evaluation framework. The benchmark spans 51 datasets and evaluates six GFMs alongside 15 supervised methods under two label-availability regimes. Shared data partitions, validation-only model selection, controlled hyperparameter searches, and multiple predictive metrics make comparisons systematic, while workflow measurements account for adaptation, training, tuning, and inference. The results favor carefully tuned graph neural networks overall. GraphPFN reaches third place by Elo when more labels are available, yet its relative strengths vary substantially with dataset properties. Efficiency comparisons further qualify the benefits of pretrained reuse: GVT and GraphPFN appear on the Pareto frontiers when supervised methods are represented by their default and fully tuned configurations. Adding intermediate tuning budgets removes this advantage for GVT and leaves GraphPFN extending the estimated frontier in the label-rich setting alone. Thus, reusing pretrained parameters does not yet provide a broadly reliable route to either stronger predictions or cheaper workflows. We release the evaluation pipeline, run-level records, and an open leaderboard at https://github.com/nums-ai/nodeground.

cs.LG↗

Polylog-depth Quantum Thermal Simulation via Local Recovery Channels

We establish sufficient conditions for preparing quantum thermal states of noncommuting local Hamiltonians with polylogarithmic circuit depth in arbitrary fixed spatial dimension. Our conditions combine locality and stability bounds on the effective interactions of reduced density matrices of a Gibbs state with a quantitative high-temperature condition and access to coarse classical reference Hamiltonians. When these references can be generated locally, the classical preprocessing cost is $N^{1+o(1)}$ for $N$ sites at inverse-polynomial global accuracy. We construct the preparation circuit from spatially localized Petz recovery maps. Quantum corrections derived from the microscopic Hamiltonian enable accurate recovery while the classical references remain coarse. After preprocessing, the resulting circuit prepares the canonical purification using $N\operatorname{polylog}(N/\varepsilon)$ gates and qubits, where $\varepsilon$ is the preparation error. These results link two fundamental questions: how correlations are organized in thermal equilibrium, and how efficiently the corresponding states can be realized through operations allowed by quantum mechanics. By translating static equilibrium structure into explicit preparation circuits, they give equilibrium locality a constructive computational interpretation and a physically grounded role in quantum algorithm design.

quant-ph↗

The starting point of a directed geodesic is special

For any $t\in [0,1/2)$, consider the process $η_t:[0,1/2]\mapsto\mathbb{R}$ defined by $η_t(s)=Π(t+s)-Π(t)$, where $Π$ is the directed geodesic from $(0,0)$ to $(0,1)$ in the directed landscape. Let $0\leq t 0$. This proves Conjecture 14.5 in Dauvergne, Ortmann, and Virág, 2022 in the affirmative.

math.PR↗

A Counterexample to the Liu-Luo-Luo Conjecture on the Harmonic Landau Radius

We disprove a conjecture of Liu, Luo, and Luo (2020) concerning the Landau radius for bounded planar harmonic mappings. We construct an explicit harmonic mapping $F_M$ satisfying the standard normalization conditions, which fails to be locally univalent strictly within the classical holomorphic Landau disk for all $M > M^* \approx 4.451$. By combining this counterexample with the lower bound established by Ponnusamy, Kalaj, and Vuorinen (2014), we demonstrate that the exact asymptotics of the harmonic Landau radius is $\fracπ{8M}$.

math.CV↗

Finite free position of maximal abelian $\ast$-subalgebras of the matrix algebra

Finite free convolution is obtained by averaging characteristic polynomials over Haar unitary conjugation. We ask when two maximal abelian $\ast$-subalgebras of the complex matrix algebra ${\sf M}_n$ can be placed in finite free position: that is, when their relative position realizes this averaging exactly for every pair of elements, one from each subalgebra. Writing such a pair as ${\sf D}_n$ and $U{\sf D}_nU^*$ with $U$ unitary, we characterize finite free position, for either additive or multiplicative convolution, by the condition $|\det U[I,J]|^2=\binom{n}{r}^{-1}$ for every $1\le r\le n$ and all $I,J$ with $|I|=|J|=r$. We show that this condition holds if and only if $n\le3$ and $\sqrt{n} U$ is a complex Hadamard matrix. To quantify the failure of exact realization for $n\ge 4$, we introduce the uniform-minor discrepancy $δ_r(U)$. We identify it with the mean-square error in the $r$-th coefficient of finite free multiplicative convolution for two diagonal matrices whose diagonal entries are independent and uniformly distributed on the unit circle. We establish the symmetry $δ_r(U)=δ_{n-r}(U)$ and the monotonicity $δ_1(U)\leδ_2(U)\le\cdots\le δ_{\lfloor n/2\rfloor}(U)$. For flat unitaries, we derive an explicit formula for $δ_2$, yielding $δ_r(U)\ge\frac{n-3}{2n}$ for $2\le r\le n-2$. Equality for $r=2$ holds precisely when the entrywise square of $\sqrt{n}U$ is also complex Hadamard.

math.OA↗