Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,729 records · Page 96Linked to original sources

Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?

Agent harnesses often improve language models with human-designed workflows, but as models grow more capable, unreliable guidance can increasingly constrain their execution. We call the ability to benefit from useful guidance while overriding unreliable guidance thinking outside the box. We introduce Box$^2$-Bench, which holds the model and task fixed while varying workflow reliability to isolate how models regulate their reliance on guidance. On Box$^2$-Bench, frontier models often benefit from reliable guidance but remain vulnerable when it is misleading or becomes unreliable. To test whether this capability can be learned, we train two open-weight models using bad workflows, reserving good workflows for evaluation. We explore two complementary training strategies: counterfactual supervised fine-tuning improves robustness, while outcome-based reinforcement learning can shift the balance toward greater use of helpful workflows. We further find that this behavior extends beyond workflows to other forms of external information, improving peer correction and robustness to corrupted memory. Together, our results identify selective reliance on fallible external information as a dimension of agent reliability not captured by task performance alone.

cs.CL↗

AVERT-VLN: Abstention-aware Visual Error Recovery and Training for Vision-and-Language Navigation

Deploying vision-and-language navigation (VLN) agents in unseen environments remains challenging because unfamiliar layouts and visual conditions can cause execution to go off track. Rather than relying on continuous human supervision, a practical strategy is to selectively request corrective guidance, recover the ongoing task, and reuse corrective interactions to improve subsequent navigation. We propose Abstention-aware Visual Error Recovery and Training for Vision-and-Language Navigation (AVERT-VLN), a closed-loop framework that uses a plug-in vision-language Monitor for online human-assisted recovery and offline preference learning. The Monitor operates separately from navigation decision generation and assesses instruction-execution consistency from the instruction, visual history, and current observation. To train the Monitor for deviation recognition, we construct LOSTNAV DATASET with 20K counterfactual risk trajectories and rule-based deviation labels. The Monitor is first fine-tuned on 40K normal trajectories to assess instruction progress and then jointly fine-tuned on normal and risk trajectories to recognize semantic deviations. At runtime, Asynchronous Sidecar Monitoring evaluates execution alongside the navigation model. When the controller accepts a LOST verdict, it suspends autonomous execution and requests human guidance for recovery. For offline policy improvement, Trajectory-Anchored Preference Learning converts deviation-associated failures into decision-level preference pairs under shared decision contexts, restricting supervision to the decisions targeted for correction. Under human-assisted evaluation, the full AVERT-VLN system achieves success rates of 76.2% and 66.3% on the val-unseen splits of R2R-CE and RxR-CE, respectively. The same monitoring and human-assisted recovery interface also improves success rates across the three evaluated navigation architectures.

cs.AI↗

On wellposedness and limiting behavior of generalized SQG equations

We consider the two-dimensional generalized surface quasi-geostrophic equations in $\mathbb{R}^2$, given by \[\partial_t θ+u\cdot \nabla θ=0,\quad u=-\nabla^{\perp}Λ^{β-2}θ,\,β\in [1,2).\] When $β=1$, the equation defines the SQG equation and for $β>1$, it defines a family of more singular active scalar equations. We prove that if the interval of existence of the smooth solution to the generalized SQG equations for some $β_0\in[1,2)$ is $[0,T]$, then with the same initial data, the interval of existence of the generalized SQG equations for $β$ close to $β_0$ also contains $[0,T]$. To prove these results, we develop estimates for a conservation law with flux modified around the generalized SQG equations. Furthermore, we also prove some new commutator estimates with bounds uniformly bounded as $β\to 1$ that may be of independent interest.

math.AP↗

Robust Transfer Learning for Paper ECG Recognition

Paper ECG recognition is challenging because real-world ECG images vary in layout, physical artifacts, and label availability. We introduce RobECG-CL, a rank-aware contrastive learning framework for robust paper ECG representation learning. Starting from standard 12-lead ECG recordings, we construct progressively degraded paper ECG views with heterogeneous layouts and train the model to balance same-recording invariance with degradation-aware ordering. Across synthetic stress tests on CODE-II and EchoNext, RobECG-CL improves robustness under severe degradation and few-shot transfer, outperforming contrastive learning baselines and surpassing the waveform-based foundation model, ECG-FM, in the 1% labeled setting. On 312 samples of hospital data with 37 labels, RobECG-CL achieves the best macro AUROC.

cs.LG↗

On the Role of Diffractive Production in Precision Studies of W and Z Bosons at the LHC

The increasing precision of measurements at the Large Hadron Collider requires a detailed understanding of all contributions to electroweak boson production. We estimate the impact of single-diffractive W and Z production on precision observables at $\sqrt{s}=5$, 7, 8, and 13 TeV. The non-diffractive baseline is calculated with DYTurbo, while Herwig provides the relative single-diffractive contribution and its transverse-momentum dependence. Factorization breaking is incorporated through a $p_T$-dependent rapidity-gap survival probability calculated with the dynamic multiparton-interaction model of Pythia8. The diffractive Asimov spectrum is constructed by deterministic bin-by-bin reweighting of the same high-numerical-precision DYTurbo cross section used for the non-diffractive baseline. With the non-perturbative parameters profiled, the nominal relative shifts in $α_s$ are -0.0072%, -0.043%, -0.070%, and -0.067% at 5, 7, 8, and 13 TeV, respectively. Five Pythia8 Pomeron-flux and diffractive-PDF configurations give a maximum spread of $9.4\times10^{-6}$ in $Δα_s$. In the W-mass study, the largest bias from an unmodelled survived single-diffractive contribution is 1.58 MeV, and the largest residual after applying the corresponding model-matched template correction is 1.14 MeV. These shifts are below the relevant fit precision, and no phenomenologically significant impact on $m_W$ is expected within the tested models.

hep-ph↗

MADGRAV: a multilevel anomaly-detection pipeline for gravitational-wave searches applied to LIGO data

We present the results of \textbf{MADGRAV}, a deep-learning-based search for high-mass compact binary coalescences, applied to the data collected by the LIGO interferometers during the third observing run and during the first and second part of the fourth observing run. The \textbf{MADGRAV} pipeline consists of a series of sequential convolutional neural networks that perform anomaly detection, glitch classification, coherence testing, and signal ranking. Data from the Hanford and Livingston LIGO detectors are studied (both individually and in coherence) by way of 1 second Q-transform windows. Of the candidates that survive every stage of the pipeline, 48 reach the significance threshold, and we report 47 gravitational wave detections characterised by a false alarm rate below $1\,{\rm yr}^{-1}$ with a probability of astrophysical origin $p_{\rm astro}>0.9$. Of the 47 detections, 44 are shared with the minimally modelled coherent WaveBurst search. The observed total source-frame masses, extracted from official gravitational wave transient catalogues, are in the $14-236 M_{\odot}$ range with a median of $69 M_{\odot}$, and a median SNR of 16. We note that the recovered fraction of confident detections rises with mass: for LIGO detectors network SNR $>10$ the pipeline recovers $8.1\%$ of confident catalog events below $30 M_{\odot}$, $39.8\%$ between $30$ and $100 M_{\odot}$, and $53.3\%$ above $100 M_{\odot}$, corresponding to $33.3\%$, $45.5\%$ and $53.3\%$ of the events detected by coherent WaveBurst in the same bins. These results suggest that anomaly detection pipelines can serve as an independent detection channel complementary to matched filtering in the high-mass high-SNR regime.

gr-qc↗

Cybersecurity in Edge Computing: A Trust-Aware Federated Hybrid Intrusion Detection Framework

Edge computing has emerged as a critical computing paradigm in modern distributed systems by migrating data processing closer to end users and Internet of Things (IoT) devices. While this paradigm decentralizes processes, minimizes latency, and reduces backhaul bandwidth congestion, it exponentially enlarges the cyberattack surface. Heterogeneous, resource-constrained edge devices deployed across unmanaged administrative domains present highly vulnerable targets. To address these vulnerabilities without compromising global data privacy regulations, this paper proposes a novel Trust-Aware Federated Hybrid Intrusion Detection Framework (TA-FHIDF). The proposed framework integrates an Autoencoder, a 1D Convolutional Neural Network (1D-CNN), and a Bidirectional Long Short-Term Memory (BiLSTM) model into a unified, localized deep learning engine capable of autonomous spatial and temporal feature extraction. Model training is performed collaboratively via federated learning, ensuring raw network telemetry remains isolated at local gateways. Furthermore, to defend against adversarial model poisoning attacks, we introduce a robust server-side trust-aware aggregation mechanism that evaluates client reliability using a cosine similarity metric before global model integration. Empirical evaluations across multi-vector benchmark datasets (UNSW-NB15, CICIDS2017, and Edge-IIoTset) demonstrate the framework's superior detection accuracy, rapid convergence, and high Byzantine fault tolerance under adversarial attack scenarios.

cs.CR↗

A Generalizable and Explainable Framework for Synthetic Video Detection Using First-Digit Gradient Statistics

AI video generators have not only become harder to detect but are used to generate a diverse set of scenarios from landscapes to street views to animal videos. This creates a problem where CNN-based detectors are effective but offer no insight into their inner workings, while forensics-based detectors are often pretrained for a set scenario or become too complex to derive meaningful insights. We present a novel approach to AI video detection using Sobel gradient values analysed with the first-digit law. Using linear discriminant analysis, we visualise the discriminatory signal, while a multi-layer perceptron is used for classification. The detection method has no generator- or scenespecific features, and the model has no knowledge of container formats, codec, bitrate, or compression artefacts. The model is trained and tested on GenBuster-200K, GenBusterBench, GenVA, FaceForensics++ C23, and CelebDF. We also show how zero-shot detection fails even though the feature set carries a discriminatory signal.

cs.CV↗

How Shaped Waves Propagate in Scattering Media

The radiative transport equation provides a powerful framework for describing wave propagation in scattering media. However, it cannot describe coherent waves whose incident wavefront is deliberately tailored. Here we establish a transport theory for wavefront-controlled waves based on a matrix transport equation. Unlike conventional radiative transport, which propagates a scalar radiance, our theory propagates a complex two-by-two matrix that retains phase information and obeys a nonlinear transport equation, despite the underlying wave dynamics being linear. We use this matrix transport equation to derive the spatial profiles of the energy and current densities of transmission eigenchannels, from closed to open channels, in arbitrary diffusive systems. Remarkably, these profiles can be expressed analytically in terms of the solution of the conventional radiative transport equation for a random wave in the same medium. The theory further extends to energy loss, including absorption and leakage into uncontrolled channels, revealing symmetry breaking of transmission eigenchannels in complex structures. Validated against numerical solutions of the wave equation, we establish matrix transport as a framework for describing and controlling coherent wave propagation in complex media.

physics.optics↗

Gauge-energy preservation under congestion-controlled network repair

We study preservation of finite gauge-energy flows under local network repair after Bernoulli edge failures. A macroscopic demand network records terminal pairs to be routed, while a microscopic physical network contains local backup routes, bypasses and shared corridors. We prove a deterministic gauge-energy repair theorem: if the usable demand network $\mathcal{B}^{\sharp}$ carries a finite $Φ$-energy flow $θ$, then the repaired physical network $H$ carries a lifted finite $Φ$-energy flow $Θ$ with $$ \mathcal{E}^Φ_H(Θ) \le L\,β_Φ(K)\,\mathcal{E}^Φ_{\mathcal{B}^{\sharp}}(θ), $$ where $L$ bounds route length, $K$ bounds routing congestion, $\mathcal{E}$ marks energy, and $β_Φ$ is the gauge dilation constant. We then convert this comparison into probabilistic repair criteria: finite-dependent local repair is handled via domination by product measures, and random repair lengths via a variable-cost formulation compatible with chemical-distance estimates. As a main application, we prove a finite-dependent local bypass theorem: any macroscopic network whose supercritical percolation cluster supports a finite gauge-energy flow remains gauge-energy stable after bounded-range local reinforcement, provided the local repair probability is sufficiently high. This yields reinforced lattice and wedge-type examples and provides a potential-theoretic framework for random network repair beyond tree-like or edge-disjoint constructions.

math.PR↗

KilometerVision: A New Frontier for Large-Scale Spatial Intelligence in VLMs

We push the frontier of large-scale spatial intelligence in Vision-Language Models (VLMs) and introduce the first benchmark that probes geographical layout understanding from real-world videos, spanning up to 1km distances. Inspired by the cognitive science literature, we evaluate models against the hierarchical stages of human spatial awareness: anchoring via landmarks, connecting them through routes, and integrating these into global mental maps. Extensive experiments reveal a fundamental divergence in how current AI models process spatial information. Instead of utilising true path integration or forming geometric survey knowledge, we find that VLMs rely almost entirely on 2D visual recognition and text-matching to bypass complex spatial reasoning. The benchmark is publicly available at https://perception-test-challenge.github.io/kilometervision.html.

cs.CV↗

Rate 1/5 Non-Malleable Codes against Entangled Split-State Tampering

We construct efficient information-theoretic non-malleable codes for classical messages that are secure against two noncommunicating local quantum tampering operations with arbitrary pre-shared entanglement. For every sufficiently small fixed $ξ>0$ and all sufficiently large first-share lengths $n$, the codes have rate at least $1/5-ξ$, perfect correctness, and error $2^{-n^{Ω(1)}}$. Security holds for every message, with a single message-independent simulator for each attack. This resolves the constant-rate question for worst-case classical messages in the entangled two-split-state model. Our construction builds on the permutation-based two-split construction of Batra, Boddu, and Jain, which achieves rate approaching $1/5$ for uniformly random messages. We retain their architecture but replace the uniform message input to the permutation with a prescribed message concatenated with fresh uniform padding. Our main contribution is a worst-case security reduction for this modification.

quant-ph↗

SPOON: Towards Coherent Compositional 3D Scene Generation from Uncalibrated Multi-view Images

Compositional 3D scene generation aims to recover complete 3D object shapes and their spatial arrangement from visual observations. Recent image-conditioned 3D generators provide strong priors for producing high-quality object geometry, making the generation of complex scenes increasingly practical. A central challenge is therefore to spatially organize these generated assets into a globally coherent scene while remaining consistent with multi-view observations. Existing approaches either entangle scene layout with object generation or separately estimate spatial placement from view-specific observations, where pose hypotheses may remain ambiguous and inconsistent across views, often resulting in an incoherent object-camera soup. We introduce SPOON, a framework that reformulates multi-view compositional 3D generation as scene-level, geometry-grounded pose reasoning. Rather than treating view-specific object pose hypotheses independently, SPOON coordinates them using reconstruction-derived multi-view geometry through a Guide-Route-Reconcile paradigm. This progressively organizes object poses and camera configurations into a coherent scene-level spatial arrangement. Extensive experiments on ARSG-110K and MIDI-3D-Front demonstrate consistent improvements in object placement and scene composition across varying numbers of input views. On ARSG-110K, SPOON reduces scene-level and object-level Chamfer distances by 12.7% and 17.7%, respectively, compared with a strong baseline.

cs.CV↗

Structural Limits of the Information-Theoretic Uncertainty Decomposition

Uncertainty estimation in machine learning typically decomposes uncertainty into aleatoric uncertainty (AU) and epistemic uncertainty (EU) using the standard information-theoretic framework. However, in practice, two critical issues arise: entanglement (AU and EU are highly correlated) and epistemic collapse (EU magnitude shrinks with increasing model capacity). We analyze this framework on a functional level and discover that significant portions of the assumed AU, EU range are infeasible in finite settings, and cannot be attained with any class probabilities. We characterize how this infeasible region scales with the number of classes and Monte Carlo samples $N$ (e.g., from ensembles with $N$ members), revealing it is bounded by $\text{AU} \leq \log(2)/N$. Crucially, the infeasible region's boundary helps explain epistemic collapse: when model confidence is high, $\text{AU} > \text{EU}$ is guaranteed by this fundamental structural limitation. Our findings show that increasing ensemble size mitigates epistemic collapse by reducing the infeasible area. Lastly, we caution against interpreting AU and EU as independent quantities in low AU regimes, since we show they are coupled when $\text{AU} \leq \log(2)/N$.

cs.CV↗

The Complexity of Single-Interaction Hamiltonians

We study the complexity of Local Hamiltonian problems under the restriction that every local term is an unweighted occurrence of the same fixed interaction. In the resulting Single-Interaction Hamiltonian (SIH) model, once the interaction is fixed, the Hamiltonian is specified entirely by its ordered interaction hypergraph. Our main technical tool is an exact compiler that converts any fixed finite alphabet of local interactions into a single interaction. On a designated auxiliary subspace, the compiled Hamiltonian reproduces the source Hamiltonian exactly, while the complete spectrum below a chosen separation scale is preserved with multiplicity. Combined with a finite-alphabet normalization, this gives a universal polynomial-time reduction from $k$-Local Hamiltonian to SIH$_{3k+3}$ up to a known global energy scale and inverse-polynomial approximation. Sharper constructions give QMA-completeness of SIH$_3$ and of geometrically local SIH$_8$ with bounded degree and interaction range. In the stoquastic setting, we obtain StoqMA-completeness of Stoq-SIH$_3$ and QMA-completeness of 1-Pinned Stoq-SIH$_4$. We also establish hardness results for uniformly system-size-dependent interactions and derive single-interaction formulations of the Hamiltonian quantum PCP conjecture, frustration-free, exponentially precise, and guided Local Hamiltonian problems. These results show that independently chosen local matrices and independently tunable coupling strengths are not necessary for a broad range of Hamiltonian-complexity phenomena.

quant-ph↗

Plasticity in graph metric spaces and their hyperspaces

A metric space is plastic if every bijective nonexpansive self-map is an isometry. We prove that the hyperspace of nonempty compact subsets of every connected, locally finite, regular graph, equipped with the Hausdorff metric associated with the path metric, is plastic. The proof identifies the sets with smallest unit balls as the nonempty subsets of adjacent-twin classes and shows that every nonexpansive bijection induces an automorphism of the quotient graph preserving the sizes of these classes. Singletons are preserved when there are no adjacent twins, but need not be preserved in general. We also prove that $\mathcal{K}(K\times G)$ is plastic for every compact connected metric space $K$ and every connected, locally finite, regular graph $G$, with the supremum metric on the product, and give a more general criterion for such products. Further results include a plastic metric space whose hyperspace is not plastic, plasticity of every tree, and plasticity of the hyperspace of a connected, locally finite graph with only finitely many vertices of minimum degree.

math.GN↗

Neuro-Symbolic Predicate Learning for Semantic Safe Robot Control

As robots are increasingly deployed in everyday environments, ensuring their safety has become a central challenge. Existing methods often encode safety requirements as opaque mathematical/logical formulations or dense cost functions. While effective in specific tasks, they remain difficult to interpret, tightly coupled to individual tasks, and offer limited insight into why a robot action is considered safe or unsafe. To address this limitation, we propose ``Neuro-Symbolic Predicate Learning for Semantic Safe Robot Control'' (NEUPRO), which leverages a differentiable reasoner that can learn reusable safety representations from human-specified safety knowledge. NEUPRO allows practitioners to express task-related safety requirements as transparent symbolic rules, while enabling gradients to propagate through these rules to a feature extractor that maps raw observations to safety-relevant concepts. As a result, the learned feature extractor is (softly) grounded in human-understandable semantics, supports transparent constraint evaluation, and is transferable across tasks. By coupling interpretability with differentiability, NEUPRO moves beyond opaque cost design toward reusable safety reasoning. To evaluate NEUPRO's capability, we collect and release REASON, the first real robot benchmark dataset for interpretable robot safety specification. Experiments on REASON show that NEUPRO learns safety-critical features that generalize across tasks, mitigate the interpretability limitations of conventional black-box cost formulations, and provide explicit explanations of safety violation.

cs.RO↗

Convergence of Practical Muon with Finite Newton-Schulz Iterations and Nesterov Momentum

Practical Muon maintains momentum and performs a small, fixed number of Newton--Schulz iterations separately for each parameter matrix, often with a Nesterov correction. We analyze these layer-wise finite-step updates jointly on a coupled nonconvex objective, rather than replacing them by exact polar factors or one global orthogonalization. Under gradient-dependent $(\mathcal L_0,\mathcal L_1,q)$-smoothness and conditionally unbiased stochastic gradients with bounded layer-wise variance, we establish an $\mathcal O(T^{-1/4})$ bound on the expected average Frobenius gradient norm. The analysis retains the Nesterov recursion and requires neither bounded stochastic gradients, symmetric noise, nor a uniform positive lower bound on the nonzero output singular values. Its constants contain no explicit matrix-dimension or rank factors when the number of blocks and problem constants are fixed. The proof follows a descent inequality and a decomposition of the momentum tracking error into initialization, noise, and drift. For the original five-step quintic, we verify the required scalar-map bounds analytically; the result also allows step-dependent coefficients satisfying the same bounds. A complementary nuclear-norm result quantifies rank dependence under a stronger spectral condition. The vanishing rate uses coupled learning-rate and momentum schedules, including the standard single-coefficient Nesterov rule.

cs.LG↗