Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,009 records · Page 56Linked to original sources

Asymmetric Scout-Worker Reconnaissance for Route Validation in Unknown Environments

This paper studies asymmetric scout-worker reconnaissance in unknown environments, where a small, agile autonomous scout explores routes for a larger worker robot that must visit an ordered sequence of goal locations. Because the scout has a smaller footprint and greater mobility, a scout-traversable route may be infeasible for the worker; worker feasibility must therefore be inferred from scout observations. This setting is not explicitly addressed by existing exploration and replanning methods, which typically assume a single traversability model and seek optimal paths for the same robot performing the exploration. We introduce a symbiotic scout-based framework that exploits the scout's superior mobility to explore only the portions of the unknown environment needed to identify worker-feasible path segments connecting the ordered goals. Evaluations in simulated and real-world settings demonstrate that the proposed approach validates feasible routes, repairs blocked segments with validated worker-feasible detours, and substantially reduces scout travel compared to baseline exploration and planning methods. A real-world indoor deployment further demonstrates the scout navigating narrow corridors to identify a worker-feasible route.

cs.RO↗

Frustrated Gd3+ Double Perovskites as High-Performance Magnetocaloric Materials for Sub-100 mK Adiabatic Demagnetization Refrigeration

Achieving temperatures below 100 mK is essential for advancing quantum technologies and exploring fundamental quantum phenomena. While paramagnetic salts have traditionally enabled adiabatic demagnetization refrigeration (ADR), their limitations have driven the search for more effective alternatives. In this work, we present Gd3+-based double perovskites, Ba2GdSbO6 and Sr2GdSbO6, as high-performance magnetocaloric materials. Starting ADR from 2 K and 5 T, these compounds reach 67 mK and 68 mK in small finite magnetic fields, and 83 mK and 78 mK in zero field, respectively, which are the lowest reported ADR temperatures for Gd3+ magnets (S = 7/2) under these conditions. The frustrated geometry and the complex interplay of exchange and dipolar interactions suppress their antiferromagnetic ordering temperatures to 100 and 166 mK for Ba2GdSbO6 and 190 mK for Sr2GdSbO6, while maintaining an outstanding entropy density of 189 mJ K-1 cm-3 and 201 mJ K-1 cm-3, respectively. Notably, Ba2GdSbO6 sustains sub-100 mK cooling even at finite fields of 0.5 T, making it particularly promising for practical ADR applications.

cond-mat.mtrl-sci↗

Invariance is Compositional for Continuous-time Systems: From Sleekness to Lebesgue Density

This work establishes the first bidirectional compositional invariance result: robust forward invariance of a global Cartesian product set is shown to be equivalent to robust forward invariance of each local subsystem under coupling inputs from its neighbors. To facilitate this, we introduce the notion of tangential Lebesgue-density, a new condition that is weaker than classical sleekness but sufficient to ensure that the product of individual tangent cones equals the tangent cone of the product set. This equivalence reduces the curse of dimensionality by allowing the verification of a high-dimensional global system to be decomposed into a series of local sub-checks. The framework's scalability is demonstrated through a DC microgrid numerical example, confirming that the verification complexity grows only linearly with the number of subsystems.

eess.SY↗

Reactive Real-Time Flow Policies via Asynchronous Distribution Alignment

Generalist robot policies such as vision-language-action models (VLAs) have achieved remarkable generalization, but their inference delays can conflict with the demands of real-time control. Asynchronous execution avoids pauses between action chunks by predicting the next sequence of actions while the robot carries out the previous one. In this paper, we study whether asynchronous execution produces the same action distribution as the original VLA. We find that, for non-Markovian demonstrations, asynchronous execution can produce a fundamentally different action distribution, which can limit the policy's reactivity. In our method, we seek to restore this reactivity by aligning the asynchronously produced action distribution with that of the original VLA through two complementary mechanisms. First, Recursive Flow-Field Distillation trains the asynchronous policy using the VLA's action-generation flow. We characterize the learned distribution theoretically and show experimentally that our asynchronous policy can generate nearly the full range of actions the original VLA would produce, while existing asynchronous methods recover only a fraction of that range. Second, Propose-Resolve prepares multiple action sequences asynchronously and uses the latest observation to select among them based on a lightweight approximation of their likelihood under the VLA's action distribution. Our resulting method matches the original VLA's success on LIBERO and retains about 80% of its success on RoboMimic, about 30 percentage points more than existing asynchronous methods.

cs.RO↗

Exponential Advantage of Quantum over Classical References in Leakage Detection

Leakage from a known encoding subspace can be detected by projection. However, the corresponding projector is unavailable when the encoding subspace is not classically specified. Here we show how independent quantum references prepared by the same encoder enable leakage detection without a classical description of the encoding. A coherent measurement on the references and a single message leaves every state within the two-dimensional encoding subspace, including its entanglement with a remote system, exactly unchanged. We derive the exact detection law for $M$ ideal references and prove optimality among tests with zero false alarm for every encoding. For orthogonal leakage, the miss probability is asymptotic to $4/M$, independent of the ambient dimension $d$. Measuring all references first, even collectively, gives zero detection at every finite budget under the same zero-false-alarm requirement. For a fixed detection target between zero and one, the optimal measurement-first cost is $Θ(d/ε)$ at sufficiently small tolerance $ε$ on normal false alarm and conditional disturbance; a dimension-independent coherent budget suffices. For one logical qubit encoded in three physical qubits, seven coherent references achieve at least $50\%$ orthogonal-leakage detection, whereas any measurement-first receiver needs at least $588$ under the same $1\%$ normal tolerances. Quantum references thus support leakage checks without classical reconstruction, with a sample advantage exponential in the number of physical qubits.

quant-ph↗

Emergent Symmetries in Ensemble Averages

It is widely believed that there are no exact global symmetries in quantum gravity. We review recent works showing that global symmetries can nevertheless emerge after ensemble averages of theories, as encountered in holography. As a precision case study, we discuss the ensemble average of two-dimensional Narain-type conformal field theories associated with an even quadratic form of general signature, which is computed by the Siegel-Weil formula and interpreted as a sum over geometries in a three-dimensional Abelian Chern-Simons theory. Global symmetries of the bulk anyons can emerge after the average, as T-dualities relating different theories in the ensemble are "folded" into symmetries of a single theory (duality origami). We discuss the relation to the Swampland program, including the emergence of global symmetries at infinite distance. Finally, we recast strong-to-weak spontaneous symmetry breaking (SWSSB) of mixed states as an emergence of symmetries in ensemble averages; in holography, wormholes connecting the two sides of a thermofield-double construction provide its bulk realization. This motivates a new Swampland conjecture.

hep-th↗

Planning Oriented 3D Scene Completion via Coupled TUDF Occupancy Representation Learning from Partial Observations

Partial observability remains a fundamental challenge in robotic navigation, where limited sensor coverage and occlusions leave large portions of the environment unobserved. Existing scene completion methods primarily focus on improving incomplete mapping or reconstructing partially observed 3D structures, but rarely investigate how scene completion can be designed to benefit downstream tasks such as path planning. In this work, we propose a path-planning-oriented 3D scene completion framework that moves beyond pure occupancy modeling toward a coupled geometric formulation. Specifically, given partial LiDAR observations as input, the proposed framework jointly predicts completed Truncated Unsigned Distance Field (TUDF)-based continuous geometric representations and voxel-wise occupancy maps. This coupled representation allows the network to better reason about obstacle boundaries and free-space geometry. To fully exploit the synergy between the two representations, we introduce a bidirectionally coupled learning scheme, where TUDF features provide dense geometric guidance to improve occupancy reconstruction, while occupancy features in turn offer complementary structural constraints that refine distance-field estimation. Consequently, the proposed network directly predicts complete occupancy and TUDF representations, allowing seamless integration of TUDF into trajectory planning without post-processing. Extensive experiments on unseen environments demonstrate that the proposed method consistently improves both geometric reconstruction quality and downstream planning performance.

cs.RO↗

DraftTrace: A Multi-View Analytics Environment for AI-Integrated Writing

Generative AI has changed how students produce writing assignments. The final artifact is no longer sufficient to understand the process through which it was produced. We introduce DraftTrace, a writing environment that jointly captures three complementary views of writing: the final product, the writing process and interactions with an integrated AI-assistant. DraftTrace reconstructs how a document develops over time and organizes these signals into submission, longitudinal, and class-level analytics for instructors. We deployed DraftTrace in a graduate NLP course with 81 students and compared their sessions with LLM-generated responses entered by automated tools and with copy-typed responses. While product measures distinguish differences in text formulation, process measures distinguish differences in how text is entered. Considering both views together helps characterize cases such as copy-typing. Interaction traces show that students use the assistant differently across stages of writing: to clarify the question at an early stage and to verify answers at a later stage. A preliminary instructor survey highlights the importance of multi-view writing analytics and their interpretability.

cs.CL↗

SCCM: Spherically Consistent Coarse Matching for ERP Dense Feature Correspondence

Equirectangular projection (ERP) is the standard representation for 360$^\circ$ imagery, and robust dense feature matching on ERP underpins panoramic stereo, view synthesis, and omnidirectional SLAM. Dense matchers trained on flat images degrade systematically on ERP because the chart introduces three coupled distortions -- topological, metric, and area -- that standard coarse matching and visibility estimation do not explicitly model. We show that correcting the three distortions at the coarse-stage interfaces where they arise -- pairwise distortions in attention, per-pixel distortion in covisibility gating -- improves PCK@$1^\circ$ from 0.229 to 0.275 on Matterport3D under a fixed coarse scaffold, with the refiner architecture unchanged -- our central result. Concretely, SCCM (Spherically Consistent Coarse Matching) augments a chart-naive cross-attention/dual-softmax coarse matcher with two sphere-derived priors: Spherical Positional Attention (SPA) pairs a yaw-periodic RoPE (topology) with a tangent-plane bias (metric), and Area-Aware Covisibility (AAC) applies a pre-sigmoid log-area correction (area). The chart-naive scaffold serves as a controlled reference, separating the scaffold-replacement effect from the spherical-prior effect. Instantiated in the RoMa V1 framework with the same frozen encoder, refiner architecture, and loss, SCCM also outperforms the ERP-native EDM (0.163) and an ERP-retrained RoMa V1 (0.198) under a unified ERP dense matching protocol, while perspective-trained matchers largely fail on ERP. It further transfers zero-shot to Stanford2D3D and, when trained on outdoor Holo360D, leads there as well.

cs.CV↗

Interactive-Policy Distillation with Bidirectional Propose-and-Verify

On-policy distillation (OPD) trains a student model on its self-generated trajectories with dense token-level teacher feedback. However, naive OPD may suffer from teacher unanchoring, where the student's reasoning trajectory drifts far from the teacher, causing the teacher to be queried on states it would hardly visit and thus provide unreliable supervision. We propose Interactive-Policy Distillation (IPD), which applies adaptive teacher intervention to the student rollout. Under a bidirectional propose-and-verify state machine, the student and teacher alternately exchange their roles as proposer and verifier, and collaboratively generate mixed-source trajectories. Then different supervisions are applied according to the source of each token. This bidirectional propose-and-verify mechanism and the source-split loss make IPD not only a more performant distillation method, but also a unified bridge between on-policy and off-policy paradigms. To make the interleaved dual-model rollouts more efficient, we also design a dedicated fused inference engine that co-hosts both models in one serving instance with separate KV caches and instantiates the state machine model to distribute, collect, and process requests. On math reasoning tasks and across multiple teacher-student model pairs, student models trained with IPD not only outperform those trained with OPD, but also demonstrate higher data efficiency. Specifically, when distilling Qwen3-30B-A3B into Qwen3-1.7B-Base, IPD brings a +3.28 mean@8 and a +3.28 best@8 benchmark-averaged accuracy improvement compared with OPD. Besides, IPD only consumes about 1/4 of the training examples and steps to outperform OPD trained on the whole training dataset for one epoch. We also investigate the impact of different loss variants and takeover / handback configurations, and demonstrate the robustness of IPD on different training data.

cs.LG↗

Single-Voxel Wireless NeRF for Spatial Spectrum Prediction

Wireless channel measurements across multiple spatial directions are crucial for AI-driven applications, such as RF digital twins and integrated communication and sensing. However, collecting channel data across large scenes is labor-intensive. Wireless NeRFs address this challenge by learning propagation behavior from sparse measurements and synthesizing channel spatial spectrum magnitude at unseen locations. However, existing wireless NeRFs inherit dense volumetric sampling from vision NeRFs, which requires substantial computation. This paper asks whether such dense sampling is necessary for predicting magnitudes of the wireless spatial spectrum. We empirically show that wireless NeRFs are over-parameterized for this task and introduce SV-INGP, a sparse volumetric sampling variant of Instant Neural Graphics Primitives (INGP). Across real-world and simulated datasets, SV-INGP matches the median Structural Similarity Index Measure (SSIM) of the NeRF2 baseline while reducing training time by 184x. These results generalize across LoS and NLoS scenes, sub-6 and millimeter-wave frequencies, and antenna array tapering configurations, suggesting a simpler and more efficient design path for RF digital twins

eess.SP↗

Classical Shadows of Higher-Form BRST Anomalies

We formulate a precise sense in which a BRST anomaly may possess a classical phase space shadow. It is the non-equivariance class of a Hamiltonian symmetry action on covariant phase space. To organize this statement, we introduce the Weil covariant phase space bicomplex, whose basic subcomplex controls equivariant Hamiltonian lifts. We then compare this purely classical class with local BRST descent. For a Hamiltonian admissible mixed transgression, the ghost-number-two descendant $a^{2}_{d-1}$ defines a deghostification map to the charge algebra, which is generated by the failure of the Cartan representative to be equivariant. The induced cohomology class is independent of Bardeen type changes of descent representative. Thus the classical anomaly is a shadow of the BRST anomaly in the precise sense that both are realizations of the same Weil transgression class, while remaining objects of different theories. As an explicit example, we analyze a five dimensional inflow transgression and show that the classical charge cocycle and the BRST descendant arise as two realizations of the same Weil transgression class.

hep-th↗

MedForge-RSI: Medical Deepfake Detection via Recursive Self-Improvement

Text-guided image editors can generate high-fidelity medical deepfakes, challenging the reliability of clinical imagery. Although reasoning-based detectors perform strongly in distribution, they degrade substantially under deployment shift. MedForge-Reasoner, an 8B vision-language model trained with supervised fine-tuning and reinforcement learning, achieves 99.2% accuracy on its target distribution, yet misclassifies 40% of authentic scans, reaches only 77% accuracy on unseen generators, and falls to 59% under transmission distortion. Adapting such models through conventional retraining is costly, requiring large-scale supervision and expert-designed guidelines. We introduce MedForge-RSI, a recursive self-improvement framework that enables a deployed detector to adapt while keeping its model weights frozen. Over 20 rounds, the detector analyzes verified errors, accumulates reusable experience, and autonomously develops image-analysis tools, while an independent acceptance test retains only validated improvements. Across 49 registered configurations, MedForge-RSI increases average accuracy over four test sets from 75.0% to 84.4%. On a held-out 4,000-image evaluation, it improves clean accuracy from 76.5% to 87.9% and transmission-distorted accuracy from 59.0% to 70.5%, with the largest gains in authentic-image recall. Controlled analysis across all 49 trajectories identifies which self-improvement mechanisms replicate across seeds, shows acceptance testing to be the largest individual contributor, and reveals a taxonomy of failed adaptations. We release the complete trajectories, including all rejected and rolled-back changes.

eess.IV↗

Grounded Revision vs. Prior Injection: Probing Retrieval-Augmented Patent Claim Amendment

Retrieval-augmented generation is widely used in professional writing, yet whether retrieval grounds revision or merely injects templates is rarely tested where "correct" has a definable meaning. Patent claim amendment supplies that signal: the examiner names the attacked limitation and cites prior art, providing per-case ground truth. We release three artifacts: (i) a corpus of 7,385 USPTO prosecution cases with XML-aligned pre/post claims, rejection, and cited prior art; (ii) a seven-probe battery comparing random and structural-match retrieval as two policies under a fixed prompt scaffold; (iii) a deterministic five-channel metric (C1-C3 and C5 in main, C4 supplementary) requiring no LLM evaluation. Across 9,600 pre-registered calls on four frontier LLMs (Claude Sonnet 4, Claude Haiku 4.5, GPT-5.4, GPT-4o-mini), no tested model exhibits detectable classical prior-injection behavior; retrieval effects are small and direction-inconsistent between random and structural retrieval, and the null is unchanged under a dense (semantic) retriever, across retrieval depths k in {1,3,5,10}, and under a paraphrase-sensitive grounding metric. Revision locality reveals a model-specific difference that the template channel misses. The four-cell taxonomy, which we treat as exploratory, leaves the prior-injector cell unoccupied.

cs.CL↗

Degree conditions for $k$-strong orientations of digraphs

Jackson and Thomassen conjectured that every $2k$-strong digraph contains a spanning $k$-strong oriented subdigraph. We prove sharp degree conditions for the existence of such a subdigraph. For every fixed positive integer $k$ and all sufficiently large $n$, every $n$-vertex digraph $D$ with $δ^0(D)\ge\lfloor(n+k-1)/2\rfloor$ admits a $k$-strong orientation. This threshold is best possible even for the weaker conclusion that $D$ itself is $k$-strong. We also prove a sharp Woodall-type analogue for every fixed positive integer $k$ and all sufficiently large $n$: if $d_D^+(x)+d_D^-(y)\ge n+2k-2$ for every missing arc $xy$, then $D$ admits a $k$-strong orientation, and this bound is again best possible. As a further consequence, we determine the sharp minimum total degree threshold. Finally, the semi-degree result also remains valid when $k\leαn$ for every fixed $0<α<0.094882\ldots$.

math.CO↗

SERA: Scale-Equalized Rollout Allocation for Maximum Likelihood Reinforcement Learning

Maximum Likelihood Reinforcement Learning (MaxRL) targets prompt-wise log-success and has shown strong performance on reasoning tasks. Under finite rollout budgets, however, the estimator used by MaxRL attenuates each prompt's likelihood gradient by a factor that depends on its success probability and rollout count. Under uniform rollout allocation, the common rollout count fails to compensate for success-dependent attenuation, leaving low-success prompts more strongly attenuated and distorting their relative contributions to the expected aggregate gradient. We introduce SERA (Scale-Equalized Rollout Allocation), which redistributes a fixed rollout budget to approximately equalize these finite-rollout scaling factors. Building on our theoretical analysis of how finite rollouts distort prompt-wise likelihood gradients, we formulate the allocation as a fixed-budget max--min problem, derive a waterline solution to its continuous relaxation, and introduce a multiplicity correction to remove the additional prompt weighting induced by heterogeneous rollout counts. Experiments show stronger alignment with exact likelihood gradients in a controlled ImageNet setting and improved multi-sample solution coverage over MaxRL on maze navigation and mathematical reasoning under matched training rollout budgets.

cs.LG↗

Learning to Explore Hidden Kinematics for Articulated Object Manipulation

The kinematics of an articulated object is often ambiguous from vision alone. Interaction resolves the ambiguity, and active perception methods exploit this by searching for the single action that most sharpens a belief over the kinematic parameters at each step. Such greedy search cannot be extended over a horizon without forward models of the contact and inertial dynamics, which are themselves unknown. We instead amortize action selection into training. We maintain a belief distribution over joint type and parameters, initialized from a generative prior and updated by Bayesian filtering on the observed part motion. To condition the policy on this belief, we render it as a per-point articulation flow field, the motion that the current posterior predicts for every point on the object. Carrying the inductive bias of articulated motion, this representation generalizes better than a latent encoding of the belief or flow tracked from observation. We train the policy with reinforcement learning, rewarding the entropy that each interaction removes from the posterior, so that informative exploration becomes learned behavior rather than a search at every step. Our method outperforms previous approaches across door and drawer manipulation on the PartManip benchmark, and reaches 61.7% success on ArticuRiddle, a new dataset of objects whose appearance implies the wrong articulation, against 44.4% for the best previous method. Project Website: https://hiddenkinematics.github.io/

cs.RO↗

Multicomponent anyons in one-dimensional optical lattices

We investigate the ground-state and dynamical properties of multicomponent interacting anyons confined in a one-dimensional (1D) optical lattice. Adopting the Anyon--Hubbard model, we explore spin and charge correlations as functions of interaction strength and the anyonic statistical parameter. Our results demonstrate that multicomponent effects combined with fractional exchange statistics substantially reshape both charge and spin correlations. For fractional exchange statistics, the canonical symmetric shell structure with prominent singularities in the spinor fermionic momentum distribution undergoes notable structural reconstruction, exhibiting emergent asymmetry, new spectral peaks, and broadened singular features. Unlike the Tonks--Girardeau bosonic gas, pseudobosons, representing a limiting case of anyons, cease to display a dominant zero-momentum peak. Instead, a quasi-fermionic shell structure emerges with finite-momentum peaks, whose locations are modulated by lattice site occupancy and interaction strength. The structure factor of spin correlations uncovers antiferromagnetic ordering in spinor anyon systems, with correlation magnitudes tunable by anyonic exchange statistics. Furthermore, quench dynamics analysis reveals the statistical phase parameter as an effective tuning knob. It enables the switching of dipole oscillations between underdamped and overdamped relaxation regimes and governs overall cloud expansion dynamics. Our findings pave the way for exploring statistics-driven ground-state and dynamical phase transitions in spinor anyon systems.

cond-mat.quant-gas↗