Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,747 records · Page 97Linked to original sources

MVP-SLAM: Multi-Camera Visual-Inertial Floorplan-Prior SLAM

Indoor building construction sites are demanding environments for visual SLAM, where variable lighting and repetitive, low-textured structures make the system drift over long trajectories, though structural elements such as walls remain distinguishable despite these conditions. These buildings are constructed according to their as-planned floor plans, available from the design phase, and although the actual as-built site can differ from this design, floor plans still provide a metric reference, both to localize the system in the building and to correct drift. Existing methods often use the floor plan to correct an already-built trajectory offline, and those that instead correct it online typically rely on depth sensors. We instead present MVP-SLAM, an online visual-inertial SLAM on two opposite-facing fisheye cameras that corrects drift from cameras alone by matching walls detected in its map to the floor plan, through a drift-aware policy. A multi-stage integration then turns each matched pair incrementally into a persistent correction, so the trajectory stays corrected and localized within the floor plan as it is built. MVP-SLAM was validated on the multi-floor construction sites of the Hilti-Trimble SLAM Challenge 2026, ranking 2nd of 22 teams in the Localization task (0.29 m mean RMSE) and 5th of 62 teams in the SLAM task (0.24 m), the top-ranked one in both tasks among those that operate online, integrate the floor plan, and localize within it.

cs.RO↗

Thermodynamics of Ahn--Doherty--Landahl Continuous Quantum Error Correction

Continuous quantum error correction (CQEC) replaces discrete syndrome measurements and recovery operations with continuous syndrome extraction and real-time Hamiltonian feedback. Here we investigate the thermodynamic resources required by this process by formulating measurement-based continuous quantum error correction as an information engine. We distinguish the system-side power associated with energy transfer between the feedback field and the protected system from the controller-side power associated with the feedback Hamiltonian. The framework is first developed for the one-qubit Ahn-Doherty-Landahl (ADL) protocol and subsequently extended to the three-qubit repetition code under continuous stabilizer monitoring. Numerical simulations show that increasing the feedback strength improves the steady-state fidelity and reduces the conditional-state entropy, while both energetic contributions continue to increase in magnitude after the fidelity begins to saturate. These results expose a direct tradeoff between logical stabilization and the energetic resources required for continuous feedback, closely related to previously established energy--precision tradeoffs in quantum measurement and quantum-Zeno stabilization.

quant-ph↗

The holographic QCD critical point is sensitive to quark flavors

We study the phase structure of dense QCD matter by applying the flavor-dependent holographic V-QCD model to two distinct physical environments: charge-neutral, beta-equilibrated matter, and strangeness-neutral matter with a fixed charge-to-baryon number ratio, n_Q/n_B = 0.4, relevant for heavy-ion collisions. The quark sector of the model incorporates the realistic mass hierarchy of two massless light quarks and a massive strange quark. We find that the resulting phase diagram depends sensitively on both the procedure used to tune the model against lattice QCD thermodynamics and the choice of physical environment. Using our preferred fitting procedure, we find that beta-equilibrated matter exhibits a first-order phase transition terminating at a critical endpoint located at significantly lower densities than in earlier holographic studies. However, this phase transition disappears entirely under heavy-ion conditions. This pronounced environmental dependence may help explain why recent net-proton cumulant measurements from Phase II of the RHIC Beam Energy Scan have not yet revealed clear non-monotonic fluctuation signatures. We calculate higher-order baryon number cumulant ratios along the chemical freeze-out curve to quantitatively compare the model predictions with experimental data. Finally, we also present global phase diagrams obtained by matching V-QCD to Hadron Resonance Gas models implemented in the Thermal-FIST package, and argue that this comparison further supports the conclusions drawn from the V-QCD model alone.

hep-ph↗

Text-to-3D Policy: Fine-Grained Language-Behavior Alignment for Unseen Specification Generalization

3D visuomotor policies provide a strong foundation for spatially precise manipulation, yet current text-to-3D policies struggle to follow unseen fine-grained behavioral specifications beyond those covered by demonstrations. We study this challenge as unseen specification generalization, where language specifies behaviorally significant variations, such as target position, displacement, or articulated state, that are absent from policy training. We find that pretrained language representations and conventional global behavior-language alignment capture coarse task semantics but often blur nearby specifications that require distinct behaviors. We introduce T3DP, a Text-to-3D Policy framework for fine-grained language-behavior alignment. Rather than compressing each instruction and demonstration into a single global embedding, T3DP preserves their local structures and establishes bidirectional token-level correspondence between linguistic elements and behavioral segments. This directly grounds subtle linguistic variations in the behavior components they affect, preventing closely related specifications from collapsing in the representation space. The resulting specification-sensitive language representation conditions a point-cloud-based 3D diffusion policy, enabling more precise control over unseen behavioral specifications without modifying the underlying policy architecture. Across Meta-World, ManiSkill, and RoboTwin, T3DP improves average held-out-specification success over global language-behavior alignment by +11.0-14.2 points, with gains on all 15 task families; on real-robot tasks, it further raises average success from 47.5% to 65.0% (+17.5 points). Representation and action-probe analyses show that fine-grained alignment better preserves specification geometry and action-relevant variation, linking local behavior grounding to downstream control.

cs.RO↗

GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed

Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply an intrinsic left-to-right generation order. This distinction makes bidirectional diffusion a natural fit, allowing spatial hypotheses to emerge in parallel and be jointly refined through iterative denoising. We introduce GroundAnything, a 4B-parameter grounding foundation model that reconciles fast parallel decoding with precise localization through blockwise denoising. Training combines grounding pretraining from public datasets and dedicated data engines, direct AR-to-diffusion conversion with joint AR and diffusion objectives, supervised fine-tuning, and GRPO-based reinforcement post-training. Across 30 grounding benchmarks, our autoregressive variant, GroundAnything-VLM, establishes a new overall state of the art among similarly sized models at 72.42%, remaining competitive with GPT-6 Astra (71.35%). With entropy-guided decoding, GroundAnything also surpasses the prior state of the art at this scale, averaging 61.75% versus 53.32% for the fast MTP-based LocateAnything model. We further explore decoding strategies, showing that an optional self-speculative mode achieves a $4.51\times$ speedup over the AR counterpart with a 0.74 percentage-point drop in COCO F1mIoU. Infrastructure experiments show that progressive inference optimizations translate parallel decoding into practical speedups. These support efficient visual grounding in latency-sensitive real-world systems.

cs.CV↗

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

Precise grounding matters. It specifies which object is the target and where that object is, even in clutter and for tiny objects, and it has to be fast enough for closed-loop control. Yet vision-language-action (VLA) and world-action models (WAMs) take perception from general-purpose vision-language and video-generation backbones, which still fail in these settings. We introduce GroundingPI, a 4B grounding foundation model that generates points and boxes as quantized coordinates in a shared vocabulary. Training combines multimodal and spatial pretraining, supervised fine-tuning, and reinforcement learning with GRPO, using supervision from public datasets and dedicated data engines. Against 44 baselines across 34 grounding benchmarks spanning 11 perceptual capabilities, GroundingPI establishes a new state of the art, averaging 73.68%, above the larger GPT-6 Astra (71.54%). As a downstream visual backbone, GroundingPI improves performance on robotic manipulation and autonomous driving. On RoboTwin 2.0, it outperforms every mainstream backbone we evaluate in all four out-of-distribution settings, by up to 24.8% relative to the strongest backbone. On RoboCasa-GR1, GroundingPI trained with 50% of the demonstrations outperforms those baselines trained with 75%. On nuScenes, used as the visual backbone, GroundingPI attains an average open-loop L2 error of 0.296 m. We systematically analyze GroundingPI's pretraining in scale and data composition. Downstream autonomous driving and robotic manipulation improve as the pretraining is scaled. Analyzing the data recipe across these 11 perceptual capabilities shows dense grounding's substantial benefits for both, and OCR's potential as a catalyst for perceptual learning. These results support grounding as a perceptual foundation, and dedicated perceptual pretraining as a promising direction for foundation models of physical intelligence.

cs.CV↗

Warm-Start Iterative QITE for Distribution Network Reconfiguration via Branch-Exchange Encoding

We present a hybrid quantum-classical algorithm for distribution network reconfiguration, a combinatorial optimization problem on power distribution networks, that combines a radiality preserving branch-exchange encoding with iterative warm-start quantum imaginary-time evolution to minimize active line losses. The encoding uses a fixed-width binary representation of sequential branch-exchange actions, ensuring that every register outcome decodes to a radial configuration. The quantum subroutine is informed by a surrogate model fit to classically-solved alternating current power flow (ACPF) labels. After a set budget of ACPF solves and circuit samples is exhausted, the algorithm returns the lowest-loss configuration seen. We employ our algorithm on ten test cases, including seven commonly-used benchmark cases and three cases we modified from these and similar standard systems. We include two methods to reduce problem size, which enable extensions of the algorithm to larger networks. We successfully reach minimal-loss configurations for six cases and find configurations with losses from 0.2 percent to 2.18 percent above the reported reference minimal loss for the other four.

quant-ph↗

An LSST-DESC Precursor Project: Hyper Suprime-Cam Year 1 $3\times2$pt in Harmonic Space

We present the first fully photometric joint analysis of weak gravitational lensing and galaxy clustering (3$\times$2pt) in harmonic space with LSST-DESC analysis pipelines applied to Hyper Suprime-Cam Year 1 (HSC Y1) data. The HSC Y1 dataset, with imaging depth and galaxy number density similar to those expected for the first year of LSST observations, is an ideal testbed for validating DESC measurement and inference tools. We measure the full set of angular power spectra - cosmic shear, galaxy clustering, and galaxy-galaxy lensing - perform null tests, and correct for the impact of observing conditions on both the galaxy density and shear fields via mode deprojection. We validate our likelihood pipeline by reproducing the official HSC Y1 cosmic shear cosmological constraints with the DESC inference code. Jointly analysing all three two-point functions, we constrain $Λ$CDM and $w$CDM cosmologies, reporting results for cosmic shear, the combination of galaxy clustering and galaxy-galaxy lensing (2$\times$2pt), and the full 3$\times$2pt. From the joint 3$\times$2pt $Λ$CDM analysis we find $Ω_m = 0.249^{+0.061}_{-0.051}$, $σ_8 = 0.904^{+0.098}_{-0.091}$, and $S_8 \equiv σ_8\sqrt{Ω_m/0.3} = 0.821^{+0.018}_{-0.021}$. This analysis validates the complete DESC 3x2pt infrastructure on survey data, providing the foundation for the forthcoming LSST Year 1 cosmological analysis.

astro-ph.CO↗

Why Do Conventional World Models Fail to Learn Cellular Automata?

Although conventional world models - auto-regressive or diffusion models based on transformers or convolutional networks - may learn surface statistics of world dynamics, can they learn the exact world dynamics from its observed history? Leveraging cellular automata as a simple testbed, we find the answer to be no in many cases. Conventional architectures predict most pixels correctly yet rarely complete a rollout: a CNN predicts 96.3% of cells but completes 18.9% of rollouts; a joint diffusion model completes none. We trace the gap to three failure modes of these world models - namely, they fail to exactly capture spatial locality, temporal locality or temporal stability. Simple changes repair each: (1) for spatial locality, two-dimensional rotary positions lift a transformer from 39.1% to 100% on the Game of Life; (2) for temporal locality, handing each token its cell's previous-frame neighbourhood lifts the same transformer from 25.8% to 99.9% on unseen rules; (3) for temporal stability, causal freezing lifts the same diffusion weights from 42.2% to 99.9%. None of the three changes touches the architectural backbone; each only modifies the information flow within it. We also compare joint and ordered sampling on billiards and, in an exploratory study, on a simulated Burgers equation.

cs.AI↗

FOMO: Forget the Concept, Don't Miss Out on the Scene in Selective Video Unlearning

The rapid advancement of generative video models has enabled the synthesis of increasingly realistic and temporally coherent videos, while also raising concerns about the generation of harmful content. The reliance on large-scale web datasets during training inevitably exposes these models to undesirable material, making concept unlearning an essential mitigation. Existing methods mainly target static visual concepts, such as objects, identities, or unsafe appearance, largely overlooking motion unlearning. Furthermore, these approaches often pay little attention to preserving the surrounding scene. As a result, successful concept removal may unintentionally alter the background, composition, or overall video dynamics. We argue that effective unlearning should ideally change only what is targeted, while minimizing unnecessary changes to the remaining scene. In this work, we introduce FOMO, to the best of our knowledge the first training-based selective video unlearning method that directly treats preservation of the original scene as a priority. We formulate unlearning around two complementary objectives: what to change and what to preserve. Our method localizes concept-related representations and modifies them, while the preservation mechanism maintains non-target scene information without requiring auxiliary data. Beyond simply erasing unwanted concepts, FOMO explicitly redirects the generation toward a specified safe alternative. We further extend this formulation to motion unlearning, where the concept is defined by temporal behavior rather than a fixed spatial region. Our solution achieves effective unlearning across unsafe content, object, and motion concepts, while achieving the best trade-off between concept removal and scene preservation. Code: https://github.com/gmum/FOMO Project Page https://gmum.github.io/FOMO

cs.CV↗

Regularity of Critical Points of Scale-Invariant Geometric Energies for Even-Dimensional Submanifolds of $\mathbb R^m$

We consider scale-invariant curvature energies for immersions of closed manifolds of even dimension $n=2h$ into $\mathbb R^m$, with principal term $\int_Σ \big|\nabla^{(h-1)} \vec{\mathrm{I\!I}}\big|_g^2\,d\text{vol}_g$ and arbitrary lower-order polynomial extrinsic invariants of the same scaling. Following the four-dimensional approach developed in joint work with Bernard, Martino, and Rivière, we prove that every weak critical immersion in the natural Sobolev class $W^{h+1,2}$, whose induced metric and its inverse have $L^\infty$ coefficients, is real-analytic in harmonic coordinates. The proof combines geometric conservation laws, additional structural identities, and elliptic estimates with critical Sobolev coefficients to obtain Morrey decay and bootstrap to full regularity.

math.AP↗

Pretext: Defeating Malicious Skill Detection Frameworks for AI Agents

Skills extend an agent's capabilities by injecting instructions and information into the context, and are widely used by agents such as OpenClaw and Claude Code. Prior work shows third-party marketplaces host malicious skills that give attackers direct influence over the victim's agent. The emerging defense scans skills before installation, pairing deterministic static checks with an LLM-based semantic judge, as in NVIDIA's SkillSpector. We show that such defenses fall to an attacker who knows the detector. Our white-box LLM attacker, Pretext, iteratively crafts skills that evade detection while still delivering the payload and performing the benign task: moving the payload from code into natural language leaves static analysis inert, while framing it as the skill's legitimate purpose and splitting instructions across files keeps the LLM stage below its blocking threshold. Across three open-source models, Pretext achieves up to 97\% and 77\% against a frozen detector and a co-adaptive one, respectively, revealing major gaps in current skill scanners.

cs.CR↗

Is This Evidence Decision-Critical? Learning to Verify Rule-Governed Decisions

Rule-based reasoning, as in eligibility checks and contract reviews, requires language models to assess evidence against individual conditions and combine their judgments under explicit rules. Errors in evidence assessment can leave a decision unchanged, but misinterpreting or overlooking decision-critical evidence can reverse it. Identifying such evidence allows more capable models to focus on checking the corresponding condition judgments, supporting accurate and safe decisions. Recognizing the evidence's criticality requires understanding how evidence affects a condition judgment and how that judgment affects the decision. To achieve the goal, we propose a INTERvention-based imPACT learning framework (InterPact), which enables counterfactual verification of evidence criticality in rule-governed decisions. Specifically, its evidence intervention constructor generates training pairs for a propagation verifier by editing case facts with a frozen language model while holding rules and non-target conditions fixed. Human-reviewed labels record the resulting condition and decision changes, while complete state-to-decision mappings supervise consequences beyond the observed edit. During training, the verifier weights learned conditional decision predictions by evidence-based condition probabilities through a fixed composition operation, propagating decision-change supervision into the base model. At inference, the trained base model directly judges criticality from the original case and target evidence, without human or stronger-model supervision. On single-case evidence criticality verification over adapted rule-governed decision cases, InterPact achieves 68.28% accuracy, outperforming all six baselines. These results support learned decision sensitivity as a basis for prioritizing evidence checks.

cs.CL↗

Complete total-transmission modes of Kerr black holes

We construct the complete spectrum of gravitational total-transmission modes (TTMs) of Kerr black holes and find four globally continuous families, $n_\infty=1,2,3,4$. Compared with the three-family classification of Cook and Lu [Phys. Rev. D 107, 044043 (2023)], our global continuation separates complex conjugation at fixed $m$ from the mirror symmetry connecting the $m$ and $-m$ spectra, yielding a uniform four-family classification without additional symmetry-unrelated roots. The $n_\infty=3$ family approaches the Schwarzschild algebraically special frequency, whereas the $n_\infty=1,2,$ and $4$ families diverge as $ω\propto a^{-4/3}$ along lower-half-plane directions $-150^\circ$, $-90^\circ$, and $-30^\circ$. High-precision data up to $\ell=32$ recover the Cook--Lu small-spin asymptotics and show how the divergent branches are embedded in the global four-family spectrum. The spherical-limit polar labeling and large-$|aω|$ angular ordering are related in a family-dependent manner. For axisymmetric perturbations, we show that the $n_\infty=2$ and $n_\infty=3$ branches coalesce at exceptional points and subsequently evolve toward distinct small-spin limits, rather than exhibiting the overtone-multiplet splitting proposed by Cook and Lu. Finally, we identify a previously unreported anomalous proximity between the $n_\infty=3$ TTM family and the unconventional Kerr quasinormal-mode sequence, with separations reaching order $10^{-8}$ while remaining numerically well resolved.

gr-qc↗

Circuit-Based Dispersion Analysis of Periodic Cross-Shaped Unit Cells with Dirac Characteristics

In this paper, the dispersion behavior of periodic cross-shaped unit cells is investigated with emphasis on Dirac-type dispersion using an equivalent circuit modeling approach. Two unit-cell geometries, including a simple cross structure and a modified configuration with a central patch, are analyzed using full-wave eigenmode simulations and scattering-parameter-based dispersion extraction. An equivalent circuit model is developed to capture the dominant coupling mechanisms and accurately reproduce the dispersion characteristics near the $Γ$-point. The results demonstrate that the proposed circuit model exhibits good agreement with full-wave simulations and provides a physically intuitive framework for analyzing stopband behavior and mode degeneracy. Furthermore, the limitations of S-parameter-based dispersion extraction in the presence of low-Q radiative modes are highlighted.

eess.SY↗

Towards Agile Vision-Based Multi-UAV Flight: Revisiting State Estimation

Agile multi-UAV flight requires accurate and low-latency onboard estimation of the kinematic states of neighboring UAVs for collision avoidance, motion coordination, etc. Most vision-based approaches rely on position-only measurements, inferring velocity and acceleration indirectly from displacement. We show that this introduces a fixed structural delay in the estimation of higher-order states, which limits the achievable agility. To address this, we propose to integrate tilt measurements, provided by a state-of-the-art visual detector, which inform about the thrust direction of co-planar multirotor UAVs. We benchmark four position-only and five pose-aware estimators, including a novel formulation of a linear thrust-constraining Kalman filter, on two real-world and one high-fidelity photorealistic simulated dataset over different levels of agility (3-21 m/s^2). In our setup, pose-aware estimation consistently reduces the average velocity and acceleration estimation errors by 40% and 57% across the three datasets with the proposed KF formulation outperforming the other estimators. Position-only filters exhibit a constant ~300 ms delay in acceleration step response independent of agility, whereas the tilt-constrained estimators operate near the physical response limit given by the camera frame-rate by observing the change in thrust direction before the displacement accumulates. In a closed-loop leader-follower simulated experiment with NMPC control, position-only estimation of the leader's state fails to facilitate stable hovering of the follower, while the proposed estimator enables tracking of lateral maneuvers exceeding 2g of acceleration.

cs.RO↗

Anisotropic wavevector-dependent damping of thickness-quantized magnons

Magnon damping is a key factor governing spin-wave transport and nonlinear dynamics of multimode magnon systems. However, many descriptions rely on the assumption of an effective mode-independent parameter, which can mask wavelength-dependent relaxation processes that depend on propagation geometry and mode profile. Here, we employ high-resolution parametric-instability spectroscopy to probe thickness-quantized spin waves with wavelengths down to about a hundred nanometers in micrometer-thick yttrium iron garnet films. The instability threshold exhibits a regular sawtooth dependence on the magnetic field, arising from switching between discrete thickness modes. Comparison with dipole--exchange theory reveals anisotropic wavevector-dependent damping that increases with mode number and depends differently on the in-plane and out-of-plane wavevector components.

cond-mat.mes-hall↗

Hybrid Methods for Robust Tabular Data Imputation

Missing data are a fundamental challenge in statistical analysis and machine learning, as the choice of imputation method substantially impacts downstream inference. In this work, we propose two hybrid imputation methods called NuclearForest and SoftForest, which combine nuclear-norm-based low-rank initialization using Singular Value Thresholding (SVT) and SoftImpute, respectively, with a non-iterative Random Forest refinement. For the SVT-based component, we further introduce an adaptive step-size rule, prove adaptive step-size bounds, and establish convergence for the corresponding zero-initialized iteration. The low-rank initialization provides a structured warm start that captures the global covariance patterns in the data, while the subsequent Random Forest step recovers residual nonlinear signals encoding local dependencies. We conduct an extensive benchmark on diverse datasets from different application domains, comparing the proposed methods with seven established imputation methods under the Missing Completely at Random (MCAR), Missing at Random (MAR), and Missing Not at Random (MNAR) mechanisms across varying missingness rates. Our results demonstrate that NuclearForest and SoftForest match or exceed the imputation fidelity of state-of-the-art iterative methods such as MissForest, while significantly reducing computational cost. In particular, they achieve speedups of approximately 5.81 times and 9.52 times over MissForest by replacing iterative cycles with a single refinement step. Our approach effectively exploits the low-rank structure of real-world tabular data and accommodates mixed-type variables, providing an efficient and robust solution for data imputation in bioinformatics, economics, and beyond.

cs.LG↗