Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,693 records · Page 94Linked to original sources

Comparative study of adapting pre-trained models for driving behavior video captioning

This report examines and compares some of the many fine tuning and prompting methods existing, applying them within the domain of autonomous driving. The idea is to compare these methods by adapting a Large Language Model (LLM) on a video dataset. LLM's have become extremely good at achieving a good understanding of different forms of data and this study aims to induce a low dimensional understanding of driving situations into our primary test model SpaceTimeGPT. Experiments on BDD-X (Berkeley DeepDrive eXplanation) dataset demonstrate good performance of the full fine tuning framework on some automatic metrics, and in some metrics, it even surpasses the baseline. We also try Low-Rank Adaptation (LoRA) and prompt engineering on VideoLLaVA model and discuss its limitations.

cs.CV↗

Phenomenology of dynamical trapped regions

Regular black holes (RBHs) and horizonless black hole mimickers (BHMs) are often studied as stationary alternatives to classical black holes, although their interiors may be unstable or undergoing relaxation. We investigate the phenomenological consequences of this evolution by numerically evolving linear scalar perturbations on prescribed, time-dependent Hayward-like geometries. We consider BHM--RBH--BHM histories in which an initially horizonless object temporarily develops a trapped region before returning to a horizonless configuration, through either a single bounce or a sequence of repeated bounces. Using horizon-penetrating Painlevé--Gullstrand coordinates, we follow the perturbations through the formation and disappearance of trapping horizons and compute the resulting waveforms and scalar-field energy. For histories with identical initial and final geometries, we find that changing the duration of the intermediate regular black hole phase produces differences in the amplitude and phase of the late-time signal, including its echoes. Longer trapped phases also yield greater energy amplification, consistent with the blueshift of outgoing modes near the inner trapping horizon. These results show that the late-time response of a dynamically relaxing BHM need not be determined by its final configuration alone: it can retain a memory of the dynamical history of its interior.

gr-qc↗

Growing an Agent/Prover Interface: Evolutionary Tool Design for Cost-Efficient Theorem Proving in Rocq and Lean

Recent achievements in AI-assisted mathematics require intensive interaction of agents with proof assistants to generate machine-checked proof certificates. Agents interact with proof assistants such as Rocq or Lean through an interface that controls what the agent receives from the prover and the cost of these interactions. Today, these interfaces are adapted from tools designed for humans and not optimized for agents. We propose an evolutionary method where a frontier model incrementally proposes new features and only keeps the ones that improve the overall performance of smaller models. We demonstrate the effectiveness of our method by growing, on a curated set of mathematical problems, \rme, a new MCP server for the Rocq prover. On the held-out \texttt{test} split of miniF2F-Rocq, an agent equipped with \rme outperforms both the baseline that only exposes the Rocq compiler and an established MCP server, across four models from two families, in success rate, cost per solve, and time per solve. Although evolved for Rocq, the resulting server transfers to Lean, improving cost and time per solve on a subset of PutnamBench. We release \rme and its port to Lean.

cs.AI↗

Stability for coupled second order evolution equations with indirect general memory-dampings without the equal-wave-speeds-type hypothesis

We study the stability for a system of coupled second order evolution equations with indirect general memory-damping without the equal-wave-speeds-type hypothesis in a Hilbert space, where the damping only appears just in one equation, the memory kernel can not necessarily be nonnegative and nonincreasing, and the ``wave speeds" implied by the system can be different. Taking advantage of new processing ideas and with the help of the properties of the Generalized Positive Definite Kernel, we overcome difficulties caused by the different ``wave speeds", the lack of the decreasing and nonnegative property for the memory kernel and the system has only one equation with damping, and obtain an optimal polynomial stability result for the energy, which covers the previous related polynomial stability results for second order coupled equations (abstract or concrete) in the literature. Moreover, applications of the abstract result are given.

math.AP↗

FAIR-Compliant Architecture for Heterogeneous Astronomical Data: KazVO Framework

This paper presents the architecture of the Kazakhstani National Virtual Observatory (KazVO) - an International Virtual Observatory Alliance (IVOA) compliant node that unifies the heterogeneous observational datasets of the Fesenkov Astrophysical Institute (FAI). The architectural foundation of the system is built upon a partitioned Data Lake, a German Astrophysical Virtual Observatory (GAVO) Data Center Helper Suite (DaCHS) publishing backend, and a PostgreSQL database that maps heterogeneous metadata to the unified IVOA ObsCore standard. We follow Open Science by introducing metadata-only model featuring based on a four-class data embargo mechanism. Machine-to-machine services deployed via IVOA protocols are listed in the global Registry of Registers, enabling analysis of Kazakhstani observational assets within external clients such as Tool for OPerations on Catalogues And Tables (TOPCAT), Aladin, and PyVO. As a result, KazVO framework provides a scalable platform driven by a developed end-to-end pipeline that bridges two fundamentally distinct data types - digitized historical glass-plate heritage (1950-1997) and live operational photometric and spectroscopic digital streams from telescopes at the Assy-Turgen and Tien-Shan Observatories - opening FAI's combined data datasets to the global scientific and time-domain astrophysics community.

astro-ph.IM↗

Learning Reliable GUI Agents under Imperfect Priors

GUI agents built on large language and vision-language models still struggle on unseen applications and complex multi-step tasks, as completing real GUI tasks depends on app-specific, temporally volatile operational knowledge that is scarce in pretraining corpora. Retrieval-augmented execution offers a natural remedy but faces two coupled bottlenecks: knowledge at scale is hard to acquire, and self-collected priors inevitably drift from the live environment due to version updates, promotions, ads, A/B tests, and personalization. We therefore argue that GUI agents should not pursue perfect knowledge but learn to act correctly under imperfect priors, and propose our framework that couples knowledge acquisition with noise-robust utilization: a structured exploration strategy traverses interactive elements, builds a UI state-transition graph, and synthesizes (task, trajectory) pairs via a VLM without human annotation; a noise-aware training strategy, grounded in a taxonomy of real GUI drift patterns, injects five types of realistic errors into self-explored trajectories to teach the agent to assess prior reliability before acting. Experiments on physical devices and online emulator benchmarks show that our method discovers more unique screens, covers more benchmark tasks, and more effectively rejects erroneous priors while leveraging correct ones, with accuracy gains that transfer across datasets.

cs.LG↗

Learning Normal Diffusion Dynamics for Backdoor Defense in Text-to-Image Models

Backdoor attacks pose a serious threat to the secure deployment of text-to-image (T2I) diffusion models. Existing defenses typically detect backdoors from specific abnormal patterns in internal representations, which may limit their generalizability with the emergence of increasingly diverse attack mechanisms. In this paper, we study backdoor defense of T2I diffusion models from a transition-dynamics perspective. We observe that benign diffusion trajectories exhibit structured and timestep-dependent transition patterns from cross-attention, latent and noise spaces, whereas backdoor attacks tend to induce deviations from such normal evolution. Motivated by these observations, we propose Normal Diffusion Dynamics Learning (NDDL), a novel backdoor defense framework that learns the normal transition dynamics of diffusion trajectories utilizing only benign samples. NDDL constructs compact multi-space trajectory representations and trains a timestep-conditioned dynamics model to predict the diffusion evolution. In the inference phase, deviations between the observed and predicted transitions are exploited to quantify dynamics inconsistency for backdoor detection. NDDL further enables trigger localization without any prior knowledge of the embedded backdoor by performing substitution with low-semantic words. Extensive experiments for diverse backdoor attacks demonstrate the effectiveness and generalizability of our proposed NDDL.

cs.CV↗

Speculative Safety Honeypot: Toward Proactive Defense Against Multi-turn Agent Attacks

As Large Language Model (LLM) agents are increasingly deployed in complex environments, multi-turn interaction attacks have become a significant security challenge. Existing detection methods typically rely on historical context. However, this retrospective logic struggles to identify deep malicious intents that are split across turns to hide future risks. Inspired by speculative decoding, we propose the Speculative Safety Honeypot (SSH) framework. SSH uses a multi-agent simulation system composed of small LLMs to build an action-level speculate-and-verify workflow. In the speculation stage, SSH predicts future behaviors of the target agent and asynchronously builds a trajectory tree to expose potential risks in advance. In the verification stage, the system uses the target agent's real actions to calibrate and prune the trajectory tree, effectively reducing false positives. As a plug-and-playable component, SSH provides existing detectors with rich decision redundancy beyond the current interaction slice. By judging risk based on the evolution of the entire trajectory tree rather than a single point in time, the system reduces the reliance on the absolute precision of individual detection components. This improves the defense resilience and the warning lead-time of agent systems against complex temporal attacks.

cs.CR↗

Hyperbolic Prototype Routing for Rehearsal-Free Class-Incremental Learning

Class-Incremental Learning (CIL) aims to continually learn new classes while preserving prior knowledge. Parameter-efficient fine-tuning with pre-trained models enables CIL with minimal parameter updates, but existing approaches still suffer from catastrophic forgetting caused by cumulative interference and suboptimal module-sample matching at inference. We propose Hyperbolic Prototype Routing (HyPro), a rehearsal-free framework for continual learning. HyPro allocates a dedicated LoRA-Expert module to each incremental task for isolated representation learning, then projects routing features onto a Poincare ball and performs geodesic nearest-prototype matching for reliable task-level discrimination. Extensive experiments on standard CIL and Few-Shot CIL benchmarks show that HyPro consistently improves average and final-stage accuracy over strong baselines.

cs.LG↗

RankEvolve: A Reliable Multi-Agent Auto-Research Harness for Evolving Ranking Models

Auto-research agents, LLM systems that propose, implement, train, and evaluate model changes across iterations, promise to automate applied ML's experimental loop. Over long horizons, execution accuracy is a binding constraint: a change can silently leak held-out data, omit normalization, disconnect a gradient, or leave a train/eval flag unwired, invalidating expensive runs and compounding error across iterations. We present RankEvolve, an auto-research framework for evolving generative ranking models. An Executable Operating Protocol (EOP) declares phases, gates, branches, and loops, and the runtime enforces the compiled state machine. A meta-meta-harness composes complete black-box coding-agent products, including Claude Code and Codex, as execution-graph nodes that review and repair one another's work. In a budget-matched evaluation, heterogeneous composition raises all-oracle execution accuracy from the best single-product baseline of 45.8 percent to 62.5 percent (paired +16.7 points, 95 percent CI [6.6, 26.7]) while achieving a 10.4 percent silent critical-defect rate. An implemented knowledge layer carries findings, including negative results, across iterations. In a twelve-iteration deployment on the open-source HSTU recommender, RankEvolve reported NDCG@10 of 0.2192 on MovieLens-20M LARGE (+4.48 percent over the published anchor) and 0.1948 on BASE (+2.80 percent). ExecML-HSTU, seeded by incidents from that deployment, provides the oracle benchmark for the execution-accuracy evaluation. A pre-specified LitGPT transfer split replicates the heterogeneous-composition effect beyond recommendation (+12.5 points, 95 percent CI [3.0, 22.0]), and a paired ablation isolates per-step from full-protocol instruction injection. These results characterize when runtime-controlled composition of coding-agent products improves execution accuracy.

cs.AI↗

Ghost in the Encoder: Decodable Artist Identity Representations in Lyrics-to-Song Generation

Text-to-song generation models can be prompted to imitate specific artists or regurgitate entire songs from their training data. Although these phenomena have been documented behaviorally on small datasets, little is known about the internal representations that may give rise to them. Prior interpretability work on generative audio has focused on locating semantic concepts such as genre or time signature within model activations. In this work, we show that a trained model can be probed for linearly decodable representations of artist identity from song lyrics alone, without any additional identifiers. Through a controlled case study of ACE-Step 1.5 spanning 2,000 songs across 100 artists, we demonstrate that the artist associated with a given set of lyrics can be identified within the model's internal activations, and that this conditioning signal propagates from the lyric encoder to the diffusion backbone during inference. These findings indicate that lyrics constitute an artist-level conditioning channel not addressed by prompt-side replication safeguards. More broadly, our work highlights how latent-space analysis can be used to audit what generative music models have implicitly learned from their training data.

cs.SD↗

EffGS: Efficient and High-Fidelity Gaussian Splatting

3D Gaussian Splatting (3DGS) enables real-time novel view synthesis, but existing general-purpose acceleration methods suffer severe rendering quality degradation when extended to more complex, large-scale scenes. To address this issue, we propose EffGS, a more general acceleration framework that improves training and rendering efficiency while maintaining reconstruction quality comparable to or better than vanilla 3DGS across bounded and large-scale scenes. EffGS combines frequency-aware guidance, localized density control, and adaptive primitive scale modulation. First, an importance scoring mechanism combines pixel-wise reconstruction errors with a difference-of-Gaussians mask scheduled over training to provide stage-dependent spatial guidance. Second, localized densification and pruning restricts density modifications to Gaussians with valid projected footprints in the sampled views. Third, learnable per-Gaussian scale modulation adjusts effective primitive extent during optimization while retaining the Compact Box rasterization rule. Extensive experiments on bounded and large-scale scene datasets demonstrate a favorable balance between reconstruction quality, training time, and primitive count. Component ablations and matched-primitive-budget comparisons further support the effectiveness of the framework.

cs.CV↗

A Thomas-Yau-Joyce result for Lagrangian spheres in K3 surfaces

We consider objects in the Fukaya category of a class of K3 surfaces defined by certain Lagrangian spheres. Assuming that homological mirror symmetry for K3 surfaces holds with sufficiently strong (expected) properties, we prove that, in this case, stability with respect to a suitable Bridgeland stability condition implies the existence of an isomorphic special Lagrangian sphere, as predicted by the general Thomas-Yau-Joyce conjectures. In particular this holds unconditionally for suitable quartic surfaces in $\mathbb{P}^3$ or sextics in $\mathbb{P}(3,1,1,1)$ and their mirrors. A variant holds on the Calabi-Yau threefolds obtained by taking the product of our K3 surfaces with an elliptic curve. These seem to be the first results of this type on compact manifolds.

math.DG↗

Judd-Ofelt analysis of holmium-doped alumino-silicate optical glass prepared by MCVD combined with nanoparticle doping

We present a detailed Judd--Ofelt (JO) analysis of Ho$^{3+}$-doped alumino-silicate optical fiber preforms in a wide range of compositions and prepared by both the standard solution doping method as well as the more advanced nanoparticle doping. The absorption spectra were measured, the absorption cross sections were calculated for each observed transition, and the Judd--Ofelt analysis was conducted. The preforms exhibited the values of JO parameters of $Ω_2 = (7.2--10.6) \times 10^{-20}\,\mathrm{cm}^2$, depending on the Al$_2$O$_3$ content, and $Ω_4$ and $Ω_6$ around $2.5$ and $1.2 \times 10^{-20}\,\mathrm{cm}^2$, respectively. The radiative transition parameters, such as transition probabilities, branching ratios, and radiative lifetimes, were calculated. The radiative lifetime of the ${}^5I_7 \rightarrow {}^5I_8$ transition, in the $16.1--17.2\,\mathrm{ms}$ range, combined with a measured value in the $0.9--1.9\,\mathrm{ms}$ range, was used to calculate the quantum efficiency, which was in the $5--11\,\%$ range. The preforms prepared by nanoparticle doping containing high Al$_2$O$_3$ contents above $8\,\mathrm{mol.}\,\%$ exhibited superior values of measured lifetime (above $1.5\,\mathrm{ms}$) and quantum efficiency (above $8\,\%$) compared to the standard solution-doped samples. To the best of our knowledge, this work is the first comprehensive report on the JO analysis of Ho$^{3+}$-doped alumino-silicate glass. The calculated parameters may be used in various calculations, simulations, and modelling of holmium-doped fiber lasers and other devices.

physics.optics↗

Exact Maximum Likelihood Decoding beyond Treewidth via Rank-Decomposition Dynamic Programming

Maximum-likelihood decoding provides an optimal decoding strategy for quantum error correction under stochastic Pauli noise. However, computing logical-class probabilities is challenging, and the leading exact tensor-network contraction requires time exponential in treewidth. In this work, we introduce a new decoding algorithm based on rank-decomposition dynamic programming (Rank DP). We express decoding partition functions with independent local fault factors as quadratic sums of powers and apply Rank DP. The resulting exact evaluator has arithmetic complexity polynomial in the input size and exponential in the Tanner graph's rank-width with respect to the fault partition, including decomposition construction. After Gaussian elimination, it gives polynomial-time decoding for punctured quantum Reed-Muller codes and a Steane-concatenated family with growing distance under independent single-qubit Pauli noise. Standard contraction of the corresponding tensor networks requires superpolynomial time. A nonnegative realization gives relative floating-point error bounds under explicit arithmetic assumptions. Numerical experiments demonstrate runtime advantages over the tested tensor-network implementations on selected code-capacity and circuit-level instances, with full likelihood evaluation for Reed-Muller codes up to $1{,}023$ qubits. We further use these likelihoods to learn circuit noise parameters from syndromes, evaluate rare postselection probabilities, and quantify decoder optimality gaps. Our work opens new avenues for exploiting algebraic structure in quantum decoding and noise characterization.

quant-ph↗

Consensus for Compressed Static Functions

The Consensus technique marked a breakthrough in the construction of minimal perfect hash functions (MPHFs), reaching a linear tradeoff between construction time and space overhead relative to the optimum. Consensus provides a clever scheme to search for and encode seeds of tasks in random data structures. We apply Consensus to the related field of compressed static functions (CSFs). These data structures store a function $f: S \to Σ$ such that querying a key $x \in S$ returns $f(x)$ and querying $x \not \in S$ returns an arbitrary value. CSFs do not need to store the keys $S$ and only need space close to the zeroth-order empirical entropy of the multiset of values. Often, some values are much more common than others. In these cases, CSFs can use less space than their non-compressed counterparts. CSFs are a useful building block, for example in database design and bioinformatics. We introduce Consensus-CSF, which can reach arbitrarily close to the empirical entropy $n H_0$, with a construction time of $n \exp(\tilde{\cal{O}} (\sqrt{1 / δ}))$ for space usage of $n H_0 (1 + δ)$ when assuming some parameters of the value distribution to be constants. This tradeoff beats previously implemented approaches that can only reach some fixed threshold above the entropy lower bound. We enable Consensus in the setting of CSFs, which is less structured than MPHFs, with the introduction of task insertions. Our approach randomly distributes the keys into one-bit Consensus tasks and then strategically inserts additional tasks in places where the construction would get stuck otherwise. We provide an implemented version of our algorithm which reaches the same order of magnitude in space overhead as competitors but is not competitive in practice. Beyond these results, we present a new way to think and reason about Consensus, which may also be applied to other problems.

cs.DS↗

General Performance Guarantee for Human Torque Estimation-Based Task-Agnostic Assistive Exoskeleton Control

Accurate human torque estimation is crucial for enabling task-agnostic control in robotic exoskeleton systems. However, estimation errors may cause mismatches between the robot assistance and the human intention, degrading controllability and task performance. In this paper, we address this issue by formally defining matched assistance as scenarios in which the robot positively contributes to human movement. Based on this definition, we develop a theoretical framework to design the robot's desired interaction torque that guarantees a lower bound on the matched assistance probability. Importantly, the proposed guarantee holds over the entire torque distribution, including unseen data beyond the training tasks. This provides our method with strong reliability and generalization, both of which are critical for effective exoskeleton control. The proposed strategy is implemented on the ABLE upper-limb exoskeleton and evaluated in a multi-task setup. Experimental results validate the theoretical guarantees and demonstrate that the proposed strategy achieves effective general performance across several tasks, guaranteeing movement smoothness while reducing human physical effort.

cs.RO↗

Divide and Collapse: MAPF-Collapse via Exact Decomposition into Independent Sub-Instances

In this work we study the problem of MAPFC, a post-optimization step for Multi-Agent Path Finding (MAPF) plans where we are given a feasible plan produced by a modern MAPF solver and are tasked with removing avoidable moves while preserving feasibility. This NP-hard problem naturally arises when using learning-based state-of-the-art (SOTA) solvers which construct plans that contain redundant moves that can be removed. Recently, Tang et al. presented Judgelight, which uses Integer Linear Programming (ILP) to solve MAPFC. Importantly, the ILP is constructed over all agents jointly, so its cost is governed by the full instance rather than by the small coupled residue that actually requires joint reasoning. Our key insight, motivating this work, is that MAPFC instances naturally decompose into independent sub-problems, most of which involve a single agent and can be solved without any inter-agent reasoning. To this end, we first identify which agents need to coordinate their motion and partition the instance into sub-problems accordingly. For the cases where no coordination is required, we introduce an extremely lightweight solver that is $\approx\!1{,}900\times$ faster than Judgelight. For cases where coordination is required, Judgelight can be used but we introduce an alternative CBS-like solver which is more efficient on easier problems. The resulting framework is exact, uses no commercial ILP solver, and matches Judgelight's quality while running substantially faster on the coordination-light majority of instances; on the coordination-heavy instances we propose a regime-aware hybrid planner that falls back to Judgelight. Over all benchmarks tested, this planner achieves a median $10.5\times$ per-instance speedup over Judgelight.

cs.AI↗