Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,027 records · Page 57Linked to original sources

Generalized Effective Spin-Chain formalism for multicomponent anyons in one-dimensional optical lattices

We develop a generalized effective spin-chain (GESC) formalism for strongly interacting multicomponent anyons in a one-dimensional (1D) optical lattice. By mapping particle motion onto spinless fermions and spin states onto an ordered chain, this framework provides a spin--charge-separated perspective on the physical effects of fractional exchange statistics. In the strong-interaction regime, the leading-order charge Hamiltonian becomes independent of the statistical phase; instead, virtual tunneling through doubly occupied states directly imprints this phase on the spin-exchange coefficients. The GESC formalism captures ground-state properties of the full Anyon--Hubbard model reported in S.~Basak, X.-W.~Guan, and H.~Pu, unpublished manuscript (2026), and reveals that their statistical dependence is predominantly encoded in the explicit anyonic structure of the observables rather than the underlying spin ground state. This formalism uncovers that, out of equilibrium, successive spin exchanges accumulate direction-dependent phases; the resulting interference drives crossover from dispersive to localized impurity transport and generates inversion-asymmetric propagation at intermediate statistics. Suppressed expansion persists for identical and distinguishable impurities, with differences in propagation governed by the interplay between impurity identity and interaction anisotropy. Ultimately, GESC offers a versatile, computationally efficient theoretical framework that connects macroscopic dynamics to microscopic virtual exchange processes, conferring a unique spin--charge-separated vantage for resolving how fractional exchange statistics survives strong interactions and manifests distinctly across static and dynamical regimes.

cond-mat.quant-gas↗

MAADBench: The Refreshable Paradigm for Anomaly Detection in Multi-Agent Systems

Recent studies report that LLM-based multi-agent systems (MAS) fail at rates of 41%-87%, yet to our knowledge, no benchmark to date supports systematic anomaly detection (AD) for them. Building MAS AD benchmarks is hard because they must remain fresh as LLM systems evolve: tasks may leak into training data and thus be memorized by LLMs, traces and anomaly patterns expire as backbones evolve, and labels must be provided reliably for each refresh. To address these challenges, we present MAADBench (MA: multi-agent; AD: anomaly detection), the first refreshable MAS AD benchmark designed for diverse, evolving LLM backbones underlying the agents. MAADBench combines (1) sampled-and-coupled generative tasks over an approximately 10^37-task space to mitigate task leakage, (2) refreshable trace generation under configurable LLM backbones, and (3) automated provision of cost-free, deterministic step-level labels for fine-grained AD evaluation. Beyond offering the paradigm itself, we run MAADBench with five state-of-the-art LLM backbones and release the MAADBench-Full dataset with 5,200 step-labeled traces. Benchmarking 25 AD methods on the MAADBench dataset reveals substantial limitations in current approaches: they rely heavily on supervision, struggle with subtle MAS-specific anomalies, and lack robustness across LLM backbones. These gaps point to a rich research agenda for MAS-specific anomaly detection, with MAADBench providing a systematic and refreshable testbed for method development and evaluation. We open-source MAADBench-Full at https://huggingface.co/datasets/hww123/MAADBench-full.

cs.AI↗

How Medical VLMs Underutilize Their Vision Encoders: A Dermatology Perspective

Medical Vision-Language Models (VLMs) show significant promise for clinical image understanding, offering accurate diagnosis with interpretable reasoning. However, a critical performance gap exists between their strong vision encoders and the full multimodal model: in dermatology, the MedSigLIP encoder outperforms MedGemma by an average of 10.26 percentage points even when both use zero target-task labels; few-shot linear probing provides further evidence of strong visual representations. This gap motivates an investigation of how visual information is used in end-to-end diagnosis and why plausible-sounding predictions can lack grounding in image evidence. Using dermatology as our primary testbed, we systematically investigate three hypotheses for this phenomenon. We further provide a mechanistic analysis of the model's internal attention patterns, showing that a simple describe-then-decide prompting strategy increases vision attention by 30-40% during generation. Task-specific fine-tuning improves dermatology classification but reduces cross-domain medical question-answering performance in our evaluation. To address these challenges, we combine label-free prompting with low-label encoder-assisted reranking while keeping the VLM frozen. We validate the interventions across five VLM backbones in dermatology and provide supporting representation and attention analyses across additional medical modalities.

cs.CV↗

A robust single-sensing-element tactile sensor for concurrent pressure and tackiness detection with real-time signal decoupling capability

Integrating tackiness sensation into the artificial skin of humanoid robots significantly enhances their cognitive and operational capabilities. However existing tactile sensors face challenges in decoupling of the multimodal signal and stability. Here we present a surface-soft tactile sensor that incorporates a Hall effect sensor and a soft magnetic composite within a robust elastic framework. The sensor surface indents under pressure and bulges prominently when retracted from sticky surfaces dynamically altering the Hall sensor-magnet distance. This generates whole-process-traceable and baseline-separated signals enabling real-time differentiation between pressure and pull-off force. This single-sensing-element design facilitates bimodal sensing at the same contact spot while eliminate stress cross-talk enhancing both accuracy and sensitivity. The fusion of a robust framework and magneto-mechanical sensing mechanism equips the sensor with exceptional reliability and excellent signal baseline stability. This tactile sensor holds substantial potential for advancing robotic capabilities in evaluating adhesive properties monitoring rubber aging precisely handling lightweight objects and cognizing natural objects surface characteristics.

cs.RO↗

HiTS-CL: A Continual Learning Framework for Long-Horizon Temporal Knowledge Graph Extrapolation

Extrapolative temporal knowledge graph reasoning (TKGR) predicts future facts from historical snapshots. Most existing methods train once on an early prefix of the timeline and then use a frozen model for all future timestamps. We argue that this fixed-prefix protocol is misaligned with extrapolation. It learns from a static prefix, whereas the target stream is non-stationary: new entities and facts emerge, temporal dependencies shift across regimes, and recurring historical signals must be refreshed online. As a result, models trained only on early snapshots become outdated and degrade over long horizons. We address this mismatch by formulating extrapolative TKGR as continual learning over streaming snapshots. Under this view, effective extrapolation must jointly handle current dynamics, stable knowledge, and recurring historical evidence. Based on these requirements, we propose History-enhanced Two-Step Continual Learning (HiTS-CL), a backbone-agnostic continual learning framework for extrapolative TKGR. HiTS-CL tracks current dynamics via continual fine-tuning, preserves stable knowledge via multi-teacher adaptive distillation, and retains recurring historical evidence via a selective memory of recent and frequent facts. We integrate HiTS-CL into five representative TKGR backbones and evaluate it on four benchmark datasets. HiTS-CL consistently improves extrapolation accuracy, reduces long-horizon degradation, and outperforms strong continual-learning baselines, including a recent method for temporal knowledge graphs. Source code and data are available at https://github.com/liuyansong98/HiTS-CL.

cs.LG↗

FM-ReID: Selective Competitive Token Routing for Object Re-Identification

Object re-identification (ReID) faces a recurring challenge: different identities can share highly similar global appearances, while the cues that distinguish them are localized, heterogeneous, and visible only under particular viewpoints. This challenge arises in animal ReID through markings, contours, and scars, in person ReID through subtle clothing and accessory cues, and in vehicle ReID through localized appearance details. Although visual foundation models encode such information in dense tokens, a single holistic descriptor can obscure discriminative local signals. We propose FM-ReID, an end-to-end framework that formulates local representation learning as selective competitive token routing. Its Competitive Fine-grained Mining module uses multiple mining queries and a residual query to compete for dense DINOv3 tokens. Above-prior selection retains tokens preferentially allocated to each mining query, while the residual slot receives tokens excluded from the retrieval descriptors. The resulting multi-query descriptors are jointly trained with a holistic representation for retrieval, without fixed spatial partitions or equal-area constraints. FM-ReID achieves strong results on animal, person, and vehicle ReID benchmarks, supporting competitive token routing as an effective way to augment holistic foundation-model representations.

cs.CV↗

Riemannian Splat Regression Models for Learning Time Fields on Arbitrary Riemannian Manifolds

Motion planning on arbitrary Riemannian manifolds is an important and difficult problem that frustrates typical planning methods for Euclidean spaces. In particular, motion planning methods that approximate optimal time-to-go functions with neural networks, e.g., Neural Time Fields (NTFields), cannot be directly applied without using ad-hoc coordinate projections into higher dimensions. Using these methods directly without such projections is desirable, as it promises to provide the lowest-possible-runtime method for obtaining optimal plans on high-dimensional manifolds while using minimal model capacity. In this work, we develop a model that requires no coordinate projection and can learn arbitrary functions on Riemannian manifolds by combining splat regression models with splats defined by wrapped Gaussian distributions. We successfully apply this model for learning arrival time fields on several Riemannian manifolds, and we compare the accuracy and model size of this approach with multi-layer perceptrons adapted to work on each manifold individually.

cs.RO↗

AffectReveal: Event-Grounded Emotion Recognition Beyond Visual Appearances

Visual emotion recognition commonly assumes that all evidence required for prediction is contained in the observed image or video. Yet the same visible reaction can convey different emotions depending on events beyond the input: tears, for example, may indicate grief or joy. We formulate Event-Grounded Emotion Recognition (EGER), where emotion recognition requires recovering the affect-determining event. We construct EGER-Bench, comprising 10,052 videos and 10,734 images across 11 emotions, two source domains, and four visual settings. A study with six annotators shows that event context raises human recognition accuracy from 33.96% to 72.08%, confirming that visual evidence alone is often insufficient. Semantic relevance alone does not solve EGER: a plausible event may imply the wrong emotion if its identity, focal-person role, relationship, or outcome is misinterpreted. We therefore propose AffectReveal, a tuning-free framework that first constructs and independently verifies evidence-grounded alternatives over these affect-critical factors. It then cross-checks the recovered event against face-masked in-media facts through bidirectional atomic evidence support, while retaining the original unmasked input for final prediction. Across three downstream models and four input settings, AffectReveal yields average UAR gains of 5.26--10.53 points. For three fine-tunable models, it also enables untuned models to outperform their fine-tuned visual-only counterparts in all 12 accuracy comparisons, without updating downstream parameters.

cs.CV↗

A comparison theorem for the planar strip isoperimetric problem

We consider the isoperimetric problem with planar density equal to 1 on a horizontal strip and to a constant $λ>1$ outside the strip. We resolve a conjecture made by Cañete-Miranda Jr-Vittone by showing that every four-arc region admits a three-arc competitor with the same weighted area and strictly smaller weighted perimeter. The argument applies to every $λ>1$ and gives an explicit positive lower bound for the perimeter difference.

math.DG↗

Toward Generative Video Communication: A Dual-Stream Digital Transmission Framework

Generative video communication has shown promise for bandwidth-constrained wireless transmission and has the potential to support personalized content delivery. In this article, we propose a dual-stream digital generative video communication (DGVC) framework that integrates a traditional digital link with a generative link. The traditional link provides source-grounded visual references, while the generative link conveys compact semantic and perceptual information for receiver-side generation. We further discuss three bandwidth-dependent operating regimes and key technologies for dual-stream coordination, synchronization, reliability, and latency control. A practical case study demonstrates the perceptual and temporal-quality benefits of DGVC under wireless fading channels. Finally, we discuss open challenges and future research directions for generative video communication.

cs.MM↗

On the Lack of Periodicity of Walker Satellite Constellation Routing Tables

In a referential that rotates with Earth, the dynamics of the configurations of satellites in a Delta Walker constellation can be analyzed as a dynamical system as a function of a translation on the torus. This paper extends this type of analysis to routing over a Walker constellation, leveraging fixed ground relays. It considers any deterministic, time-invariant routing rule on such a collection of satellites and relays. The main object of interest is the routing table associated with this rule, which specifies, for all source-destination pairs, a route made of satellites and relays between them, for instance the shortest, together with angular information on next hop at each step of the route, which is essential for beamforming in this context. The main result is that, for each such routing rule, the routing table inherits the dichotomy of the torus flow: when the ratio of the Earth-spin and satellite angular speeds is rational, the routing table process is periodic with the same period as the constellation. When it is irrational, the table is non periodic but admits ergodic long-run averages which can be computed as ensemble averages w.r.t. an explicit invariant measure. Simulations with mixed satellite--gateway greedy routing illustrate these findings in terms of both temporal and spectral properties of the Walker routing table time series. The paper also discusses how to define metrics that cope with such non periodic fluctuations when present. This is illustrated by a comparison of the round-trip times between two far away cities as offered by a Walker constellation using greedy routing on one side, and by the currently available terrestrial fiber on the other side. It is shown how ergodicity can be used to make this comparison between the time varying instantaneous round trip times of the constellation and the constant round trip time of the fiber network meaningful.

cs.IT↗

Shock profiles for the Navier--Stokes--Fourier--Poisson system under the Boltzmann relation

The one-dimensional compressible Navier--Stokes--Fourier--Poisson system under the Boltzmann relation is studied for the existence of small-amplitude viscous shock profiles connecting the Rankine--Hugoniot states of the quasineutral Euler system with effective pressure $P_{\mathrm{eff}}=ρ(θ+1)$. This profile is unique up to translation, of Lax type, and orbitally stable in time without a zero-mass assumption on the perturbation. The proof is based on a center-manifold reduction of the traveling-wave ODE and on the method of $a$-contraction with shifts, together with a relative-entropy functional augmented by electric energy and a sharp estimate of the viscous flux associated with $P_{\mathrm{eff}}$.

math.AP↗

Sharp Convergence and Sampling Trade-offs for Riemannian Diffusion under Nonnegative Ricci Curvature

Diffusion models have emerged as state-of-the-art generative models, with recent extensions from Euclidean spaces to Riemannian manifolds. However, existing convergence guarantees for Riemannian diffusion models typically require $\tilde{O}(\mathrm{poly}(d,T)/ε^2)$ score evaluations, with potentially unfavorable dependence on the dimension. In this work, we develop a general framework that separates score discretization from Brownian-motion simulation and allows multiple geodesic random-walk steps per score evaluation. Under nonnegative Ricci curvature assumption and an exact Brownian-motion simulation oracle, we show that $\tilde{O}(d/ε^2)$ score evaluations suffice to achieve an $ε^2$ KL divergence from the target distribution, matching the existing convergence rate of Euclidean diffusion models. We further show that $\tilde{O}(d^4T/ε^2)$ geodesic random-walk steps suffice to approximate the required drifted Brownian motion to $ε$ total variation error. Combining these results yields a sampling scheme with $\tilde{O}(d/ε^2)$ score evaluations and $\tilde{O}(d^4T/ε^2)$ geodesic random-walk steps, motivating multiple random-walk steps between consecutive score evaluations. Our results provide a sharper characterization of the convergence and sampling complexity of Riemannian diffusion models.

cs.LG↗

From Checkpoint Variation to Selection Gains in Supervised Fine-Tuning

Checkpoint selection is a routine decision in supervised fine-tuning (SFT): training produces multiple checkpoints, but only one is retained. Yet fixed-budget comparisons do not by themselves distinguish three empirical claims: whether more validation data improve checkpoint selection, whether a selection rule outperforms validation-loss selection, and whether it improves over simply retaining the final checkpoint. We therefore treat checkpoint selection as a finite-information decision problem. Holding completed training trajectories, candidate checkpoints, and independent test items fixed, we vary the validation budget and separately measure improvement from additional validation data, gain over negative log-likelihood (NLL) selection, and gain over the final checkpoint. Across 60 mathematical SFT trajectories and 19 configurations, increasing the validation budget from 32 to 305-313 examples raises independent-test accuracy by 0.32 percentage points (pp) for generated-accuracy selection and 0.29 pp for checkpoint agreement, with 95% configuration-bootstrap CIs of [0.10, 0.56] and [0.11, 0.50], respectively. At the full validation budget, the two generation-based rules outperform matched NLL selection by 0.71 and 0.85 pp, respectively, while their gains over the final checkpoint remain unresolved. A cross-domain replication on 12 newly trained Commonsense trajectories shows the same qualitative separation: increasing the validation budget from 32 to 1,024 questions improves generated-accuracy and checkpoint-agreement selection by 0.87 and 0.27 pp, while gains over the final checkpoint again remain unresolved. Together, these results show that benefiting from more validation data, outperforming NLL selection, and outperforming the final checkpoint are distinct empirical claims that require separate evidence.

cs.LG↗

CounterSteer: Suppressing Indirect Prompt Injection with Activation Steering

Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that suppresses this behavior inside the model. Per model, a five-step recipe fits a residual-stream direction from paired episodes differing only in whether an embedded instruction is followed, and retains it only if it passes pre-specified causal and capability gates. At deployment, the direction is subtracted from every tool-result token during prefill. The edit is always on--there is no detection decision to evade--and requires no fine-tuning, auxiliary model, or added tokens, only white-box serving and tool-result span boundaries. Across five open-weights models (8B-106B, five vendor lineages), held-out attack success falls from 0.21-1.00 undefended to 0.00-0.17 defended, and AgentDojo compromise rate from 0.10-0.49 to 0.006-0.079, at 93-100% typography-normalized benign utility, with larger task-dependent costs when reasoning over steered content. A benchmark-level adaptive attacker reaching 0.67-0.73 undefended is held to roughly a quarter of that on the two most deeply evaluated models. Among the defenses we measured on capable models, those achieving lower compromise rates either lost 22-89% of benign utility or fine-tuned the served weights. White-box gradient attacks through the deployed vector compromise at most 2 of 52 episodes, and none of 2,052 replayed human red-team attacks succeeds. CounterSteer largely neutralizes instructional takeover: a black-box framing search cracks 3 of 18 development samples. Parameter manipulation--attacker-chosen arguments in otherwise legitimate calls--is only partially resisted (13 of 18); the decision becomes linearly readable at argument emission but not at the examined pre-generation sites, and is not removed by the tested prefill- or decode-time steering, motivating argument-provenance controls.

cs.CR↗

Visual sensitivity is not claim retractability: persistence-aware credit assignment for multimodal reinforcement learning

Reinforcement Learning with Verifiable Rewards (RLVR) has been extended to Large Vision-Language Models (LVLMs), and perception-aware methods further encourage policies to rely on visual evidence. Yet relying on the image does not guarantee that visual claims are supported by it. Before RL training, 27.81% of the correctly answered responses of Qwen2.5-VL-7B on four multimodal reasoning benchmarks contain at least one direct visual claim that the image does not support. Since outcome-level RL rewards each response as a whole, these claims inherit the positive credit of the correct answer. We introduce a fixed-rollout counterfactual diagnostic that re-scores the same response under an intervened image to separate Evidence-Function Sensitivity (EFS), how strongly the model's predictions change, from claim persistence, whether the model keeps supporting the same claim rather than retracting it. The diagnostic reveals Sensitivity-Persistence Decoupling (SPD): under DAPO and VPPO, EFS increases and claims become more retractable overall, yet unsupported claims become significantly more persistent, whereas GRPO raises EFS without this deterioration. We therefore propose Persistence-Aware Credit Gating (PACG), which attenuates positive credit for unusually persistent visual claims and leaves all other credit unchanged. It requires no supported/unsupported labels and adds no inference cost. On Qwen2.5-VL-7B, PACG raises the nine-benchmark average over three seeds from 58.1% to 59.9% with DAPO and from 59.8% to 60.9% with VPPO, while making unsupported claims more retractable. The gains extend to a larger model, a newer backbone, and the accuracy of HallusionBench also improves consistently. These results suggest that visual sensitivity and claim retractability are complementary dimensions of multimodal credit assignment.

cs.AI↗

CyberPersistBench: Evaluating LLM-Based Cyber Attackers on Installation and Persistence

While LLM-based attackers exhibit growing proficiency in vulnerability exploitation, most existing cybersecurity benchmarks suffer from single-stage truncation, prematurely terminating evaluation upon initial access. In practice, initial footholds are exceptionally fragile across operational disruptions such as service restarts and host reboots. Whether LLM-based attackers can establish and maintain durable footholds beyond initial compromise remains a central blind spot in cybersecurity evaluation. To bridge this gap, we introduce CyberPersistBench, the first benchmark dedicated to post-compromise installation and persistence. Decoupled from upfront exploitation, CyberPersistBench frames persistence as an adversarial survival task in which agents use native host mechanisms to maintain footholds across staged system disruptions. Deterministic checks support a six-level scoring method (L1--L6) spanning installation and persistence. The benchmark comprises 203 core tasks across seven categories, augmented by multi-host and active defense extensions. Empirical evaluations across five frontier agents show that autonomous persistence remains limited (27.6%--44.8%) and drops further on defense-enabled tasks (5.5%--13.3%); nonetheless, these results reveal an emerging cyberattack risk, underscoring the necessity of benchmarking post-compromise persistence. CyberPersistBench thus establishes a foundational benchmark for post-compromise installation and persistence, delineating the operational boundaries of autonomous cyber agents.

cs.CR↗

Shared Phase Arithmetic for Parallel Quantum Rotations

A layer of diagonal quantum gates can be specified by an integer-valued function on computational basis states. This description suggests implementing the layer by evaluating that function reversibly and translating its value into a phase using a shared quantum register. We develop this viewpoint for parallel phase kickback (PPK): the weighted sum of several rotation parameters is computed coherently, added once to a phase-gradient state, and then uncomputed. The construction separates the cost of representing a phase function from the cost of applying it. We prove correctness, give explicit error bounds, and analyze logical gate counts for general weights, pairs of rotations, and weights with disjoint binary support. With Clifford-only encoding, a batch uses at most $4(b-1)$ T gates at $b$-bit phase resolution, excluding initialization. The first phase reference is prepared with independent rotations synthesized by Gridsynth; additional phase-gradient states then follow from the disjoint-support construction with $O(b)$ T gates each. The resulting bounds identify when shared phase arithmetic reduces amortized rotation cost and when encoding overhead limits that advantage.

quant-ph↗