Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,405 records · Page 78Linked to original sources

When the Wrong Key Wins: Understanding and Detecting Hallucinations in LLMs

Large language models can hallucinate even when the knowledge required for a correct answer is already available. We study this failure through a latent-key view of inference, where answer selection depends on competition among associations acquired during pretraining. We show that model predictions can be highly sensitive to individual query keywords, that these influential keywords exhibit entity-specific binding, and that their effects are systematically shaped by pretraining frequency. Multiple bindings can also compete and exhibit higher-order interactions within the same query. Based on this mechanism, we introduce a two-stage keyword-perturbation method for hallucination detection. By removing influential keywords and measuring how the model reorganizes its prediction, the method distinguishes errors caused by misleading key associations from correct decisions supported by diagnostic evidence. Across multiple models and benchmarks, perturbation provides a strong and transferable detection signal, reaching $0.910$ AUROC on probe-known ScientistQA. Finally, we extend the same probabilistic framework to four hallucination regimes: knowledge deficit, wrong knowledge, context distraction, and unstable inference. Their operational distributions across benchmarks provide diagnostic context for why different detector families succeed in different settings.

cs.CL↗

HarnessVLN: Unifying Training-Free Embodied Navigation through an Agent Harness

Embodied navigation requires agents to ground instructions or object goals in spatial observations and translate plans into successful execution. As multimodal large language models (MLLMs) become increasingly capable, they offer stronger support for navigation without task-specific training; however, improved semantic reasoning alone does not ensure that proposed actions remain consistent with spatial evidence, task progress, and execution outcomes. We introduce HarnessVLN, a zero-shot, training-free framework that unifies instruction-following and object-goal navigation through a shared Agent Harness. The Harness coordinates perception, memory, and execution tools through a unified interface, validating planner proposals for evidential support, geometric feasibility, and subgoal consistency before dispatch. It jointly manages hierarchical event memory and a persistent Spatiotemporal Graph to track task progress, preserve spatial evidence, and contextualize failures. Structured execution feedback updates this shared state, guiding subsequent planning, recovery, and termination. Across R2R, RxR, HM3D-v2, and HM3D-OVON, HarnessVLN achieves success rates of 59.6%, 51.4%, 76.0%, and 59.3%, respectively, outperforming prior training-free methods. Humanoid robot deployment further demonstrates its applicability to both navigation tasks in real-world environments. The project page is available at [https://agibot-harnessvln.netlify.app/].

cs.RO↗

Hecke structure of quaternionic modular forms mod $p$

We relate the systems of Hecke eigenvalues arising from the (mod $p$) modular forms on the Shimura curve attached to the indefinite quaternion algebra $B$ of discriminant $δ$ over $\mathbf{Q}$ to the systems of Hecke eigenvalues arising from the (mod $p$) algebraic modular forms attached to the definite quaternion algebra $D$ of discriminant $pδ$. Moreover, we discuss details of the Hecke structure on the latter spaces, following ideas of Serre. The entire setup can be seen as a Shimura curve analogue of Serre's letter to Tate on quaternions and modular forms. The bijection between the sets of systems of Hecke eigenvalues is a special case of recent work of Terakado and Yu; the novelty of this paper is the explicit nature of the construction, allowing for finer control of its behaviour with respect to weights, as well as the results on the Hecke module structure.

math.NT↗

A 3-regular counterexample to the Bilu--Linial signing conjecture

We construct a finite connected simple cubic graph $F$ such that every signing of its edges yields a signed adjacency matrix with an eigenvalue outside $[-2\sqrt2,2\sqrt2]$. This disproves the Bilu--Linial signing conjecture for general regular graphs. The graph $F$ is not Ramanujan, and the conjecture restricted to Ramanujan base graphs remains open.

math.CO↗

IROH: Insightful Ranking Of Humor using Multi-Stage Hybrid Retrieval with Rationale-Distilled LLM Judges for JOKER 2026 Track Task 1 English

Our team, VANGUARD, presents IROH (Insightful Ranking of Humor), a three-stage retrieval system for JOKER Task 1 English at CLEF 2026, achieving first place on the leaderboard with 0.6347 MAP. Our pipeline combines hybrid sparse-dense retrieval, cross-encoder reranking, and a LoRA-adapted Large Language Model judge ensemble. We employ Gemma 4 to generate query-aware rationales under two prompt strategies, generic and typed, and produce up to four types of structured hard negatives for training data construction. Through an ablation across three cross-encoder architectures, four dense embedders, and eight judge configurations, our key findings are threefold: (1) the rationale-distilled judge is the primary driver of ranking quality, whereas appending rationales to the first-stage index contributes negligibly; (2) structured hard negatives degrade generalisation in nearly all configurations despite inflating local validation scores; and (3) across the components we ablate, the lighter, better-calibrated model is competitive with or stronger than its larger counterpart, with the generic-rationale Qwen2.5-7B judge (0.6055 MAP) outperforming every Gemma-4-31B configuration, and the advantage of generic over typed rationales is concentrated almost entirely in the smaller model.

cs.IR↗

Accelerating Transfer-Learning-Based Autotuning with Predictive LLVM IR Performance Ranking

As the complexity of High Performance Computing (HPC) ecosys- tems continually increases, achieving optimal performance becomes a challenge. Traditional performance autotuning techniques pro- vide promising means to navigate this complexity, these techniques remain computationally intensive and require many evaluations to find optimal configurations. This work proposes an autotuning framework that designs a machine learning-based ensemble LLVM Intermediate Representa- tion (IR) ranker, Neural Configuration Scorer (NCS). NCS ranks the performance of IRs sampled by a transfer-learning-based autotuner, improving the efficiency of the tuning process by reducing tuning overheads and circumventing subpar evaluations. By leveraging knowledge from related tasks, we are able to effectively exploit the transfer relationship to access high-performing configurations in fewer samples than traditional techniques that rely upon itera- tive refinement. Our framework can achieve similar performance improvements as state-of-the-art autotuning techniques with up to 61.67% fewer evaluations, averaging 27.85% fewer evaluations across various HPC benchmarks.

cs.PF↗

SWB-DM: A Calibrated Sliced-Wasserstein-Barycenter Aggregator with Delayed-Momentum Caching for Byzantine-Robust Federated Learning under Partial Participation

Robust aggregation methods for federated learning quietly rest on a fragile assumption: that whoever shows up in a given round is a fair sample of the full population. In practice, they rarely are. When only a handful of clients participate per round, even a modest fraction of adversaries can dominate that sample and silently invalidate the finite-sample guarantees that coordinate-wise median, Krum, Bulyan, and trimmed mean all depend on. We introduce SWB-DM to address this directly. SWB treats each slice of a client update as a one-dimensional distribution, computes a trimmed Wasserstein barycenter across clients, and recovers coordinate identity via a medoid-based gauge-fixing step -- a heuristic we developed and do not claim it belongs to standard optimal-transport theory. DeMoA-style delayed momentum then caches updates across the full client population each round, decoupling robustness from whoever happened to be sampled. Trim ratio calibration is not cosmetic: under-trimming causes collapse at corruption levels a properly calibrated model survives. Across 448 CIFAR-10 configurations, plus CIFAR-100, FEMNIST, and a 500-client scalability run, we find several mechanistically distinct failure modes. Even-sample coordinate-wise median degrades to a deterministic wrong answer. Krum silently violates its own n greater than 2f+2 precondition and diverges without warning. Bulyan's n greater than or equal to 4f+3 threshold produces a sharp pass/fail boundary. On attacks, IPM defeats order-statistic defenses -- including SWB -- more reliably than ALIE, confirmed through delta-space measurements against a convergence bound. SWB-DM's cache carries a real warm-up cost, but extending all baselines to the same round budget shows its CIFAR-10 gains are disproportionately large. On CIFAR-100, FLTrust benefits more -- for reasons entirely unrelated to caching.

cs.LG↗

Multiple-Mediator-Affected Ultra-High-Energy Neutrino Attenuation: Hints for 5D ${U(1)}_{L_μ- L_τ}$ in IceCube

Recent IceCube observations point to the ultra-high energy neutrino flux exhibiting a very soft spectral nature beyond tens of TeV, which is unexpected from the standard cosmic-ray neutrino connection. As a potential explanation for this behaviour, we investigate neutrino self-interactions in an extra-dimensional $ {U(1)}_{L_μ-L_τ} $ gauge theory, where symmetry breaking leads to a tower of Kaluza-Klein gauge bosons in compactified four dimensions. These multiple gauge bosons may enhance the attenuation of astrophysical neutrinos as they propagate in the C$ν$B medium. Unlike single-mediator scenarios, for example, which arise from broken $ {U(1)}_{L_μ- L_τ} $ gauge symmetries in four dimensions, the presence of a Kaluza-Klein tower of mediators produces multiple closely spaced resonances whose interference gives rise to a rich energy-dependent behaviour of the scattering cross-section over a vast range of incident neutrino energies. Alongside $s$-channel resonances, off-resonant $t$- and $u$-channel contributions also become important. We explore the possibility that repeated resonant scattering between astrophysical neutrinos and those in the C$ν$B, mediated by these new multiple gauge bosons may increasingly attenuate the former at higher energies, thereby addressing the possibility of softening its spectral nature within the consideration of a single power law.

hep-ph↗

Symmetry Descent in M-theory, Part I: A Twelve-Dimensional Parent Theory

We initiate a symmetry descent procedure for M-theory engineered quantum field theories. Starting from a higher form BF theory supplemented by a cubic bulk topological interaction, we construct a gauge-invariant bulk-boundary system whose edge modes acquire generalized Maxwell--Chern--Simons dynamics after the introduction of a metric-dependent boundary action. The nonlinear contribution to the boundary equation is induced entirely by the cubic bulk interaction. In the case of twelve dimensional bulk parent, the resulting boundary conditions reproduce the local flux equations of the eleven-dimensional supergravity $C$-field, including the gravitational $I_8$ correction. Thus, the electric--magnetic pairing and the nonlinear Maxwell--Chern--Simons dynamics descend from a single 12D topological model. Using Hypothesis H, we identify the bulk equations with the Sullivan model of $S^4$ and interpret their nonlinear gauge transformations as homotopies. We then propose a twisted 4-cohomotopy quantization of the bulk fields and a homotopy pullback description of the coupled bulk--boundary field space. The construction provides both a local dynamical mechanism for M-theory symmetry descent and a candidate global characterization of its fields.

hep-th↗

Local Langlands functoriality for Yu's supercuspidals I: Kaletha's parametrization

This is the first of two papers dedicated to the explicit computation of the Fargues--Scholze correspondence. In this first part, we construct (inspired by, and extending, Kaletha's explicit Local Langlands Correspondence for non-singular supercuspidal representations) semisimple inertial L-parameters for all cuspidal representations arising from Yu's constructions, with arbitrary coefficients. We then establish a characterization of this correpondence in terms of certain functoriality properties. This characterization will be used in the sequel paper in order to compute the Fargues--Scholze correspondence after restriction to inertia.

math.NT↗

Local Langlands functoriality for Yu's supercuspidals II: Fargues--Scholze's parametrization

This is the second of two papers dedicated to the explicit computation of the Fargues--Scholze correspondence. We compute the Tate cohomology of cuspidal representations arising from Yu's construction. Combining this with modular functoriality in the Local Langlands Correspondence, independence of $\ell$, and the partial characterization of the Local Langlands Correspondence established in the first paper, we compare the Fargues--Scholze and Kaletha parametrizations. Among the consequences, we deduce that the Fargues--Scholze correspondence has finite fibers and is surjective onto inertial L-parameters.

math.NT↗

Fully differentiable framework for inverse identification of geometry and material parameters with application to determining stress-free configuration of soft tissues

Medical images of soft tissues typically depict loaded configurations rather than true unloaded (stress-free) states, which can bias biomechanical simulations and inverse material identification. A gradient-based inverse finite element framework is presented that jointly reconstructs an effective unloaded reference geometry and estimates hyperelastic material parameters from two or more observed deformed configurations. The formulation is fully differentiable and leverages exact end-to-end gradients to enable unified, simultaneous optimization of geometry and constitutive parameters. The objective function combines a nodal-position misfit with a deformation-gradient-based mismatch term, improving robustness under large deformations. Benchmark studies quantify the influence of loading diversity and observation count and demonstrate reduced sensitivity to poor material initialization. Finally, application to an MRI-derived breast model shows accurate recovery of the unloaded configuration and constitutive parameters from multiple gravity-loaded states. The framework provides a unified and scalable tool for inverse biomechanics with potential applications in personalized modeling, elastography, and surgical planning.

physics.comp-ph↗

Redesigning the linear--quadratic--Gaussian cost function for feedback cooling of a quantum harmonic oscillator

Linear--quadratic--Gaussian (LQG) control is optimal only with respect to a prescribed cost function, the choice of which dictates the physical objective of the control. We consider feedback cooling of a continuously monitored quantum harmonic oscillator by shifting the minimum of its trapping potential. In this setting, the physically relevant cooling objective can be defined as minimizing the oscillator's energy relative to the feedback-shifted potential. In contrast, conventional LQG control evaluates the energy from a fixed origin and thus fails to directly optimize this quantity. To address this problem, we introduce a redesigned cost function that explicitly accounts for the feedback-induced shift of the potential. We then derive the corresponding optimal feedback law and obtain an analytic expression for the minimum achievable steady-state phonon occupation number. The redesigned LQG control achieves a lower occupation number than low-pass-filter (LPF) feedback formulated for the same cooling objective. While this improvement is minor at detection efficiencies currently attainable in experiments---indicating that LPF feedback already delivers near-optimal cooling performance---the advantage becomes pronounced as the detection efficiency approaches unity. In this regime, the redesigned LQG control provides an increasing advantage for reaching the motional ground state at a finite measurement strength. We clarify that the conventional and redesigned cost functions represent distinct control objectives rather than different implementations of the same optimization problem.

quant-ph↗

Understanding Dynamic Scenes at Gigapixel Scale: Wide-Area Spatio-Temporal Perception from UAVs

UAV-borne imaging has advanced from megapixel to gigapixel sensors, shifting aerial perception from recognizing individual targets to understanding entire dynamic scenes. We characterize this demand as Wide-area Spatio-temporal Scene Understanding (WSTU), which requires wide-area coverage, per-target resolution, and temporal continuity at once, a combination existing datasets lack. To fill this gap, we introduce an ultra-High-resolution (12768x9564) Airborne Remote-sensing Dataset (HARD) annotated at three levels for object detection, multi-object tracking, and scene-level visual question answering. Ultra-high-resolution imagery raises per-frame processing time to seconds. At that scale latency can no longer be ignored in evaluation. Thus, we propose a latency-aware metric for multi-object tracking called streaming-HOTA (s-HOTA). Extensive baseline experiments show how ultra-high-resolution processing reshapes each task. For detection, the end-to-end pipeline affects accuracy and speed as much as the detector itself does. For tracking, high latency charges the association axis far more unevenly than the detection axis, and association is where pipelines diverge. As a result, the pipeline that performs best offline can lose its lead under s-HOTA. For VQA, vision-language models remain weak at cross-frame identity binding and cannot transfer their single-frame gains to it. Together these findings show that the baselines we evaluate fall short of WSTU. HARD provides the data and the systematic baselines to advance it.

cs.CV↗

Variational computation of anharmonic ground and excited vibrational eigenstates using bound Quartic Force Fields: Application with MCTDH and ElVibRot

In this work we introduce the use of Quartic Force fields (QFF) potential expansions in the context of variational calculations. Such potentials are commonly employed in molecular Vibrational Second-Order Perturbation Theory (VPT2) studies, for which equations explicitly dependent on the QFF parameters exist. However, QFF are unbound potentials for more or less large displacements from the reference point and, most of the time, this prevents their use in conjunction with variational wavepacket-based calculations. In this work, we propose a general correction to QFFs and introduce a fully automated numerical approach to avoid their unbound character. Our corrected potentials, bound QFF (bQFF), do not exhibit appreciable modification of the local topography around the region of interest for infrared spectroscopy. As a consequence of this, we can affirm that the vibrational eigenstructure (eigenvalues, eigenstates) remains essentially unaltered by our correction. To illustrate their numerical stability, we have interfaced our bQFF routines in combination with to two well-established quantum simulation software packages MCTDH and \textsc{ElVibRot} which feature variational approaches. More specifically, our bQFFs are separable and hence directly expressible as MCTDH operators. Furthermore, concerning the size of our bQFF expansion, we show that it is possible to tensor-decompose our bQFF in Canonical Polyadic form (CP-bQFF). We use the Monte Carlo Canonical Polyadic decomposition algorithm for this. CP-bQFF results are virtually identical to uncompressed bQFF, but the computational efficiency is largely improved. Our approach paves the way for the automated variational study of anharmonic eigenstates in molecular systems within the reach of QFF-based potentials, using either time-dependent or time-independent schemes.

physics.chem-ph↗

A parity obstruction to completeness of object cotorsion pairs

Fu, Guil Asensio, Herzog and Torrecillas asked whether a complete ideal cotorsion pair of object ideals in an exact category induces a complete cotorsion pair of objects. We give a negative answer in a Hom-finite, weakly idempotent complete Frobenius exact category. Our example consists of bounded complexes of finite-dimensional vector spaces with even total cohomology dimension. Two classes defined by cohomological support generate a complete ideal cotorsion pair, whereas the corresponding object cotorsion pair is neither special precovering nor special preenveloping. The obstruction is that cohomological truncations need not remain in the category, although their doubles do. Conceptually, this obstruction reflects the failure of the standard $t$-structure on the ambient derived category to restrict to the stable category of our example.

math.RA↗

Decodable but Misrouted: Sparse Features Uncover a Readout Gap in Vision-Language Models for Harmful Meme Detection

When large vision-language models misclassify harmful memes, the failure may reflect missing internal evidence or an inability to route represented evidence to their outputs. We distinguish these cases in Gemma-3 and Qwen3.5 using sparse autoencoders, role-conditioned probes, causal interventions, and recovery experiments across six harmful content benchmarks, with additional Spanish and Hindi-English code-mixed evaluations. Sparse readouts outperform native prediction on all six primary binary tasks: Qwen averages $0.740$ versus $0.432$ for native macro-F1, residual reconstruction reaches $0.486$, and Gemma improves from $0.532$ to $0.714$. These gains measure how accessible the label is to a supervised readout; they do not show that the model's native generation already applies such a decision rule. Under the evaluated scales, Qwen silent-feature ablation is $24-63$ times more probe-sensitive, whereas routed-feature patching on literal yes/no tasks is $16-140$ times more output-sensitive. Native-only threshold calibration explains much, but not all of the gap: on five tasks with matched probe scores, it recovers $69.8$\% of the raw native-to-probe difference, while direct routing adds $0.094$ mean macro-F1 beyond calibrated native scoring. Joint gold-label, probe-KL, and pairwise LoRA supervision improves dedicated FHM prediction, but a gold-only adapter performs better on the shared seven-task mean. A case study of Gemma-3-12B on the Facebook Hateful Memes dataset finds a distributed rank-32 image-prompt interaction, reaching $0.756$ versus $0.685$ native macro-F1. Robustness controls show that the signal is not explained solely by accompanying OCR and depends on paired visual evidence, and that it extends beyond English. In many of the errors we study, the evidence is represented but does not reach the answer; therefore, routing is a common bottleneck in harmful meme classification.

cs.CV↗