Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Sample-based quantum diagonalization approach for open-shell transition-metal complexes in gas and implicit-solvent

Open-shell $3d$ transition-metal complexes challenge electronic-structure methods because competing spin states, charge transfer, and solvation jointly determine their energetics. Here, we combine sample-based quantum diagonalization (SQD) with the integral-equation-formalism polarizable continuum model (IEF-PCM), extending SQD to correlated open-shell transition-metal systems in a dielectric environment. We investigate the octahedrally coordinated $\mathrm{[Co(H_2O)_5CO_2]^{2+/3+}}$ complex across two oxidation states, four spin multiplicities, and a metal-ligand dissociation coordinate. We study the Co(III) singlet and quintet states and the Co(II) doublet and quartet states, incorporating open-shell references into SQD-IEF-PCM through an outer self-consistent reaction-field loop. Using samples collected on an IBM Heron quantum processor and active spaces of up to 50 qubits, SQD reproduces coupled-cluster and heat-bath configuration-interaction benchmarks within the same active space in the gas phase and implicit solvent, with a largest observed deviation below 9 $mE_h$. Along the dissociation coordinate of high-spin quintet $\mathrm{[Co(H_2O)_5CO_2]^{3+}}$, SQD resolves an avoided crossing caused by internal charge transfer; this feature is absent in the singlet and the lower oxidation state of the complex. Relative to the gas phase, implicit solvation stabilizes for the quintet state the neutral CO$_2$ dissociation and suppresses the avoided-crossing feature. To our knowledge, this is the first hardware demonstration of SQD for an open-shell $3d$ transition-metal complex in gas phase and implict solvent. These results establish SQD as a robust quantum-centric approach for transition-metal chemistry where spin state ordering, charge transfer, and environmental effects are strongly intertwined.

quant-ph↗

Semantically Similar, Yet Not Answerable: Diagnosing the Semantic-Answerability Gap in Table RAG

In retrieval-augmented generation (RAG), semantic relevance asks whether a source matches a query in meaning, while answerability asks whether it contains sufficient information to answer the query. A Semantic-Answerability Gap (SAG) may arise in retrieval when a retriever can reach semantically relevant sources yet fail to identify those that are uniquely answerable. We uncover this gap using tables as a controlled setting, where shared schemas and entities provide strong semantic signals while localized content and row-column bindings distinguish answerable from non-answerable sources. Using TCR-Bench, a controlled sibling-table benchmark, we find that dense retrievers achieve only 18.2% top-1 target retrieval, reducing QA F1 from 0.755 with the oracle table to 0.330 with retrieved top-5 tables. Controlled diagnostics show that retrievers favor semantic volume over sufficiency, respond weakly to row-column binding disruptions, and struggle to distinguish Targets from Siblings. Explicit answerability assessment substantially improves target identification, while fine-tuning shows that answerability is learnable but difficult to transfer without compromising broad semantic retrieval.

cs.AI↗

SAGE: Subgoal-Conditioned Action Generation for Latent World Model Planning

Latent world models have emerged as a powerful planning paradigm by learning action-conditioned predictive dynamics and using them as internal simulators to imagine and evaluate candidate action sequences. However, as the planning horizon grows, performance becomes increasingly constrained by proposal quality: a fixed candidate budget must search an exponentially larger action space, making it difficult to expose the world model to high-quality candidate futures for evaluation. In this paper, we introduce SAGE, a prior-conditioned planner that replaces random proposal initialization with structured guidance. At each planning stage, a goal-conditioned generator predicts the next intermediate latent subgoal for a specified duration, which is then used to condition the generation of candidate action sequences. To capture semantic information across temporal scales, we use subgoals of varying durations as priors, balancing fine-grained local control with higher-level long-horizon progress. Then the frozen world model evaluates these proposals against the same subgoal and guides their refinement before execution. Experiments on PushT and OGBench Cube show that coupling latent subgoal decomposition with prior-conditioned action generation substantially improves long-horizon planning while preserving strong short-horizon performance. To be specific, when the target offset is $150$, it raises PushT success from $4.7\%$ to $64.7\%$ and OGBench Cube success from $20.7\%$ to $67.3\%$. We further extend latent world-model planning to LIBERO, where SAGE improves full-episode success from $0\%$ with the vanilla LeWM planner to $48.7\%$ on Scene2 and $58\%$ on Caddy.

cs.AI↗

RuNNer 2.0: A Software Suite for High-Dimensional Neural Network Potentials

We present RuNNer 2.0, the "Ruhr University Neural Network energy representation", a highly optimized software suite for training and evaluating high-dimensional neural network potentials (HDNNPs) of the second, third, and fourth generation. Long-range electrostatics and charge equilibration (QEq) for the description of non-local charge transfer in fourth-generation (4G) HDNNPs are accelerated by quasi-linear-scaling plane-wave methods, reducing QEq computational complexity from $\mathcal{O}(N^3)$ to $\mathcal{O}(N\log^2 N)$ such that linear or quasi-linear scaling is achieved across all HDNNP generations. An optimized memory management strategy eliminates the training overhead traditionally associated with long-range interactions, allowing 4G-HDNNPs to be trained with the same efficiency as their local counterparts. Developed in modern Fortran (2003/2008 standards), combined with a hybrid MPI/OpenMP parallelization scheme, RuNNer 2.0 has been designed to run efficiently in any CPU environment, from cost-effective local workstations to massive HPC clusters. Its modular library architecture facilitates straightforward binding to external simulation software; native interfaces to LAMMPS and the Atomic Simulation Environment (ASE) provide full access to all its features, including built-in committee-based uncertainty quantification. The high efficiency and scalability of the RuNNer 2.0 ecosystem are demonstrated through detailed benchmarks.

physics.chem-ph↗

Do Newer Models Produce Better Patches? A Non-Functional Quality Study

Repository-level coding benchmarks typically measure progress in model capability by comparing the resolved rates of later and earlier models. However, this focus overlooks whether the non-functional quality of their generated patches has also changed across model generations. This study investigates whether later models produce functionally correct patches with better non-functional characteristics than earlier models on comparable repository-level repair tasks. We conducted two case studies involving four Claude and DeepSeek models on SWE-bench Lite. Using the same SWE-agent functional repair setting, we evaluated the generated patches with CodeQL, CodeScene, CPU time, and peak memory. Our primary analysis compared the models on commonly resolved instances. The static analysis results showed that most CodeQL paired differences were zero and that no CodeQL or CodeScene comparison remained significant after Holm correction. CPU time differences were small and inconsistent across model families, while peak memory usage was slightly higher for the later models under the benchmark test workload, with small absolute differences. Differences in individual CodeQL rules and CodeScene categories varied across model families and did not survive multiple-comparison correction. Overall, later models resolved more instances but showed no consistent improvement in the measured non-functional indicators on tasks solved by both models. Through this study, we hope to encourage a more comprehensive evaluation of models' practical software engineering capabilities.

cs.SE↗

STeP: Signal Temporal Logic for Precise Specifications for Action Generation with Vision Language Models

Natural-language robot instructions often specify more than a coarse task goal: they may impose spatial, temporal, and logical requirements that must remain satisfied throughout execution. We present STeP, a specification-based agentic framework that uses Signal Temporal Logic (STL) as an explicit interface between high-level language reasoning and low-level robot execution. Rather than encoding such requirements implicitly in a learned policy, STeP formalizes them as task specifications that can be decomposed across multi-stage manipulation, enforced during execution, monitored online, and used as structured feedback for replanning. We evaluate STeP on standard LIBERO and LIBERO-PRO, and introduce LIBERO-Constrained, a new benchmark for manipulation tasks with spatial, temporal, and logical requirements, together with real-world tabletop experiments. On LIBERO-PRO, STeP retains 54%-89% success across five of six evaluated perturbation settings, substantially outperforming VLA and code-as-policy baselines under distribution shift. On LIBERO-Constrained, STeP achieves 80% safe success across 49 task-constraint instances; on real-world tasks, it improves safe success over a specification-free model-based baseline across all four task categories, with gains of up to 45 percentage points. These results support explicit formal specifications as a practical interface between foundation-model reasoning and reliable robot execution.

cs.RO↗

MOPDA: Mixed-Trajectory On-Policy Distillation for Language-Guided Industrial Anomaly Detection

Large vision-language models (LVLMs) have shown strong potential for industrial anomaly detection (IAD) by providing image-level anomaly judgments and interpretable reasoning. However, reliably translating generated judgments into precise pixel-level localization remains challenging. We propose \textbf{M}ixed-Trajectory \textbf{O}n-\textbf{P}olicy \textbf{D}istillation for Language-Guided Industrial \textbf{A}nomaly Detection (MOPDA), the first framework to introduce on-policy self-distillation into LVLM-based IAD. For judgment learning, \method introduces \textbf{Mixed-Trajectory Supervision}, combining student-generated on-policy trajectories with evidence-conditioned teacher trajectories under a shared token-level distillation objective. Student trajectories preserve supervision on deployment-relevant response paths, while teacher trajectories provide complementary evidence-conditioned supervision. For dense localization, \textbf{Language-guided Visual Anchoring} uses the final judgment as a compact semantic condition to construct image-specific normal and abnormal anchors, which are contrasted with dense visual features to produce anomaly maps. This keeps language as semantic guidance while grounding pixel-level responses in visual evidence. Under a strict cross-dataset zero-shot protocol on five IAD benchmarks, \method outperforms the evaluated LVLM-based baselines on most detection, localization, and judgment metrics while remaining competitive with CLIP-based methods. Ablations further validate both mixed-trajectory supervision and final-judgment conditioning.

cs.CV↗

Confidently Deceptive: On the Relationship Between Confidence and Deception in LLMs

The increasing capabilities of large language models (LLMs) are being accompanied by deep-rooted risks of deceptive behaviours that cause models to produce misleading outputs in service of a contextually or experimentally induced goal. The harm posed by such behaviours depends not only on the content of deceptive outputs but also how confidently models deliver them, since confidence has a major impact on how persuasive the communication is to end users. In this paper, we provide a comprehensive study on the crucial relationship between confidence and deception across existing deception benchmarks and different model families, while covering both verbalized numerical and logit-based aggregated confidence. Through this, we reveal how confidently models behave when being deceptive. We demonstrate that when producing deceptive rather than honest responses, models exhibit a gap between their belief (how likely they think a claim is to be true) and their commitment (how firmly they assert and would defend that claim). LLMs produce persuasive deceptive claims while reporting low belief in their factual correctness. Their reported commitment to deceptive responses can easily be increased through further prompting and preference fine-tuning, with smaller and condition-dependent changes in reported belief. However, we show that low reported belief remains comparatively invariant and provides a strong signal for detecting deception in the evaluated settings. Using only an API call, our approach achieves detection scores of up to 0.99 for induced deception and 0.89 for emergent deception. This ultimately shows how confidence can be a practical tool for detecting and diagnosing deceptive behaviour in LLMs.

cs.CL↗

Interaction Dynamics Modeling and Predictive Control for Safe Steerable Catheter--Tissue Interaction

Steerable catheters are the primary tool for cardiac electrophysiology (EP) procedures including radiofrequency ablation, where the tip must be positioned precisely at target tissue while maintaining controlled, stable contact. The central control problem is therefore not merely tip tracking and not merely force regulation; it is the regulation of \emph{catheter--tissue interaction dynamics}. The interaction state must encode how the tip moves relative to tissue, how persistent friction and contact forces bias that motion, and how safety limits reshape what motion is physically allowable. Existing methods regulate these interaction dynamics through different mechanisms. Classical impedance control~\cite{hogan1985} shapes the tip port as a virtual mechanical impedance $Z(s) = M_d s^2 + D_d s + K_d$, providing passive compliance without an explicit contact model. Three complementary design requirements motivate the present formulation: \textbf{(i)}~an explicit force-related constraint, \textbf{(ii)}~compensation for steady error under persistent loading, and \textbf{(iii)}~a prediction model that can incorporate trajectory and actuator information.

eess.SY↗

Flavour current correlators and the non-Abelian hydrodynamic approximation: the charged sector

Flavor-current correlators are studied in strongly-coupled dense (holographic) matter, at finite quark chemical potential $μ_q$ and finite isospin asymmetry. The non-Abelian hydrodynamic description of the charged currents is derived in the presence of an isospin chemical potential $μ_3$. The two-point correlators of charged currents are then computed holographically at finite quark and isospin chemical potentials. In the near-extremal hydrodynamic regime, $ω, k, T, μ_3 \ll μ\equiv \sqrt{μ_q^2+μ_3^2}$, relevant for cold strongly coupled matter, the IR properties of the correlators are studied. It is shown that in this regime, the correlators agree with the non-Abelian hydrodynamic predictions. Therefore, the traditional regime of validity of standard hydrodynamics extends beyond $ω, k \ll T \ll μ$ to the so-called extended hydrodynamic regime $T\ll ω, k \ll μ$. The holographic product formula is applied to the present non-Abelian system, and is used to propose an extended hydrodynamic approximation capturing both hydrodynamic-like poles and the leading effect of AdS$_2$ poles, by resumming the low-$ω$ logarithms. The results are verified through a detailed numerical analysis of the exact correlators and quasi-normal mode spectrum.

hep-th↗

Quantifying Event-Related (De)Synchronization Variability for Brain-Computer Interface: A Unified and Interpretable Framework

Objective: Brain-Computer Interfaces (BCIs) enable the control of external devices by decoding user intentions from electroencephalography (EEG). However, substantial EEG variability within and between users remains a major challenge. To better understand this variability, we propose interpretable metrics that independently quantify temporal, spatial, and frequency variability in BCI related brain activity within and between users. Methods: We propose a framework to quantify variability by extracting EEG features and defining variability as their dispersion around their centroid using appropriate distance functions. Using two motor imagery BCI datasets (N = 133 users), we investigated the relationship between BCI performance and the variability metrics through within-user and cross-user classification experiments. Results: Negative correlations of -0.2 to -0.4 were observed across most conditions, suggesting that lower variability is associated with higher BCI performance. Moreover, the metrics revealed differences in robustness to variability between the deep learning and Riemannian-based classifiers, with the former showing weaker correlations. Conclusion: The results demonstrate the effectiveness of the proposed variability metrics and suggest that reducing variability may improve BCI performance while revealing differences in the sensitivity of classification models to different types of variability. Significance: The framework quantifies temporal, spatial, and frequency variability at multiple hierarchical levels (within-trial, between-trial, and between-trial-group), providing interpretable measures to better understand EEG variability and support more robust BCIs. It could also be used to characterize dataset variability, evaluate classifier sensitivity, incorporate variability into objective functions, and provide variability-based user feedback.

eess.SP↗

Flash EQ-Linear: Accelerating Equivariant Linear Layers via Group-wise Discrete Fourier Transform

Equivariant networks embed geometric symmetries as structural priors through weight sharing, achieving remarkable parameter efficiency across vision tasks. However, this parameter efficiency does not translate into compute efficiency: most existing implementations unroll the structured weights into dense matrices and dispatch them to generic dense kernels, so an equivariant layer costs no fewer MACs than its non-equivariant counterpart. In this paper, we observe that the equivariant linear (EQ-Linear) layer---the most fundamental and frequently used module in modern equivariant architectures---is essentially a circular convolution along the group dimension composed with a linear transform along the channel dimension. Building on this observation, we propose Flash EQ-Linear, an exact acceleration algorithm that reduces the cost to $2(T-1)/T^2$ of the original dense formulation ($T$ is the equivariant group size) by combining the Fourier convolution theorem along the group dimension with the conjugate symmetry of the real DFT. To translate these computational savings into wall-clock speedups, we further develop dedicated CUDA kernels for the $\mathrm{p}4$ group. At the operator level, Flash EQ-Linear achieves up to $2.1\times$ forward speedup over PyTorch's highly optimized F.linear; at the network level, Flash EQ-ViT achieves up to ${1.7\times}$ end-to-end speedup over both equivariant and non-equivariant baselines. As an operator-level acceleration algorithm, Flash EQ-Linear provides plug-and-play acceleration for diverse pretrained equivariant models, including EQ-ViT, EQ-Swin, EQ-VMamba, and EQ-INR, without retraining or architectural changes. Code is available at https://github.com/zhongchenzhao/FlashEQLinear.

cs.CV↗

On the Approximation of the Unitary Operator Group Associated with a Rotation Matrix and Its Applications to Abstract Hyperbolic Equations

The solution of the Cauchy problem for homogeneous abstract hyperbolic equations, together with its derivative, admits a vector representation in terms of a unitary operator group associated with a rotation matrix. A rational approximation of this unitary group is constructed and shown to possess optimal fourth-order convergence. The order of convergence is determined in accordance with the smoothness scale. Based on this rational approximation, a two-layer semi-discrete scheme is constructed for the approximate solution of Cauchy problems for nonhomogeneous abstract hyperbolic equations in both the linear and semilinear settings. The convergence properties of the scheme are examined in relation to the regularity of the solution.

math.NA↗

Priors learned from legacy reconstructions inherit undetectable overconfidence

Where truths are scarce (e.g., seismic and medical imaging), learned priors in ill-posed inverse problems are trained on archives of legacy reconstructions---i.e., an older method's outputs---and their reported uncertainty is taken as data-driven. We show that this prior is, in the population limit, exactly the regularizer that produced its archive of posterior samples, advanced one expectation--maximization step toward the truth. While the step improves the regularizer on the directions the measurements resolve, it leaves the regularizer's assumption on the operator's blind subspace unchanged. An archive of single-best reconstructions collapses the blind interval to zero width. Neither error is detectable in practice, as truths differing only on the blind subspace share the data law, and simulation-based calibration is neutral by construction. We identify from the operator alone which directions the measurements do not inform, and, given a handful of ground-truth models, build intervals there that contain the truth as often as they claim to. We validate these findings on a two-dimensional example with closed-form predictions and in controlled experiments on seismic-imaging and groundwater-flow operators, against priors trained on the truth.

stat.ML↗

ViTacWorld: Scaling Visuo-Tactile World Models for Contact-Rich Robot Manipulation

Contact-rich robot manipulation requires physical interaction cues that are often invisible to cameras, making tactile sensing essential for robust control. However, scaling visuo-tactile robot learning remains difficult because real tactile interaction data are expensive to collect, hardware-dependent, and limited in task and scene diversity. We present ViTacWorld, an action-conditioned visuo-tactile world model for scalable contact-rich robot manipulation. ViTacWorld leverages public real tactile datasets and a constructed simulation environment to scale visuo-tactile-action data, exploiting the fact that tactile signals are directly grounded in physical contact and can exhibit a smaller simulation-to-real gap than purely visual observations. The model is first pretrained with large-scale real and simulated visuo-tactile trajectories, and then finetuned with real-world policy rollouts to better match downstream manipulation behaviors. Given robot actions, ViTacWorld predicts temporally aligned visual observations and tactile feedback, enabling visuo-tactile-action rollout generation. To the best of our knowledge, ViTacWorld is the first framework that uses a world model for robot visuo-tactile-action trajectory generation and policy evaluation. It serves two roles: synthesizing rollouts to improve downstream tactile policies, and evaluating policies by predicting action-conditioned visuo-tactile outcomes under controlled action sequences. Experiments on contact-rich manipulation tasks show that ViTacWorld generates physically meaningful rollouts, improves policy performance through scalable data augmentation, and enables action-conditioned policy evaluation. Project page: https://vitacworld.github.io/

cs.RO↗

How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift

Post-training is a key mechanism for adapting large language models to downstream tasks. While prior work suggests that task adaptation can alter a model's pre-existing alignment, especially its safety behavior, its broader effects across alignment domains remain poorly understood. We address this gap through a systematic evaluation of representative task-adaptation methods, including supervised fine-tuning (SFT), KL-regularized SFT, and reinforcement learning with verifiable rewards (RLVR) across 15 alignment aspects spanning six key domains: safety, factuality, stance stability, social harm, controllability, and instructability. Our results reveal that post-training does not reshape alignment uniformly. RLVR improves task performance while inducing comparatively small, but non-zero, metric-specific shifts, while SFT leads to substantially larger alignment drift across domains. KL regularization mitigates this effect: stronger reference-model anchoring reduces alignment drift from the baseline, although KL-SFT still falls short of RLVR in preserving alignment. Representation-level analysis further supports this pattern, with shifts in alignment-relevant representations tracking behavioral drift. Together, these results show that task adaptation is not merely a capability-improving step, but an alignment intervention in its own right, motivating multi-dimensional alignment evaluation as a standard component of post-training pipelines.

cs.AI↗

DualityCert: Verifier-Gated Language-Model Repair of Broken Duality Claims in Quantum Field Theory

We present DualityCert, a symbolic verifier for candidate Seiberg-duality claims in four-dimensional N=1 quiver gauge theories. The verifier evaluates 't Hooft anomaly matching, superpotential R-charge consistency, central-charge matching, and a bounded chiral-ring proxy. A claim that passes receives a consistency certificate, which states that no tested inconsistency was found, not that the duality is proven. We use the verifier as a repair environment for language-model agents, which receive a deliberately broken claim and must edit it until it certifies. On a preregistered benchmark of 145 broken claims, with the analysis fixed before the first confirmatory model call, verifier-gated retry improves final repair success over a single attempt by +8.3 percentage points (pp) on deepseek-chat and +7.1 pp on qwen-plus (Holm-adjusted p<0.002). Under an equal budget of eleven attempts, the stop-first strategy portfolio underperforms independent verifier-filtered resampling by 10.3 percentage points on deepseek-chat but outperforms it by 14.7 points on qwen-plus, reversing the ordering of the two tested verifier-exploitation policies across the two confirmatory models. On qwen-plus, category-level verifier feedback is worth +8.7 pp over content-free retry, and interpretable obligation identities alone are worth +6.4 pp over structurally identical masked feedback. Neither effect is detected on deepseek-chat. Separately, a preregistered MiniMax-M2.5 extension again finds an iteration gain and independent verifier-filtered resampling outperforming the strategy portfolio. Which policy is better thus differs between the two models, while every winning policy uses the same cheap certificate. The verifier, benchmark, protocol, and all per-attempt records are released.

cs.CR↗

Variational Boosting for Physics-Informed Neural Networks

Physics-Informed Neural Networks (PINNs) solve differential equations by minimizing the residual of a nonlinear operator over a neural parameterization of the solution. However, monolithic PINNs often suffer from ill-conditioning, spectral bias, and optimization instability. We introduce a variational boosting framework in which solutions are constructed additively in function space. Each stage trains a weak learner whose converged correction satisfies a local orthogonality condition, equivalent to a projected functional gradient descent step onto the tangent space of the network's function manifold. Because each correction network is deliberately small, the restricted minimization admits full Newton or conjugate gradient updates, which are typically infeasible in large PINNs. The resulting method separates global nonlinear refinement into a sequence of well-conditioned subproblems while preserving the full variational structure of the operator. This framework provides a geometric interpretation of multi-stage PINNs as projected functional gradient descent and enables stable second-order optimization for nonlinear differential equations.

cs.LG↗