Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

The bare necessities of a physically reasonable mathematical model for quantum theory

What are the minimal requirements for a physically reasonable mathematical model for quantum theory or a potential extension? In addressing this question, rather than relying on the generalized probabilistic theories or other usual approaches, we propose a reset based on the following fundamental features: transition probabilities, which are typical of quantum theory, and continuous reversible dynamical processes (e.g., represented by a Lie group), which play the same crucial role in classical as well as in quantum physics. In doing so, we get back to the very elementary transition probability space framework originally introduced by Bogdan Mielnik in 1968 and combine it with the continuous reversibility condition. The primary class of valid spaces arises from the irreducible atomic JBW algebras, which encompass the irreducible atomic von Neumann algebras and the Jordan matrix algebras. The von Neumann algebras represent the model of common quantum theory. Beyond these, there is a non-standard valid space, where the exceptional Lie group $E_6$ acts transitively. While $E_6$ has been proposed as a candidate for internal symmetries in particle physics, we highlight that this model deviates from standard quantum theory in critical ways. Specifically, it fails to guarantee the existence of post-measurement states in all situations. Though the above postulates are powerful, they still allow for some physically meaningless models. A further feature of quantum theory (the correspondence between the pure states and the minimal projections) leads us to an additional requirement that rules out these models. However, a complete classification of the valid models remains one of the open issues, pointed out in the paper, which is intended for the mathematical and physical experts, inspiring them to tackle these problems.

quant-ph↗

Sharp continuity of quantum conditional entropy

We prove the sharp uniform continuity bound for quantum conditional entropy. If two bipartite states are at trace distance at most $δ$ and $d=\dim A$, the optimal dimension-only modulus of continuity is $h_2(δ)+δ\log(d^2-1)$ up to $δ=1-d^{-2}$ and $2\log d$ thereafter, where $h_2$ denotes the binary entropy. When $\dim B\ge d$, this bound is tight for every $δ\in[0,1]$. The key proof idea was developed with the assistance of ChatGPT 5.6 Sol, building on and adapting the tight classical proof of Alhejji \& Smith [IEEE ISIT (2020)], which follows a conceptually different approach.

quant-ph↗

Area-minimizing submanifolds are not generically smooth, except for geodesics, minimal surfaces, and minimal hypersurfaces

We prove that area-minimizing submanifolds in mod $2$ homology are not generically smooth, except in the case of geodesics, minimal surfaces and minimal hypersurfaces. This answers questions of White and Yau that ask about the generic smoothness of area-minimizing submanifolds. We furthermore establish a lower bound on the Hausdorff dimension of the singular sets of area-minimizing submanifolds with respect to open sets of Riemannian metrics. The lower bound is ${(d-3)}$ where $d$ denotes the dimension of the submanifold. As a crucial step, we prove that the cone over the Veronese minimal embedding of $\mathbb{RP}^2$ is mod $2$ area-minimizing, settling another long-standing open problem.

math.DG↗

A Numerical Realization of Suzuki's Weil-Quadratic-Form Operator: The Archimedean Spectral Law, its Universality, and an Operator Form of Weil's Positivity Criterion

This paper presents the first numerical realization of Suzuki's Weil-Quadratic-Form operator, a candidate for the Hilbert--Pólya program linking spectral positivity to the Riemann Hypothesis (RH). Suzuki's 2026 construction was purely theoretical; here, the operator is instantiated via P1 finite-element discretization and Richardson extrapolation. Key results include: (R1) In the prime-free regime, the spectrum follows a closed Archimedean law $A_k(a) = \log(1/a) + \log(k-2) + B_0 + O(a)$, with $B_0 = \log q - 2\log 2$, confirmed to 30-digit precision. (R2) A Mellin double-pole argument proves the head coefficient $B(ν)$ and shows $B_0$ depends only on the conductor $q$, independent of the Archimedean parameter. (R2b) The degree $d$ of an L-function appears directly as the logarithmic slope of the spectrum. (R3) Total spectral intensity follows the prime number theorem, $S(a) \sim (2a)^3/6$. (R4) Nontrivial zeros are not eigenvalues but occur in the explicit-formula error term of the prime symbol. (R5) The best-match line $σ^*(a)$ descends toward the critical line. (R6) Weil's positivity criterion is realized in operator form: bounded residual growth corresponds to all zeros on the line, while an injected off-line zero causes exponential blow-up. (R7) The lowest eigenvalue $λ_1(a)$ is strictly positive, decays superexponentially, and passes smoothly through the first prime threshold. (R8) The characteristic function $W(a,0;z)$ is computed for the first time, with all zeros confirmed real. (R9) Indirect traces of GUE statistics appear in the moment structure, even where direct detection is blocked. The authors emphasize that this work does not prove RH. All results are Archimedean and universal, with significance lying in the faithful numerical realization of classical identities rather than new arithmetic.

math.GM↗

Cross-Correlations of Metric and Monopole Perturbations from Holographic Cosmology

In this paper, we present the 1-loop calculation of the three-point function $\langle TJJ \rangle$ of the stress-energy tensor and two insertions of $SO(3)$ global currents, using a 3d toy model for holographic cosmology. By applying the holographic dictionary that relates these QFT $n$-point functions to cosmological correlators, together with the $Sl(2,\mathbb{Z})$ duality that maps the electric Noether current to a magnetic vortex current dual to cosmological magnetic monopoles, we relate the $\langle TJJ \rangle$ correlator to the cross-correlations between metric perturbations and the bulk magnetic monopole field, specifically mapping to the non-Gaussianities $\langle ξ\tilde{A} \tilde{A} \rangle$ and $\langle γ\tilde{A} \tilde{A} \rangle$. We calculate the semi-local contact terms necessary to achieve the exact factorization of the cosmological correlators in the squeezed limit, $p_1 \to 0$. Finally, we evaluate the effective non-linear parameters, showing that the scalar-monopole cross-correlation vanishes at leading order, $f_{NL}^{ξ\tilde{A} \tilde{A}} = 0$, while the tensor-monopole cross-correlation yields a non-zero $f_{NL}^{γ\tilde A\tilde A}$. These results respect the expected amplitude hierarchy of the non-Gaussian correlators, while pointing at new directions in which holographic cosmology can be tested experimentally.

hep-th↗

Asymmetric quantum information scrambling in inhomogeneous XXZ spin chains

Inhomogeneous systems have become increasingly relevant in experimentally engineered quantum many-body systems, where spatially varying interaction strengths can significantly influence nonequilibrium dynamics. Motivated by this, we numerically study the out-of-time-ordered correlators (OTOCs) in inhomogeneous XXZ spin chains. Using two types of spatially varying interaction profiles, namely, linear-gradient and stepwise profiles, we show that spatial inhomogeneity in the interactions can induce pronounced asymmetry in information scrambling, with OTOC operators located at different sites exhibiting distinct dynamical behavior. In particular, OTOCs associated with the strongly interacting side exhibit suppressed scrambling compared to those associated with the weakly interacting side, with the difference becoming more pronounced when the interaction strengths differ significantly. To elucidate the origin of the long-time saturation of OTOCs, we analyze the diagonal matrix elements of the OTOC observables in the energy eigenbasis and their overlap with the Hamiltonian, which explicitly incorporates the spatial interaction profile. The resulting analytical expression is consistent with the numerical results. Finally, we extend our analysis to periodically driven inhomogeneous spin chains and show that the interaction-gradient-induced asymmetry persists under periodic driving.

quant-ph↗

Explicit Layer Modeling for Video Object Insertion and Video Layer Decomposition

Most video editing systems still lack explicit layered video representations, limiting realistic compositing, object reuse, and consistent manipulation. This limitation is particularly evident in video object insertion and video layer decomposition, where existing methods lack direct supervision for foreground layers that capture both objects and their associated visual effects. We introduce TriLayer, a triplet video dataset containing aligned composite--background--foreground videos, where the foreground layers include both object appearance and associated visual effects. With aligned triplet supervision, TriLayer enables explicit supervised learning of layered video representations for the first time. Building on this dataset, we propose DBL-Diffusion, a dual-branch diffusion framework that jointly models scene-level RGB content and RGBA foreground layers through cross-branch interaction during denoising. We instantiate the framework in two tasks: DBL-Insert for layered object insertion, which generates explicit RGBA layers for realistic compositing and flexible post-editing, and DBL-Decompose for video layer decomposition, which recovers foreground and background layers using triplet supervision. Experiments demonstrate that explicit layer modeling substantially improves both insertion fidelity and decomposition quality.

cs.CV↗

Dirac Fermion Scattering and Conductance Response in Asymmetric Graphene Wormholes

We calculate massless Dirac scattering and the conductance response to a localized scalar ring gate in a graphene wormhole with two independent curvature scales and no axial magnetic flux. Matching Hankel and Gauss hypergeometric spinors gives the unperturbed scattering states. Interchanging the throat segments preserves the total Landauer conductance, while a gate at a fixed displaced position produces different conductance changes in the two mirror configurations. For $a=5$~nm and $r_L+r_R=13.68$~nm, a $1$~meV gate of width $1$~nm centered at $u_0=7$~nm gives relative conductance changes of $+0.0718\%$ and $-0.0918\%$ at $E=200$~meV for $η=0.5$ and $2$, respectively. Energy--position maps show how asymmetry redistributes the electrical response. Gates centered well inside both curved segments can give different response magnitudes; opposite signs occur near the mouths and farther out. We also find opposite response signs in the fixed-length comparison at the chosen operating point. For equally occupied opposite angular channels, the sublattice imbalances cancel and the conductance responses add. A spatially controlled gate thus probes throat orientation through conductance even when the unperturbed two-terminal measurement cannot distinguish the two configurations.

cond-mat.mes-hall↗

THGFM: Dual-Branch Temporal Heterogeneous Graph Fusion Model

Temporal heterogeneous graphs offer a natural abstraction for dynamic relational systems in which diverse node and relation types co-exist and evolve over time. Learning on such graphs requires jointly modeling cross-type structural heterogeneity and the temporal dynamics of interactions, yet existing methods still struggle to reconcile parameter-efficient cross-type transfer with relation-aware specialization, and typically inject time only as additive features outside the attention kernel. We propose \textbf{THGFM}, a web-scale temporal heterogeneous graph fusion model that addresses both limitations within a unified dual-path architecture. THGFM couples a \textit{Shared-Space Temporal Attention} branch for parameter-efficient cross-type transfer with a \textit{Relational Type-Partitioned Temporal Attention} branch for relation-aware specialization, and integrates them through \textit{Dual-Path Relational--Shared Fusion}, instantiated with \textit{Type-Conditioned Non-Competitive Gated Sum Fusion}: a adaptive mechanism that assigns independent, type-conditioned feature-wise gates to the shared and specialized branches, allowing both to be amplified or suppressed without zero-sum competition. To directly incorporate relative time into the attention score, THGFM further introduces \textit{Rotary Temporal Attention}, which rotates queries and keys by half-phases of relative time before matching. THGFM consistently outperforms baseline graph transformer models on academic graphs benchmarks, delivering a $+3.25\%$ six-task mean gain, with peak relative gains of $+12.37\%$ on OAG-CS PV, $+4.87\%$ on PF-$L_2$, and $+1.18\%$ on PF-$L_1$, and $+4.24\%$, $+3.73\%$, and $+4.61\%$ on OGBN-MAG, HTAG-ArXiv, and HTAG-DBLP, respectively.

cs.LG↗

Fractional Parabolic Partial Differential Equations in Anisotropic Spectral Barron Spaces: Regularity and Neural Approximation

We study fractional parabolic initial-value problems with lower-order drift and potential terms in anisotropic spectral Barron spaces, defined by weighted space--time Fourier $L^1$ norms adapted to parabolic scaling. We prove existence, uniqueness, and maximal regularity with a gain of one derivative in time and $γ$ derivatives in space, where $γ>0$ is the order of the fractional Laplacian. The evolution is defined only for $t\geq0$, whereas the finite-time norm requires a global extension with sufficient temporal Fourier decay. We construct a finite reflected semigroup extension using a Vandermonde system to match derivatives at $t=0$, obtaining temporal Fourier estimates uniform in the semigroup parameter. Combined with Fourier multiplier estimates for the damped principal operator, it yields maximal regularity. Dimension-independent multiplication estimates support a finite regularity bootstrap, while interpolation and sufficient damping absorb the lower-order terms in the base estimate. The a priori estimate and the method of continuity yield maximal regularity without smallness assumptions on the lower-order coefficients. A frequency-localized counterexample shows that a uniform-in-time spatial Barron bound on the forcing does not imply the corresponding two-derivative solution bound, even for the one-dimensional heat equation. Using this regularity, Fourier sampling yields $n^{-1/2}$ approximation rates for the solution in mixed space--time Sobolev norms using shallow networks with suitable activations. Sampling in a product Hilbert space yields a population-level PINN consistency estimate for shallow cosine networks on a bounded cylinder. There exists a single width-$n$ network for which the sum of the squared mixed-Sobolev solution error, the squared $L^2$-norm of the residual for the whole-space fractional equation, and the squared initial-data error is $O(n^{-1})$.

math.AP↗

RedFlow: Redirect Failure into Action-level Corrections for Flow-matching VLA Policy

Reinforcement learning (RL) can improve Vision-Language-Action (VLA) policies from deployment experience, but reward- and preference-based RL primarily identifies desirable behaviors without specifying how to correct failed actions, underutilizing failure trajectories and limiting sample efficiency. Can such corrections be derived from fixed rollouts? Our key insight is that rollouts with different outcomes may contain action chunks executed in similar states, enabling higher-quality chunks to provide locally supported corrective references. Building on this insight, we introduce \textbf{RedFlow}, an offline post-training method for flow-matching VLA policies. \emph{Execution-Context Matching} groups chunks using a compact representation of estimated task progress and robot proprioception. \emph{Quality-Guided Action Redirection} assigns signed chunk-quality scores and aggregates higher-quality chunks into corrective targets, reinforcing high-quality chunks, suppressing low-quality chunks, and redirecting correctable chunks toward their targets. RedFlow requires neither external HIL corrections nor online data collection during post-training. Across four LIBERO suites, RedFlow improves average success from 56.2\% to 68.2\%, outperforming the strongest evaluated offline baseline, AWR (62.3\%), by 5.9 points. Across three real-robot tasks, it improves average success from 56.7\% to 74.7\%. On LIBERO-Spatial, RedFlow reaches 75.8\% success with 1{,}536 fixed rollouts, while the evaluated online methods require 8.7--16$\times$ as many fresh post-training rollouts to reach the same threshold.

cs.RO↗

AI-based scoring systematically underestimates conceptual understanding of linguistically weak students' explanations in physics

Students' explanations of scientific phenomena provide important evidence of their conceptual understanding. However, because conceptual understanding can only be inferred through language rather than observed directly, distinguishing conceptual understanding from linguistic quality represents a fundamental challenge for assessment. This challenge can also be expected to arise when artificial intelligence (AI) is used to score students' text-based explanations. We investigated whether AI-based scoring approaches assess students' conceptual understanding independently of the linguistic quality of their explanations about a phenomenon taken from a physics context. Conceptual understanding scores assigned by nine machine learning-based scoring approaches and two large language model-based scoring approaches were compared with expert-assigned scores for 116 explanations produced by secondary-school students in Germany. Multinomial logistic regression analyses examined whether expert-rated linguistic quality was associated with under- or overestimation of conceptual understanding. Across all eleven AI-based scoring approaches, lower linguistic quality was associated with a greater likelihood of underestimation, although the strength and statistical significance of this association varied between approaches. AI-based scoring approaches may thus systematically underestimate the conceptual understanding expressed in linguistically weaker explanations. This pattern resembles language bias previously documented in STEM teachers' assessment practices. The findings highlight the importance of examining not only agreement with expert ratings but also whether AI-based scoring approaches inadvertently rely on construct-irrelevant information. Addressing language bias is therefore essential for the validity and fairness of AI-supported assessment in physics education and STEM education more broadly.

physics.ed-ph↗

Spinning down neutron-star merger remnants with the Tayler-Spruit dynamo: Global simulations reveal the formation of massive disks and neutron-rich ejecta

Magnetic-field amplification and angular momentum (AM) transport critically shape the secular evolution, lifetime, and electromagnetic signatures of binary neutron-star merger remnants. While the magnetorotational instability can operate in the outer negative-shear regions of the neutron-star remnant and accretion disk, the positive-shear, stably stratified core may instead be susceptible to the Tayler-Spruit dynamo. We present the first global, long-term general-relativistic neutrino-radiation magnetohydrodynamics simulations of a neutron-star merger remnant incorporating the unresolved Tayler-Spruit dynamo through a new mean-field dynamo subgrid prescription. Our axisymmetric simulations starting from a realistic merger remnant show that the Tayler-Spruit dynamo is primarily active in high-latitude regions of the remnant core. The resulting Maxwell stresses redistribute AM on a spin-down timescale of a few hundred milliseconds, substantially flattening the core rotation profile and transferring mass and AM from the outer remnant into the disk. This produces a more massive, extended, and strongly magnetized disk with a low electron fraction, leading to substantially more neutron-rich ejecta. Our results demonstrate that currently unmodelled Tayler-Spruit dynamo action can qualitatively alter the rotational evolution, collapse prospects, disk formation, and multi-messenger signatures of long-lived neutron-star merger remnants.

astro-ph.HE↗

TELLER: Dual-Path Iterative Preference Optimization for Table Entity Linking

Entity linking in tables matches short and ambiguous cell mentions to their corresponding knowledge-base entities. Existing approaches typically rely on data preprocessing pipelines that retain either compact or extensive table content as contextual evidence, and then formulate entity linking as a language generation task for instruction-tuned models; recent systems further incorporate explicit reasoning to disambiguate challenging mentions. However, their training supervision is usually static: fixed preference data cannot adapt to the residual errors of an evolving model, while variations in reasoning length can bias sequence-level preference learning. To address these limitations, we present TELLER: Table Entity Linking through Learning from Errors and Reasoning. We first retrieve and rank Wikidata candidates and retain reduced table evidence in the prompt. The direct-answer path applies iterative direct preference optimization and refreshes its preference data with residual errors from the updated model. The reasoning path uses filtered and compressed chain-of-thought rationales for supervised fine-tuning, followed by our iterative length-normalized regularized preference optimization. On the TableInstruct entity-linking subset, the direct-answer path improves accuracy from 94.35\% to 94.50\%; on the MammoTab V2 evaluation set, it improves accuracy from 87.59\% to 88.20\%. The reasoning path improves accuracy from 92.90\% to 92.95\% on TableInstruct and from 79.09\% to 81.85\% on MammoTab V2, while maintaining high rates of complete reasoning generation. These results show that iterative preference learning benefits both concise entity prediction and explicit reasoning.

cs.CL↗

A Poromechanics-Based Framework for Fully Coupled Reactive Transport and Geomechanics in Porous Rocks

Mineral precipitation and dissolution in rocks alter pore structure, stress, and hydraulic properties, producing tightly coupled chemo-hydro-mechanical processes that are central to subsurface systems. Yet, existing continuum-scale simulators lack a generalizable, mathematically tractable, and physically grounded chemo-mechanics formulation for modeling these coupled processes. To address this gap, this work presents a poromechanics-based framework for chemo-mechanics coupling and integrates it into coupled reactive transport and geomechanics simulations. Building upon classical poromechanics theory, the framework provides distinct descriptions for the stress states of the host rock and pore minerals, in addition to pore fluid pressure, enabling a rigorous and flexible treatment of their individual behavior and mechanical interactions during precipitation and dissolution. As demonstrated by numerical examples, the poromechanics-based method captures the expected mechanical response by translating pore-scale mineral growth into mineralization pressure. By accounting for mineral compressibility, the method also supports more physically realistic representations of deformation and induced stress. Numerical comparisons further highlight the unique capability of the poromechanics-based approach to model effective stress evolution and fracture development without explicitly representing pore geometry or other microstructural features. Overall, this framework offers a versatile, physics-based predictive approach to modeling coupled chemo-mechanical processes in reactive porous rocks and supports future reservoir-scale analyses of engineered subsurface systems.

physics.geo-ph↗

Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety

We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive Open Online Validation and Evaluation), a clinician-led platform collecting blinded pairwise preferences alongside multi-criterion rubric ratings. Clinicians assign scores on a discrete $[-2, +2]$ scale, where negative values indicate clinically unsafe or misleading content. Using 26{,}804 pairwise judgments across outputs from 13 LLMs, contributed by more than 736 clinicians across 28+ countries, we find that clinician preference is a poor proxy for safety-critical performance. Models ranking highly under pairwise preference can still exhibit substantial rates of clinically meaningful failures ($\leq -1$) on dimensions such as \emph{Harmlessness} and \emph{Accuracy}. These failures are unevenly distributed across specialties, creating domain-specific ``no-go zones'' not visible in aggregate rankings or single-number leaderboards. We further analyze contributing factors including prompt length, refusal and escalation behavior, and the relative contributions of safety-critical versus surface-level features. A substantial fraction of preference votes carry no positive safety signal, while feature decomposition shows that surface-level characteristics explain slightly more preference variation than safety-critical rubric differences. Finally, we introduce a clinically adjusted preference ranking combining pairwise preference with rubric-derived feedback, producing a more safety-aware ordering than raw Bradley--Terry strength alone. Our findings support evaluation practices that separate preference from safety, report safety-critical failure rates directly, and incorporate clinically grounded adjustments when ranking LLMs for clinical decision making.

cs.CL↗

dots.tts.edit: Precisely Controlled Speech Editing with a Continuous Autoregressive Model

Speech editing for content creation requires precise control over both what an edit should do and where it should apply. Free-form natural language provides a flexible interface for expressing edit requests, but its ambiguity may leave the intended operation, parameters, or target region underspecified. We study a precise and explicit interface for speech editing: a transcript-grounded structural edit instruction with XML-style tags explicitly specifies typed operations and localizes them to transcript spans or boundaries. This semantic timeline avoids explicit timestamp alignment and provides an externally inspectable contract for compositional edits. We instantiate the interface in dots$.$tts$.$edit, an editor adapted from the continuous autoregressive dots$.$tts foundation model. Four representative speech-creation controls cover lexical content, affective expression, pitch and speaking-rate delivery, and temporal phrasing through text, emotion, prosody, and pause editing. Task-specific data pipelines construct operation- and scope-controlled pairs while retaining source-derived context outside each target region. We further introduce doteBench, a bilingual evaluation suite that measures precise instruction following, local preservation, and audio quality across the four controls and their composition. Experiments show leading overall instruction following and local preservation across its five editing categories, while audio quality remains comparable to existing open-source systems. Across three Seed-TTS-Eval shards, the model shows negligible differences from the base model in zero-shot TTS recognition error rate and speaker similarity.

cs.SD↗

Experimental validation of an open-source low-cost single-camera 6-DOF tracking of floating-body motion in wave tanks

Accurately measuring motion of floating structures in experimental settings is both important and non-trivial. In this paper, we present a single-camera, 6-degree-of-freedom motion-tracking system that is both low-cost and simple to set up for an experimental campaign. The system uses fiducial markers and open-source computer vision tools to estimate the positions and orientations of multi-marker geometries and track them. We validate the accuracy and limitations on a precisely controlled linear actuator with static, regular, and irregular motion. The performance is quantified depending on both the camera-to-marker distance and direction of motion relative to the camera plane. Additionally, we present a practical use case for our system in a wave tank. Although the motion direction perpendicular to the camera plane shows the highest errors, the system achieves sub-millimeter accuracy at short and moderate camera-to-marker distances. For the irregular motion validation, the best cases reproduced the displacements with an RMSE of approximately 0.5 mm. Overall, the results indicate that low-cost single-camera fiducial-marker tracking can provide sufficiently accurate, non-intrusive motion measurements for a range of hydrodynamic laboratory experiments.

physics.ao-ph↗