Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 433 records · Page 24Linked to original sources

The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation

Cooperative AI agents are evaluated against other AIs, yet human cooperation relies on implicit conventions -- shared protocols for reading meaning beyond the literal message -- which AI-AI benchmarks may not capture. We propose the convention gap, the difference between the failure probability predicted from the literal content of communication and the observed failure rate, as a metric of implicit communication. In the card game Hanabi, the finite deck and deterministic hint constraints make this posterior exactly computable. We replayed about 101,000 play actions from three public datasets of human-human (an online Hanabi platform), AI-AI (HOAD), and human-AI (HanabiData) games. The gap was +26.2 percentage points (pp) in human pairs, -0.7 pp in AI pairs, and +16.4 pp in human-AI pairs, and was concentrated on plays of cards that had received no hints (+46 pp in human pairs). Within human-AI play, the literal information available to humans was similar across the three AI partners (mean predicted failure 38-41%), but human failure rates ranged from 14.4% to 34.4% and the gap from +24.1 to +6.2 pp; the partner eliciting the largest gap produced the fewest human failures. Game score carried different information: it depended on each corpus's roster composition, whereas the gap separated human from AI play at the agent level. As a known-answer check, Off-Belief Learning agents, whose convention content is controlled by construction, gave a gap of +1.6 pp at the convention-free level, rising monotonically to +21.7 pp. These results suggest that convention compatibility, rather than AI-AI performance, may predict an AI's effectiveness with human partners.

cs.AI↗

Asymptotic $q,t$-Fuss--Catalan numbers for type $B$

Let $W=W(B_n)$ act diagonally on $\mathfrak{h}\oplus\mathfrak{h}^*$, let $S=\mathbb{C}[\mathfrak{h}\oplus\mathfrak{h}^*]$, let $J\subset S$ be the ideal generated by the $W$-alternating polynomials and $\mathfrak{m}_S$ is the maximal ideal of the origin. For sufficiently large $m$ we compute $q,t$-Fuss-Catalan polynomial $Cat^{(m)}(B_n;q,t):=Hilb(\frac{J^m}{\mathfrak{m}_S J^m})_{det-part}$ and imply $Cat^{(m)}(B_n;1,1)=\binom{n(m+1)}{n}$. For proofs, we work with the $Γ$-equivariant Hilbert scheme $Y_n=nΓ$-$Hilb(\mathbb{C}^2)$, $Γ=μ_2$ and Haiman-type Koszul complex that defines the punctual locus of $Y_n$. Our formula for $Cat^{(m)}(B_n;q,t)$ is derived from a localization computaion for the Haiman-type Koszul complex.

math.CO↗

Sobolev Regularity in Mixed and Isotropic Scales for Vector Fields on the Torus

We study regularity, existence, and uniqueness of solutions to vector fields on the two-dimensional torus in Sobolev spaces of dominating mixed smoothness, which measure regularity separately in the two variables. For constant-coefficient vector fields, we obtain families of mixed smoothness estimates describing how the gain or loss of regularity can be distributed between the two variables. In the nonreal case, a gain of one derivative can be distributed between the two directions, whereas for real irrational coefficients the loss is governed by the irrationality measure of the coefficient. We also establish sharpness below the corresponding arithmetic threshold and describe the rational and Liouville obstructions. For real-valued variable coefficients, direct estimates for the periodic Fourier-mode equations yield mixed smoothness regularity results and their isotropic and classical consequences. We then use a periodic conjugation to the averaged constant-coefficient normal form. Although this conjugation introduces an additional loss in the mixed smoothness scale, it preserves isotropic Sobolev orders and therefore transfers the sharp constant-coefficient isotropic theory to the variable-coefficient setting. In the nonresonant regimes, we also obtain existence and uniqueness under the natural zero-mean compatibility condition.

math.AP↗

DuplexDrama: A Synthesized Dialogue Dataset with Scenarios, Full-Duplex Behaviors, Expressive Speech, and Sound Events

We present DuplexDrama, the first synthesized spoken dialogue dataset that simultaneously covers four dimensions: (i) complete persona and scenario settings; (ii) three full-duplex behaviors (interruption, backchannel, incomplete); (iii) expressive speech with persona-aligned emotion labels; and (iv) script-aware sound events. DuplexDrama is built via a 4-stage pipeline; quality validation on both scripts and synthesized audio confirms its quality. We have produced more than 2,000 hours audio data with a 64-voice timbre pool spanning 13 personas and 5 age buckets; 3.8% of all turns carry at least one full-duplex behavior. This data has been validated through internal full-duplex model training. We will release a curated subset of 6,400 bilingual dialogues (800 h, Chinese ~500 h + English ~300 h) to advance full-duplex spoken dialogue model research. Data samples are available at our demo page and LLM-judge evaluation prompts will be released with the dataset.

cs.CL↗

Wigner entropy below vacuum: physical counterexamples and stability limits

Positive Wigner functions can have less Shannon entropy than the vacuum, even arbitrarily close to the vacuum state. We construct physical counterexamples and identify the competition that controls their entropy: the relative entropy to the vacuum phase-space density can exceed twice the mean photon number. A two-level family exhibits a finite window in which coherence lowers entropy while preserving global Wigner positivity. A complementary construction repairs a remote negative tail with an exponentially small thermal admixture; an explicit mixing weight of $2\times10^{-21}$ preserves a rigorously established entropy decrease. Optimisation over all one-mode Wigner-nonnegative states gives the sharp low-energy scale $E^γ/[\ln(1/E)]^β$, with $γ\simeq0.7412$ and $β\simeq0.5861$. Because $γ<1$, tensor products can have vanishing total energy and trace distance from vacuum while their entropy deficit diverges. We derive the exact half-transmission threshold for universal Shannon-entropy recovery and connect it to companion results showing residual non-Gaussian structure. The counterexamples reveal distinct controls on entropy: coherence sets the local descent, mode number amplifies it, and loss restores the vacuum bound.

quant-ph↗

Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agents

LLM judges are increasingly used to evaluate and improve AI-generated outputs, yet their reliability for complex professional work remains unclear. We study this problem through Vibe Patenting, an end-to-end patent-drafting testbed for AI-agent evaluation. A separately-invoked LLM judge evaluates generated patent drafts and provides structured feedback for iterative revision. Across multiple inventions and drafting-agent configurations, judge-guided revision consistently improves judge-assessed quality, while unguided revision tends to saturate. Notably, iterative judge feedback enables a low-reasoning agent to approach the performance of a substantially more expensive high-reasoning agent. Stronger models and increased reasoning generally improve judge-assessed drafting quality, while domain-specific agentic workflows provide further gains. We validate the judge against independent evaluation by a professional patent attorney and find meaningful but strongly metric-dependent agreement and systematic calibration differences. These results highlight both the utility and limitations of LLM judges as evaluators and optimization signals for complex professional workflows.

cs.AI↗

SignMimic: Robust High-Quality Sign Language Motion Generation via Human-Shape-Oblivious Pose Transfer Guidance

We study the challenge of sign language video mimicking: given a driving video and a single reference frame, synthesize a video where the target signer reproduces the source motion while preserving identity and linguistic form. Prior pipelines entangle rigid motion, non-rigid deformation, and view-dependent completion in a monolithic generator, causing handshape drift and spatio-temporal instability. We present SignMimic, which (i) applies a TNet-based model to study SE(3) rigid canonicalization to stabilize global pose, (ii) performs non-rigid adaptation in a canonical space to preserve fine-grained articulators (hands/face) and coarticulation via NIF2D, and (iii) uses Pose-MAE-style completion before conditional video diffusion. This factorization injects geometric and linguistic priors, yielding shape and spatio-temporal consistency. On several large-scale datasets (ASL 50K, How2Sign, CSL News), SignMimic achieves state-of-the-art-level performance on video quality, identity similarity, and frame continuity while also achieving minimal loss when performing back translation (SLT) on generated videos. Ablations confirm the role of rigid canonicalization, non-rigid adaptation, and completion. Code is available at https://anonymous.4open.science/r/UniSignMimicTurbo-6088; model checkpoints and video examples will be released.

cs.CV↗

CyFM: Cylindrical Optimal Transport for Few-Step Complex-Valued Flow Matching

Complex-valued signals like MRI and audio spectrograms are typically modelled as flat two-channel Euclidean data. The inherited Euclidean metric $dA^2 + A^2 dθ^2$ vanishes at the origin, leaving phase unpenalised exactly where the signal is weakest. We replace it with the decoupled product metric $dA^2 + dθ^2$ on the cylindrical closure $[0, \infty) \times S^1$, which stays non-degenerate at $A = 0$. We measure what this substitution costs and buys. Exact analytical bridges across synthetic fields, fastMRI knee data, and LibriSpeech spectrograms show Cartesian paths induce a heavy-tailed angular velocity distribution (Pareto index $\approx 1$). Under independent coupling, 43%-49% of signal energy falls on paths turning faster than $π$ rad per unit time. Cylindrical paths never reach this speed. We formulate Cylindrical Flow Matching (CyFM) to strictly bound the angular regression target, coupling noise and data via exact minibatch Optimal Transport jointly over whole fields. This coupling reduces few-step generation error by 3%-60%. CyFM achieves lower generative error than the best Cartesian baseline at every step up to $k = 8$ on synthetic fields and speech spectrograms, with all seeds separated. On knee MRI, the single-step advantage is 1.8x. At convergence ($k = 100$), the two geometries show no significant difference. Finally, a prior-only control exposes the cost of flat parametrisation: on synthetic fields, a single Cartesian Euler step performs worse than the unintegrated noise prior (0.376 vs. 0.150).

cs.LG↗

ESAFusion: LiDAR--4-D Radar Fusion via Local Geometric Complementation and Multiscale Adaptive Interaction for 3-D Object Detection

LiDAR--4-D radar fusion combines accurate spatial geometry with motion and reflectivity cues from radar, offering a promising solution for 3-D object detection in complex driving environments. However, sparse radar observations and differences in spatial sampling between the two modalities complicate reliable cross-modal complementation. Moreover, the relative importance of modalities and feature scales varies across spatial regions, making adaptive fusion challenging. To address these challenges, we propose ESAFusion, an evidence-aware and scale-adaptive framework that combines local geometric complementation with multiscale adaptive interaction. Specifically, we introduce an Evidence-Aware Radar Selection (ERS) module to suppress radar clutter using motion and observation-quality evidence while retaining foreground confidence for subsequent fusion. Then, the Pillar-Level Complementary Encoder (PCE) improves cross-modal complementation under mismatched spatial sampling using local geometric support from neighboring LiDAR pillars. We further design an Intra- and Inter-Scale Adaptive Fusion (ISAF) module to adaptively adjust the contributions of different modalities and feature scales in bird's-eye-view (BEV) space. Extensive experiments on the View-of-Delft (VoD) dataset show that ESAFusion achieves the highest mean average precision (mAP) among the compared methods, reaching 74.60% in the Entire Annotated Area and 88.89% in the Driving Corridor. It also attains the highest average precision (AP) for Cyclist among these methods in both regions while running at 19.23 FPS. Evaluations on VoD-Fog further demonstrate robustness under progressively degraded LiDAR observations. The source code will be made publicly available at https://github.com/SenJieHu549/ESAFusion.

cs.CV↗

Optimal location of small favourable regions for Robin eigenvalues with indefinite weights

We consider the positive principal eigenvalue of an elliptic problem with a Robin boundary condition and bang--bang indefinite weight $κ\mathbf 1_D-\mathbf 1_{Ω\setminus D}$, and ask where a favourable region $D$ of prescribed small volume $|D|=δ$ should be located. Put $\varepsilon=δ^{1/N}$ and $τ_δ=α_δ\varepsilon$. We prove that there is a finite threshold $τ_*=τ_*(N,κ)$, independent of the ambient domain, such that $δ^{2/N}Λ_δ(α_δ)\toΛ_{\mathbb H}(τ)$ whenever $τ_δ\toτ<\infty$. If $τ<τ_*$, optimal small regions concentrate at the boundary. If $τ>τ_*$, their concentration centres move to distances much larger than $\varepsilon$ from the boundary, and after recentring and rescaling the favourable sets converge in measure to the whole-space optimal ball. At $τ=τ_*$, the half-space problem admits a compact boundary optimiser as well as minimising sequences escaping to infinity. In the finer regime $τ_δ=τ_*+σ\varepsilon+o(\varepsilon)$, we determine the first-order competition between the boundary and interior configurations. Tangential symmetry of compact threshold optimisers reduces the geometric correction to mean curvature. In particular, every fixed finite Robin coefficient is asymptotically in the boundary regime. Numerical experiments illustrate the boundary--interior transition, the three cases in the first-order selection law, and the curvature-dependent boundary location; for $N=2$ and $κ=1$ they suggest a transition near $τ_*\approx3.2$.

math.AP↗

How broad is that claim? Mapping Generalisation in NLP Research

Generalisations are common in scientific communication, even though they are semantically ambiguous. An automated method is needed to identify and categorise claims according to their level of generalisation, in order help detect an over-reliance on generalisations and possible misrepresentations of scientific findings. We introduce a comprehensive taxonomy of generalisations in the scientific domain, NLPGenX, which labels claims according to their level of generality and framing within the text. We operationalise this taxonomy with an LLM-powered framework, NLPGenA, that automatically classifies sentences from scientific articles into 5 different generalisation classes. We validate our framework with human annotators and use the framework to construct a large-scale dataset of NLP papers annotated according to generality, with auxiliary labels for hedging and vague descriptors (NLPGens). We use NLPGens to analyse the use of generalisations in NLP papers across multiple venues and subdomains, and to examine associations with citation counts, hedging, and vague descriptors.

cs.CL↗

An explicit solution of the five-expert prediction PDE and the exact optimality set of COMB

In this paper, we derive an explicit solution of the stationary prediction with expert advice PDE for five experts. The formula is given in three regions. In the first two regions, it is the four-expert solution plus a single integral with an elementary positive density. In the third region, it is a finite sum of hyperbolic products whose coefficients are determined by one scalar quadrature. Our formula establishes that the direction $(1,0,1,0,0)$ is optimal throughout the ordered sector, and that the COMB strategy $(1,0,1,0,1)$ is optimal only on a lower dimensional subset of the sector (where $x_1=x_2$ and $x_3=x_4$). This disproves the COMB optimality conjecture of Gravin, Peres and Sivan. The verification of the Hamiltonian inequalities is a tedious task, part of which is completed with a computer assisted proof. The verification reduces to 21 scalar inequalities, which we prove using 147 exact rational Bernstein polynomial certificates. The exact certificates and their independent arithmetic checks are included in a supplement to this paper, and a Lean 4 formalization machine-checks the verification and both main theorems, apart from the viscosity characterization.

math.AP↗

Neural-Network Solutions to Real-Space Charge Density and Generalization

The Hohenberg-Kohn theorem establishes that, in principle, the ground state (GS) charge density contains all GS information of a many-electron system, such that all GS observables can be expressed as functionals of the GS charge density. Conventional Kohn-Sham density functional theory requires iterative solution of the self-consistent-field equations at substantial computational cost, motivating the development of deep learning surrogates for electronic structure calculations and, in turn, accelerating computer-aided materials design. Here, we propose AIDEN, an Atomic-Interaction Density Equivariant Network for solving real-space charge density. AIDEN separates the element-dependent one-center density from environment-induced density redistribution and represents the latter through complementary atom- and edge-centered tensor correlations. A continuous low-rank Gaussian decoder then reconstructs the density at arbitrary spatial coordinates while reusing atomic encodings independently of the evaluation grid. AIDEN achieves state-of-the-art accuracy on periodic crystal benchmarks while remaining competitive for molecular systems, and further demonstrates zero-shot transferability across several structurally distinct out-of-distribution case studies. Furthermore, AIDEN provides substantially faster inference than both baseline models and full SCF calculations, enabling efficient charge density reconstruction for large-scale electronic structure calculations.

cond-mat.mtrl-sci↗

Improving the Last-Iterate Guarantees of Anytime Algorithms for Stochastic Monotone Variational Inequalities

We analyze a stochastic algorithm with Halpern-type anchoring for constrained convex-concave problems and monotone variational inequalities. This single-loop and single-call algorithm uses one unbiased sample of the gradient operator at every iteration, to be applicable to monotone games with noisy feedback. With $t$ denoting the iteration counter, we prove an anytime last-iterate convergence rate of $O(t^{-1/4})$ for both the gradient-mapping norm and restricted gap, bypassing the $O(t^{-1/5})$ constrained-anytime bottleneck in the literature. Specializing then to multi-point oracles, we use variance reduction to achieve the $O(t^{-1/2})$ rate with an anytime single-loop algorithm using $2$ samples per iteration. Our results allow constrained problems with a potentially unbounded feasible set; as well as a structured class of stochastic oracles whose variance need not be uniformly bounded.

math.OC↗

Klein Tunneling of Dirac Fermions through Electromagnetic Barriers

The Lorentz covariance of relativistic Dirac equations serves as a fundamental principle underlying the laws of electromagnetism across different inertial frames. Exploiting the covariance, we obtain the general solutions for Dirac fermions under both the in-plane electric $\boldsymbol{E}$ and perpendicular magnetic $\boldsymbol{B}$ fields, which reduce to either a magnetic or electric field in the inertial frame with drift velocity along the $\boldsymbol{E}\times\boldsymbol{B}$ direction. This dichotomy defines the magnetic and electric regimes, separated by the critical field ratio $E/B=v_\text{F}$ with $v_\text{F}$ denoting the Fermi velocity of Dirac fermions. Using these solutions, we revisit Klein tunneling through a heterojunction with generalized electromagnetic potentials. In the magnetic regime, the transmission exhibits oscillations governed by the Fabry-Pérot interference. In the electric regime, perfect transmission occurs at normal incidence in the drifted frame. The interference phase is further analyzed in terms of the solid angles on the Bloch sphere, providing a geometric interpretation of Klein tunneling. Finally, we briefly discuss the relation between the tilting of Dirac cones and the in-plane electric field, establishing the correspondence of the undertilted and overtilted cases to the magnetic and electric regimes, respectively. Our findings reveal the manipulation of Klein tunneling by electromagnetic fields, offering a theoretical basis for designing novel electronic devices.

quant-ph↗

Sharp Bounds on the Mean Efficiency of a Fluctuating Machine

The efficiency of a machine at the scale of thermal fluctuations is random, and the conventional ratio $-W/Q_h$ has no moment of any order: its input heat fluctuates through zero. We work instead with the exergetic ratio $η=W/(W+T_0S)$, which lies in $[0,1]$ pointwise for non-negative dissipation, and ask what the energy budget alone determines about its mean. With $α=T_0\langle S\rangle/W$ and $σ^2$ the relative variance of the dissipation, $1/(1+α)<\langleη\rangle\leσ^2/(1+σ^2)+1/\{(1+σ^2)[1+α(1+σ^2)]\}$, both ends sharp, with no distributional assumption. The floor is Jensen's inequality: fluctuating dissipation raises the mean efficiency above its deterministic value, and the mean alone gives nothing more. The ceiling is attained by an intermittently reversible law, dissipating nothing in a fraction $σ^2/(1+σ^2)$ of realisations, a prediction testable on trajectories. A third moment lifts the floor. Fixed delivered work is not required: when it too fluctuates, the bounds hold with the moments taken on $T_0S/W$, and the thermodynamic uncertainty relation on the work current converts the ceiling into a precision-efficiency frontier, whose zero-variance member is the known bound on a motor's ratio-of-means efficiency, shown here to be unsafe for the mean of the fluctuating ratio. Inside the interval lies the maximum-entropy benchmark $α^{-1}e^{1/α}E_1(1/α)$. Finally the bounds are worked out for a motor with futile cycles, observed until a fixed number of steps is delivered. There the dissipation cannot fall below the reversible cost of that work, and this floor $b$ sharpens the ceiling to $q/(1+αb)+(1-q)/(1+αc)$, $q=σ^2/[σ^2+(1-b)^2]$, $c=1+σ^2/(1-b)$, removing 40-67 per cent of the width. It is saturated when slips are rare: the extremal law is an operating regime, not an idealisation.

cond-mat.stat-mech↗

Discrete q-Hermitian Clifford analysis

We develop a discrete $q$-Hermitian Clifford calculus based on coordinatewise Jackson differences on multiplicative $q$-lattices. The resulting Hermitian Jackson--Dirac operators are nilpotent, factor a $q$-Laplacian, and have occupancy-dependent Euler anticommutators. These give explicit Fischer projectors and homotopies, but the local calculus is not controlled by total degree. We prove a two-face Cauchy--Kovalevskaya theorem for polynomials and extend it to a $q$-analytic Jackson--Fischer class. A divided-power conjugation recovers scalar Euler relations and transfers the classical joint Fischer decomposition and dimension formulas. The Jackson--Fischer completion is a vector-valued $q$-Fock space whose polarized nullspaces have projected reproducing kernels; the joint kernel is obtained by alternating projections. We also determine the linear symmetry group of the multiplicative lattice.

math.CV↗

The Troy Moment: How LLM Agents Adjudicate the Decision Point Under Impossible Tasks, Claimed Authority, and Peer Information

Recent investigations of the July 2026 OpenAI-Hugging Face incident motivate two questions about agent behavior under task failure: when an assigned task becomes impossible, does an agent persist, stop, or escalate, and can observing another agent's behavior change that decision? We study this decision point on ImpossibleBench-derived software-repair tasks with GPT-5.6 Sol, Claude Fable 5.1, and Gemini 3.8 Flash. Each task contains a genuine software defect together with a conflicting test requirement that cannot be satisfied by a behaviorally correct source-code change. If the agent modifies the protected test file, it violates the boundary, which it is not supposed to. Holding the impossible task fixed, we vary what is told to the agent: peer precedent and punishment, a forged authorization claim, instruction wording, and tool friction; we also study three-agent swarms sharing a message board. Around this shared boundary, the models exhibit distinct adjudication policies. Fable emphasizes scope and provenance, Gemini often interprets boundary-relevant cues through a security lens, and Sol largely filters lateral precedent while engaging apparent vertical authority. Our study shows that compliance is not well characterized as a property of a prompt or model in isolation. We propose conflict adjudication, the mapping from information to interpretation to action, as a useful unit for evaluating agent alignment when task pressure, authority claims, tool affordances, and social evidence conflict.

cs.AI↗