Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 397 records · Page 22Linked to original sources

Equilibrium Laws for Julia's Zero and the Hyperbolic Zero of Binary Forms

For a real binary form with no real roots, Julia's zero and the hyperbolic zero are two SL(2,R)-equivariant points in the upper half-plane. We show that they admit closely related equilibrium characterizations, with radial weights given respectively by the hyperbolic tangent and hyperbolic sine of the distances to the roots. This common framework gives geometric criteria for coincidence of the two zero maps. They always agree for binary quartics; for binary sextics they agree exactly when the three upper-half-plane roots form an equilateral hyperbolic triangle, or are collinear with one root the hyperbolic midpoint of the other two. We also obtain a corresponding result for collinear binary octics. The two equilibrium laws further reveal a sharp difference in the influence of distant roots. A strict majority of roots confined to a compact set keeps Julia's zero in a compact set, and the threshold one-half is optimal. In contrast, an escaping minority can force the hyperbolic zero to escape at linear scale. We also illustrate the computational advantages of the explicit formula for the hyperbolic zero.

math.MG↗

Characteristic localization, sharp lifespan asymptotics and global dynamics for a one-dimensional derivative wave equation

We study the Cauchy problem $$ v_{tt}-v_{xx}=\abs{v_t+v_x}^{m}\abs{v_t}^{n}, \qquad v(x,0)=ηφ(x),\quad v_t(x,0)=ηψ(x), $$ on $\R$, with compactly supported profiles $φ\in C_0^2(\R)$, $ψ\in C_0^1(\R)$, exponents $m,n>1$, and amplitude $η>0$. We show that the sign of $P_0=ψ+φ'$ decides the behaviour of small solutions, whatever the sign of $ψ-φ'$. If $P_0$ is positive at some point, the lifespan $T(η)$ obeys explicit two-sided bounds of order $η^{-(m+n-1)}$ for every $η>0$, and $$ \lim_{η\to0^+}η^{m+n-1}T(η)=\frac{2^{n}}{(m+n-1)\,(\max_{\R}P_0)^{m+n-1}} . $$ If $P_0\le0$, the solution is global for every amplitude below an explicit threshold. If moreover $P_0\not\equiv0$, the component $v_t+v_x$ decays at the universal rate $\bigl(2^{n}/((m+n-1)t)\bigr)^{1/(m+n-1)}$, and the gradient of the solution converges uniformly to that of a free wave travelling to the right; if $P_0\equiv0$, the solution is itself such a travelling wave. The analysis rests on a localization property: the zero set of $v_t+v_x$ is invariant along its own characteristics, so that the nonlinear source stays in a slab of fixed width moving with speed one, and each characteristic of the other family is forced only during a bounded time. This yields the global existence, the long-time behaviour and the exact value of the limit. The results extend to the endpoint exponents $m,n\ge1$ and to sources $f(v_t+v_x)g(v_t)$, for which $T(η)$ is asymptotic to an Osgood-type integral.

math.AP↗

Carroll-Cotton Tensors and Gravitational Radiation at Null Infinity

We provide a first-principles definition of Cotton tensors on conformal Carroll geometries by gauging the conformal Carroll algebra. On the mathematical side, this analysis yields a conformally covariant object designed to quantify deviations from conformal flatness in Carrollian geometries. On the physical side, it offers a powerful geometric tool to describe gravitational radiative degrees of freedom at the conformal boundary of four-dimensional asymptotically flat spacetimes. Unlike the Bondi news, whose fully-covariant definition requires amendment by supplementary data that we identify with the Geroch tensor, the Carroll--Cotton tensors offer the natural geometric objects to quantify deviations from stationarity due to the emission of gravitational waves.

hep-th↗

A General Superconvergence Result for Cubature on Tesselated Polytopal Domains in Two and Three Dimensions

Cubature rules, which approximate definite integrals as a linear combination of a set of function values, are ubiquitous and necessary for computational methods in the physical sciences. A superconvergence result for cubature rules on polytopal domains in two and three dimensions is developed, whereby a rule that is exact for all bivariate polynomials of a fixed even degree realizes an extra order of convergence under a decrease in the spacing between nodes.

math.NA↗

Optimal error estimates of sequential finite element method for nonlinear thermo-poroelasticity problems

This study introduces and analyzes a three-step sequential decoupling algorithm designed to address nonlinear, fully coupled quasi-static thermo-poroelasticity systems incorporating convective transport. The finite element method is employed for spatial discretization and the backward Euler method for temporal discretization. The proposed sequential method has a higher computing efficiency than the fully implicit nonlinear numerical scheme, since it does not require any internal iterations. The well-posedness of the numerical solution is discussed by introducing a cut-off operator and the stability analysis of the algorithm is performed. Rigorous analysis yields optimal convergence order estimates for both spatial and temporal discretizations. In order to confirm the theoretical results and the effectiveness of the suggested approach, numerical experiments are finally carried out.

math.NA↗

Parameter-uniform Robin uniqueness on large dilations

Berestycki and Graham proved large-dilation uniqueness for bounded positive solutions of \[ -Δu=f(u)\quad\hbox{in }κΩ, \qquad u+α\partial_νu=0\quad\hbox{on }\partial(κΩ),\] when $α$ is fixed, and remarked that the dilation threshold should not depend on $α$. We show that it does not, including at the Dirichlet and Neumann endpoints, for possibly unbounded uniformly $C^{2,γ}$ domains. The half-space linearizations have a common positive spectral gap over the compactified boundary parameter. The Dirichlet end requires a separate compactness argument because the Robin coefficient diverges there. After rescaling by its inverse, the equation has a harmonic half-space limit. A Liouville lemma rules out a nonzero limiting trace, and the resulting endpoint compactness, together with the half-space gap and localization, yields the uniform uniqueness statement. When $\partialΩ\neq\varnothing$, we further obtain convergence of the spectral bottom to its half-space value with error $O(κ^{-1/2})$. For bounded $C^{4,γ}$ domains the boundary layer also has a first mean-curvature correction, with remainder $O(κ^{-1-γ}+κ^{-2})$ on each fixed boundary strip.

math.AP↗

Symmetry-dependence in Rounding of a Convex Body

The symmetry measure of a convex body $S\subset\mathbb{R}^n$ is given by: $\mathrm{sym}(S):=\max\{α\ge0:\text{ there exists }x\in S\text{ such that }-α(S-x)\subseteq S-x\}$, where such an $x$ is called a Minkowski center. We prove that every convex body $S$ admits a $\sqrt{\frac{n}{\mathrm{sym}(S)}}$-rounding of $S$, namely, there exists an origin-centered ellipsoid $E$ and a center $c$ such that $E\subseteq S-c\subseteq\sqrt{\frac{n}{\mathrm{sym}(S)}}\,E$. This result was conjectured in 2005 by Belloni and Freund. As special cases, this recovers an $n$-rounding of $S$ (since $\mathrm{sym}(S)\ge\frac{1}{n}$), and a $\sqrt{n}$-rounding when $\mathrm{sym}(S)=1$. In the case when $S$ is a polytope given as the convex hull of points, the desired rounding is produced by a regularized minimum-volume covering ellipsoid problem where the regularization is with respect to the Minkowski center. Similarly, when $S$ is a polytope given as the intersection of halfspaces, such a rounding is produced by a regularized maximum-volume inscribed ellipsoid problem. In both of these cases, the rounding can be computed by first solving a linear optimization problem (to compute $\mathrm{sym}(S)$ and a Minkowski center), and then solving a convex optimization problem with a logarithmic determinant objective, second-order cone constraints, and one semidefinite cone constraint. We also show that the factor $\sqrt{\frac{n}{\mathrm{sym}(S)}}$ is nearly tight in its dependence on dimension and symmetry. When $\frac{n+1}{1+\mathrm{sym}(S)}$ is an integer, we show by explicit construction that the factor $\sqrt{\frac{n}{\mathrm{sym}(S)}}$ is tight. In the more general case, for every dimension $n$ and every admissible symmetry value, we construct a polytope $S$ for which every rounding factor is at least $\sqrt{\frac{2}{3}}\sqrt{\frac{n}{\mathrm{sym}(S)}}$.

math.OC↗

Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

On-policy distillation (OPD) offers dense token-level supervision as an alternative to the sparse outcome-level advantages of reinforcement learning with verifiable rewards (RLVR). However, the teacher scores student-generated trajectories that are inherently off-policy for it, so the reliability of its supervision, and hence the source of the student's improvement, remains unclear. We quantitatively analyze teacher supervision during OPD training and find substantial noise whose prevalence increases with teacher scale. Surprisingly, the student policy is insensitive to such noise, converging to comparable performance regardless of whether noisy supervision is retained or removed. Does OPD distill at all? By analyzing what drives its gains, we find that learning concentrates on low log-probability tokens, and using a single fixed negative advantage matches the performance of teacher-provided ones. This suggests that OPD works largely by suppressing low log-probability tokens, which requires no teacher. These findings motivate On-Policy Self-Adaptation (OPSA), a supervision-free method using entropy-adaptive negative advantages. It assigns stronger learning signals to high-entropy positions, suppressing tail tokens, and evenly redistributing probability mass among head tokens. Compared with the base \texttt{Qwen3-1.7B}, OPSA improves Avg@32 by 35.41 points on AIME24, corresponding to a 263\% relative gain, and more than doubles Pass@32 across all three benchmarks. It also outperforms OPD by 16.77 points in Avg@32 on AIME24. Extensive experiments and analyses across model families and tasks further demonstrate its effectiveness and generalizability.

cs.LG↗

Generalized Graph Compositions with Applications to Difference Graphs of Finite Groups

The difference graph $D(G)$ of a finite group $G$ is obtained from the edge difference between its intersection power graph and power graph, after deleting isolated vertices. This graph has already been studied, with sufficient conditions for connectedness and a diameter bound $6$ for finite groups satisfying those conditions. We use generalized graph composition to reduce $D(G)$ to a graph $B(G)$ on the cyclic subgroups of $G$, so that connectedness and diameter are determined by the subgroup structure of $G$. We obtain a general criterion for the non-emptiness of $B(G)$ in terms of branching subgroups and b-normality, and characterize its connectedness for finite $p$-groups, non-cyclic finite abelian groups, and non-abelian groups with both trivial and non-trivial center. Combined with the previously established cyclic-group case, this gives a complete characterization of non-emptiness and connectedness of difference graphs for all finite groups. The successive structural cases lead naturally to the sharp diameter bounds $2,3,4,$ and $5$. For centerless non-abelian groups, connectedness is governed either by a unique branching subgroup or by an auxiliary graph $\mathcal A(G)$; in the latter case \[ \operatorname{diam}\mathcal A(G)-1 \leq \operatorname{diam}B(G) \leq \max\{4,\operatorname{diam}\mathcal A(G)+1\}, \] and both bounds are sharp.

math.GR↗

Calibration is the Bottleneck: An Action-Class Diagnostic of Multi-Turn Tool-Calling

Multi-turn tool calling is a core evaluation scenario for large language model (LLM) agents. On public tool-calling benchmarks, open-weight models now approach or even surpass closed-source frontier models in aggregate accuracy. However, this metric averages over many different multi-turn situations and obscures whether progress is balanced across them. We propose an action-class-oriented diagnostic framework that decomposes multi-turn failures into two orthogonal modes: action-class miscalibration and action-execution failure. The framework operates over a four-class action space (TOOL_CALL/ASK/REFUSE/CONFIRM) and introduces a self-revealing upper bound Acc <= GAR (Gold Action Recall); the two modes show up as bound violation (Acc > GAR, exposing state-grader masking of miscalibration) and large bound slack (GAR >> Acc, localizing execution failure within TOOL_CALL). We validate it on a panel of tool-calling models across multiple multi-turn benchmarks. Across our panel, the diagnostic reveals action-class miscalibration as a substantial failure mode the state grader cannot see. This gap inflates standing for heavily tool-trained families, which our diagnostic separates from families with context-appropriate action choice. Calibration is reshapable through context-only perturbations, but the reshape is heterogeneous: a single perturbation moves accuracy in opposite directions across families (up to +11.5 vs -21.0 pp on the same scenario), and its effect further depends on the perturbation mechanism. We argue that multi-turn tool-calling evaluations should supplement aggregate accuracy with action-class diagnostics that expose what the model actually does in each scenario.

cs.CL↗

Parabolic Homotopy Colimits and Coxeter Descents

Let $G$ be a compact, connected, simply connected semisimple Lie group with Weyl group $W$ and simple reflections $S$. For a simplicial complex $\mathcal K$ on $S$, form the homotopy colimit $X_{\mathcal K}(G)=\operatorname*{hocolim}_{I\in\mathcal K}G/G_I$ of standard partial flag manifolds. We compute its integral homology. If $\operatorname{Des}_R(w)$ is the right descent set of $w\in W$ and $\ell(w)$ its Coxeter length, then $$ H_n(X_{\mathcal K}(G);\mathbb Z)\cong \bigoplus_{w\in W}\widetilde H_{n-2\ell(w)-1}(\mathcal K_{\operatorname{Des}_R(w)};\mathbb Z). $$ Thus the induced subcomplexes of $\mathcal K$ supply the topological data, while the Weyl group determines which subcomplex occurs and the Schubert-degree shift. The proof gives a chain-level splitting and an integral Morse reduction. We derive homotopy detection, duality and rigidity results, and recover polyhedral products, matroid--Tutte formulas, and the adjoint sphere as special cases.

math.CO↗

Fourteen and fifteen lonely runners

We prove the Lonely Runner Conjecture for fourteen and fifteen runners. Our proof combines a stronger bound on the speed product in a primitive counterexample with exhaustive computations modulo primes. To obtain the bound, we work with a projected lattice basis and use cases of the conjecture with fewer runners to bound partial sums of its squared Gram-Schmidt lengths. For fifteen runners, the product bound derived from the work of Malikiosis, Santos, and Schymura gives a logarithmic threshold of about $810$; ours reduces this to about $415$, making the computation feasible. For fourteen runners, the new bound reduces the required number of primes from $111$ to $61$. The computations start with a complete two-branch covering search, followed by binary lifting. Since $14$ and $15$ are composite, the polynomial argument used in earlier work does not apply directly. We finish the fourteen-runner case by a direct search. For fifteen runners, we use the factorisation $15=3\cdot5$ and a shifting argument when all but a few speeds share a common divisor. The code and certificates are publicly archived.

math.CO↗

Modern Transformers Are Implicit Hybrids: From Functional Differentiation to Principled Hybrid Architecture Design

Hybrid architectures combining Full Attention (FA) and Linear Attention (LA) are increasingly prominent, yet their allocation remains heuristic. We seek an evidence-grounded basis in head-level functional organization learned by RoPE-based Transformers. Behavioral probes do not yield a complete taxonomy, so we propose two intervention metrics: RoPE Frequency Importance Score (RFIS), measuring how each frequency affects a head's attention distribution, and RoPE Positional Dependence (RPD), isolating dependence on rotary positional modulation. On Qwen3-series models and Llama3.1, RFIS suggests and RPD verifies a complete taxonomy of retrieval and positional heads separated by a salient mid-low-frequency band. Controlled Transformers show that this boundary follows the training-length positional scale; we term it the Global Positional Band (GPBand). The analysis suggests a potential cause of zero-shot length-extrapolation failure and yields two principles: positional modeling should operate only locally, with global access through position-independent retrieval; and both functions should be assigned at head granularity with layer-specific allocation. We instantiate them in Head-wise Hybrid Architecture (HwH), using NoPE FA for global retrieval and LA for local positional modeling. With an FA-to-LA ratio below 1:3, HwH retains strong language modeling and commonsense reasoning while improving retrieval and substantially strengthening zero-shot long-context extrapolation over Transformer, LA, and a layer-wise hybrid baseline. Ablations validate both principles and component roles, highlighting principled hybrid architecture design as a promising route toward future foundation models.

cs.LG↗

F-sets of arbitrary finite width

Ferraguti and Micheli conjectured that non-trivial $F$-sets of every prescribed width exist over each finite field. We prove the finite-width part of their conjecture over $\mathbb{F}_q$ for every $q\neq2,3$. Fix a suitable prime $\ell$. For a degree bound $D$, we consider the saturated family of all irreducible polynomials $g(X^{\ell^j})$ whose core $g$ has degree at most $D$. A factor-descent lemma shows that every irreducible factor arising from a shifted difference belongs to the same family and has strictly smaller core degree. The saturated family is therefore an $F$-set of finite width. Dirichlet's theorem for $\mathbb{F}_q[X]$ produces core ladders of arbitrary length, and Kummer lifting reproduces each ladder at infinitely many scales. These parallel ladders force the width to be sufficiently large; an appropriate tail of the nullity filtration then has the prescribed width.

math.NT↗

Foundations of Many-Body Theory of Quantum Unified Statistics: Green functions and Linear Response Theory

We develop a comprehensive many-body theory for systems of particles obeying quantum unified statistics, or quons. After exploring the properties of the Fock space of this system, we formulate a systematic S-matrix expansion and a generalized Wick's theorem. A consistent Green-function formalism is constructed at both zero and finite temperatures, accompanied by a generalized Wick's theorem appropriate for infinite-statistics operator algebras. Within this framework, we establish diagrammatic rules for interacting quon systems. Employing the random phase approximation, we derive the dielectric function and reveal the emergence of anomalous plasmon modes that have no direct counterpart in conventional Bose or Fermi systems. We further analyze the ground-state energy, energy-loss function, generalized Thomas-Fermi screening wave vectors, and Friedel oscillations, elucidating how infinite statistics qualitatively modifies collective behavior and screening properties.

cond-mat.stat-mech↗

SCoCaT: Success Conditioned Constrained Reinforcement Learning for Spacecraft Docking

Termination-based constrained reinforcement learning is attractive for safety-critical robotic deployments: it avoids online optimization at inference, scales easily to many constraints via a single scalar per constraint, and is simpler to implement than commonly used Lagrangian methods. Instead of pricing violations through summed cost penalties, this approach makes violations structurally unprofitable by shortening the effective horizon for each violation. We identify a structural failure mode of this method class on terminal-navigation tasks: reaching a precise goal configuration while satisfying safety constraints that tighten along the final approach. When the goal sits inside the region close to where the constraints become active, the survival-weighted objective makes dwelling outside the goal region strictly preferable to entering, producing high constraint compliance with low task completion. We formalize this pathology and show that a minimal augmentation to off-the-shelf RL algorithms like PPO resolves this ``feasibility collapse''. We empirically demonstrate that adding a dense per-step success signal via an auxiliary value critic improves the task completion rate while maintaining safety-critical constraint compliance. Validation across two representative spacecraft platforms: a 6U-CubeSat spanning the mass and degree-of-freedom envelope of operational proximity operations, and a floating platform testbed for zero-shot sim-to-real transfer in our laboratory, supports the generality of these findings.

cs.RO↗

Hopf quotients of the infinite-dimensional Gaussian pyramid

We study the infinite-dimensional Gaussian pyramid and its quotients by the global sign flip and the $U(1)$-Hopf action. We resolve affirmatively a long-standing problem posed by Tomohiro Fukaya around 2014: these three limiting geometries are pairwise non-similar, meaning that no positive rescaling makes any two of them coincide.

math.MG↗

Centering Drives Normalization Gains: Price-Offset Nuisances in Cross-Sectional Return Prediction

Cross-sectional return prediction from raw intraday bars is sensitive to each instrument price level, an additive nuisance under a return-ranking hypothesis. We test whether removing this offset, rather than rescaling amplitudes or changing the encoder, explains gains on a point-in-time CSI~300 five-minute panel. We evaluate eight parameter-matched encoders with and without RevIN normalization; a parameter-free ladder then separates identity, scale-only, centering, last-value referencing, differencing, and standardization across all fields and restricted channels. Centering drives the reliable effect, while scale-only normalization does not help. All eight paired effects are positive and survive Holm correction on raw rank IC, after style residualization, and after further residualizing on short-term reversal. Among six stronger encoders, gains of 0.0376-0.0567 exceed the 0.0109 spread of normalized IC (0.0830-0.0939). Price-only standardization retains 93--101% of the all-field gain. These results place the main effect in transformed price-channel offset removal rather than amplitude scaling or encoder choice.

cs.CE↗