Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 613 records · Page 34Linked to original sources

Polynomial growth of complex polynomial Bohnenblust--Hille constants

For complex \(m\)-homogeneous polynomials, let \(D_m\) denote the optimal dimension-free constant in the polynomial Bohnenblust--Hille inequality. We prove that, for every \(B>1/2\), there exists \(K_B>0\) such that \[ D_m\le K_B m^B,\qquad m\ge1. \] More precisely, we obtain the square-root-scale estimate \[ D_m\le \sqrt m\,\exp\!\bigl(C\sqrt{\log m}\,\log\log m\bigr). \] The proof combines coefficient-preserving degree reduction with a weighted bootstrap, separating balanced and dominant degree profiles. We also determine the critical linear-dimensional asymptotics: \[ D_{m,n_m}\longrightarrow 2 \qquad\text{whenever}\qquad \frac{n_m}{m}\longrightarrow1, \] and consequently \(\liminf_{m\to\infty}D_m\ge2\). Further applications give two-sided bounds for homogeneous Sidon constants and a quantitative logarithmic remainder in the multidimensional Bohr-radius asymptotic.

math.FA↗

Critical Sobolev thresholds for openness and discreteness of Monge-Ampere gradient mappings

Let $Ω\subset\R^n$ be a domain and let $u\in W^{2,n}_{\loc}(Ω)$ satisfy \[ \det D^2u\geqδ>0\qquad\text{a.e. in }Ω. \] Guerra and Tione asked whether the gradient mapping $Du$ must be open and discrete. We give a negative answer in every dimension $n\geq4$, in a form that isolates the sharp regularity mechanism. First, a local Pogorelov model is chosen to solve the exact equation $\det D^2u=1$ a.e. It is convex, belongs to $C^{1,1-2/n}\cap W^{2,n}_{\loc}$, and its Hessian is positive definite off a line. Nevertheless, its gradient collapses that line to one point. In fact the gradient is neither open nor discrete, its branch set is exactly the collapsed line, and its Jacobian equals one a.e. We then develop a $k$-dimensional version of the construction. If $1\leq k<n/2$, there are convex potentials with a $k$-dimensional flat contact set, uniformly positive and bounded Hessian determinant, and exact critical exponent \[ p_{n,k}=\frac{n(n-k)}{2k}. \] Their Hessians belong to $L^p_{\loc}$ precisely for $p<p_{n,k}$, lie in the weak endpoint space $L^{p_{n,k},\infty}_{\loc}$, and fail to belong to any finite-index Lorentz endpoint. This matches the critical minimum-set theorem of Collins and Mooney. At the regularity required in the question, the construction produces branch sets of every integer dimension $k<n/3$. We also give a topological reformulation of the problem. In the convex branch, gradient fibers are exactly contact sets with supporting affine functions, and openness and discreteness are equivalent to strict convexity. For a general gradient under the hypotheses above, critical Sobolev mapping theory already supplies continuity and sense preservation. The question asks whether the monotone factor in the Eilenberg--Whyburn monotone--light factorization is trivial.

math.AP↗

Characterization of $\mathfrak{k}$-highest weight and $\widehat{\mathfrak{g}}$-dominant tableaux via $1$-$0$-slack recording tableaux in the quantum Littlewood-Richardson rule and lattice points in flagged hive polytopes

We have previously, for a given positive integer $n$, explicitly characterized by certain linear inequalities the $\mathfrak{k}$-highest weight tableaux of shape length $2n$ or $2n-1$ produced by $1$-$0$-slack recording tableaux in the quantum Littlewood-Richardson (LR) rule. We now extend that characterization for lower shape lengths for $n$ odd. Then using the composition of promotion operators defining the Naito-Suzuki-Watanabe bijection between $\mathfrak{k}$-highest weight tableaux and $\widehat{\mathfrak{g}}$-dominant tableaux, we also explicitly characterize by certain linear inequalities the corresponding $\widehat{\mathfrak{g}}$-dominant tableaux when the shape length is $2n$ or $2n-1$. In addition, when the given $n$ is $3$ or $ 4$ the $\mathfrak{k}$-highest weight and the $\widehat{\mathfrak{g}}$-dominant tableaux via $1$-$0$-slack recording tableaux are characterized by linear inequalities for any shape length. Since recording tableaux in the quantum Littlewood-Richardson rule are in natural bijection with Littlewood-Richardson-Sundaram (LRS) tableaux, we relate our results on the inverse quantum LR rule with other two bijections for the Naito-Sagaki conjecture and establish bijections between $\mathfrak{k}$-highest weight tableaux, $\widehat{\mathfrak{g}}$-dominant tableaux and the lattice points in a (disjoint) union of flagged hive polytopes.

math.CO↗

J-Miner: Recovering the Decision Logic of Fine-Tuned LLM Classifiers as Compact Rules

Task-fine-tuned large language model (LLM) classifiers acquire task-specific decision knowledge, but this knowledge remains implicit in distributed internal computations, making their decision logic difficult to interpret. We introduce the Executable Decision Compression (EDC) framework and propose J-Miner, which mines vocabulary-named variables from internal readouts and learns rules shared across inputs to produce executable explanations. Analysis reveals that a small set of these variables captures much of the classifier's decision behavior, holding for both varying parameter scales within a family and distinct families. Across six binary tasks, a rule using just one variable reproduces 76.7% of source-classifier decisions on average, rising to 88.8% with 16 variables. Most of the decision information retained by these variables comes from internal activations beyond literal surface matching. A lightweight text reader predicts the variable states, allowing the same fixed rules to execute independently of the source classifier.

cs.LG↗

PTXBench: Benchmarking and Adapting LLMs for GPU Kernel Optimization with Architecture-specific PTX

We introduce PTXBench, a benchmark for evaluating and adapting large language models (LLMs) to use architecture-specific PTX for GPU kernel optimization. PTXBench measures functional correctness, whether selected target instructions execute at runtime, and speedup over frontier libraries across GEMM and attention workloads on H100 and B200 GPUs. Our evaluation shows that architecture-specific PTX capability remains uneven: success rates fall substantially on complex attention backward workloads, and executing the target instructions does not necessarily translate into competitive performance. No evaluated model consistently matches frontier libraries across the suite. We further adapt Qwen3.6-27B using supervised fine-tuning. Repair-conditioned training improves several tasks, but generalization remains uneven; data coverage, balance, and the quality of the reasoning teacher matter in addition to dataset size. PTXBench provides an auditable testbed for measuring and improving LLMs' ability to exploit evolving GPU architectures.

cs.CL↗

Near-unit-root persistence of symmetric and asymmetric stable autoregressive sequences

For an AR($1$) sequence with coefficient $0 < a < 1$, the probability of staying above zero decays exponentially in the number of steps; at the unit root $a=1$ it decays polynomially, with order $n^{-1/2}$ for symmetric increments. We study this transition for AR($1$) sequences driven by symmetric and asymmetric $α$-stable innovations. We write $Λ(a,α)$ for the symmetric exponential persistence rate and $Λ(a,α,ρ)$ for the rate of a strictly stable law with positivity parameter $ρ$. The entire chain admits an exact representation through a single stable Lévy process observed on a geometrically expanding time grid. In the symmetric case, comparison with continuous half-line survival gives $Λ(a,α) \leq \fracα{2}\log{(1/a)}$. For $0 < α< 2$, this bound disproves the stable specialization of a conjecture of Hinrichs, Kolb and Wachtel concerning regularly varying innovation tails. Combining stable closure under subsampling with a monotonicity coupling yields a lower bound of the same near-unit order. This proves $Λ(a,α) \asymp \log{(1/a)}$ as $a \uparrow 1$ and shows that $Λ(a,α)/\log{(1/a)}$ converges to a limit in $(0,α/2]$ equal to the supremum of that ratio over $0 < a < 1$. A Lamperti transformation reduces identification of this constant to a dense-sampling problem for a stationary stable Ornstein-Uhlenbeck process. Existing Gaussian theory determines the sharp value at $α=2$; for $0 < α< 2$, rescued crossings between observations remain the obstacle. For asymmetric strictly stable innovations with $α\neq 1$ and admissible $ρ\in(0,1)$, the corresponding bounds are $Λ(a,α,ρ) \leq α(1-ρ)\log{(1/a)}$ and $Λ(a,α,ρ) \asymp \log{(1/a)}$. At $α=2$, only the symmetric case $ρ=1/2$ occurs.

math.PR↗

Difference-in-Differences Models in the Presence of Time-Varying Mediators and Lagged outcomes

We study difference-in-differences (DiD) designs in which a binary treatment changes an endogenous time-varying (continuous, discrete, or mixed) mediator that in turn affects an outcome. We allow for the inclusion of lagged outcomes in the model. Under our model assumptions, we show that the usual DiD estimand mixes the average direct effect on the treated, the average indirect effect, and a trend bias term. A two-way fixed effects (TWFE) regression that controls for the mediator does not recover the average direct treatment effect on the treated. We show that a DiD estimand conditional on the observed mediator path identifies the conditional average direct effect for treated units at that path, and that averaging over the treated path distribution identifies the average direct effect even when unconditional parallel trends fails. A stable average mediator effect assumption helps recover the average mediator and indirect effects. The framework extends to multivariate mediators, nonlinear DiD, and multiple-treatment-period settings. Existing doubly robust estimators can be used to conduct inference. Revisiting the effects of railroad access on agricultural land values, the specification yields a positive direct component not mediated by measured market access, while the corresponding indirect component is small and imprecise.

econ.EM↗

Cluster Representation of Renormalization Group Transformations and a Rigorous Proof for Convergence of the RG-Flow of the Ising Model to Trivial Fixed Points away from Criticality

Many rigorous results in the modern era of statistical mechanics have been obtained through geometrical representations. A much studied tool in this setting is the random cluster representation of lattice spin models. Another area of statistical mechanics that is of great interest but lacking rigorous results is the theory of the renormalization group. This paper investigates the idea to find a common ground between these two concepts in order to obtain rigorous results on the renormalization group flow. A need for negative cluster-weights weakens the success of this approach. Nevertheless, we managed to establish a relation between the scaling limit of the renormalization group flow of the nearest-neighbour Ising model away from criticality with a simple one-dimensional dynamical system that follows the cluster connectivity of the renormalization group transformation. This allows for a rigorous proof of the convergence of this flow to the zero- and infinite-temperature fixed point respectively for a large family of renormalization group transformations. Explicit results will be established on $\mathbb{Z}^2$ followed by a discussion of the available generalizations to higher dimensions.

math-ph↗

Molecular Implementation of the Machine-Learned Skala Exchange-Correlation Functional in CP2K through GauXC

Machine-learned exchange--correlation (XC) functionals offer a route to improve Kohn--Sham density-functional theory without incurring the cost of explicitly correlated electronic-structure methods. Their use in production simulation codes, however, requires a well-defined mapping between the learned model and the host-code density representation. We formulate and implement a Skala-1.1 interface in CP2K through the external GauXC library. CP2K supplies the geometry, Gaussian basis, spin-resolved atomic-orbital density matrix, and communicator, while GauXC evaluates the XC energy, atomic-orbital potential matrix, and available nuclear derivatives. The interface accepts both all-electron and valence-only density matrices. The latter may arise from separable dual-space pseudopotentials or molecular effective-core potentials. Implementation errors are isolated from functional differences by comparing the Perdew--Burke--Ernzerhof (PBE) functional evaluated through GauXC with native CP2K PBE. The resulting interface gives consistent energies, forces validated against finite-difference total-energy checks, and force-based molecular-virial diagnostics for representative molecular cases. The dietGMTKN55 benchmark suite is evaluated with an all-electron Gaussian augmented plane-wave treatment for elements up to bromine and def2 effective-core potentials for the heavier elements. The resulting aggregate mean absolute deviation of 1.255 kcal/mol is within 0.012 kcal/mol of the corresponding Skala reference value of 1.243 kcal/mol. This work establishes a validated molecular implementation of Skala in CP2K through GauXC.

physics.chem-ph↗

What is Missing from AI Post-Training AI: An Empirical Analysis

Large language model (LLM) agents can now post-train an LLM end-to-end, raising the prospect of recursive self-improvement (RSI). Yet this progress is measured by aggregate benchmark scores, which cannot tell whether an agent executes a fixed plan well or strategically revises the plan when it fails. We separate these two capabilities: execution-level capability, iterating within an established training strategy, and strategy-level capability, revising that strategy as experimental evidence accumulates. Analyzing 1,338 post-training trajectories of frontier agents, we find that agents reliably execute post-training but lock into a default strategy, which follows the agent rather than the task, and only 2.1% of transitions between adjacent training runs ever change strategy. We then test whether the agent lacks experience, reasoning, or the decision to switch. (1) Experience improves execution but not the strategy. (2) Additional reasoning compute yields front-loaded gains on easier tasks but refines, rather than revises, the committed strategy. (3) Human review before training changes which strategy the agent locks into, not whether it locks in, whereas a single mid-run instruction outperforms the agent's own continuation by up to 17.44 points under the same budget. In conclusion, what the agent lacks is the decision to reopen a committed strategy and try another one. Realizing RSI therefore calls for interaction protocols and training signals that make strategy revision an explicit, rewarded decision.

cs.AI↗

Twisted magnon frequency combs in ferromagnetic nanorings

We systematically investigate the emergence of twisted magnon frequency combs (tMFCs) and their higher-order modes arising from strong nonlinear coupling between vortex-core gyration and azimuthal spin-wave modes in ferromagnetic nanorings. The comb spacing is set by the gyrotropic frequency, which is controlled by both the size of the central hole and external magnetic fields. Remarkably, for the larger hole diameter (50 nm), an additional magnon mode emerges, leading to additional tMFC families. We also demonstrate that the selection rules still hold for different nanorings. In addition, the external in-plane magnetic field provides an effective means to tune the tMFC, while the response strongly depends on the nanostructure geometry. The nanodisk shows an approximately symmetric response under field reversal, whereas the response of nanorings depends strongly on the size of the central hole. For a small hole diameter (5 nm), the low-field response becomes asymmetric, and the tMFC spacing increases with field magnitude over the higher?field branches. A larger hole diameter (50 nm) raises the gyrotropic frequency, yielding a sparser sideband structure near the drive frequency. Our results show that ferromagnetic nanorings support geometrically and magnetically tunable tMFCs.

cond-mat.mes-hall↗

Learning how to Forget: Fine-tuning for Long-Context Sparse Attention

A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language models without excessive hardware budgets. We provide a new method for fine-tuning models with sparse attention. It works for any KV cache policy, runs on a moderate hardware budget (e.g., a single Nvidia A100 GPU with 40 GB RAM), and allows the model to co-adapt with the policy, often outperforming models trained with exact attention (sequence parallelism). We also provide an efficient implementation of H2O sparse attention (the leading policy in our experiments) with dedicated scaled dot product attention kernel support. KeysAndValues (https://github.com/awslabs/keys_values), a new open source library for long-context inference and fine-tuning, provides easy-to-use and performant code for all methods discussed here.

cs.CL↗

Optimal Skill Selection for LLM Agents with Provable Bicriteria Guarantees

Loading reusable skill documents into a bounded context window has become a primary way large language model (LLM) agents acquire task-specific capabilities, which makes skill selection a first-order determinant of task performance and token cost. Yet current agents score skills independently by semantic relevance and assemble the set by top-$k$ or greedy packing, with no quality guarantee or cost awareness on the selected set. Redundant or poorly chosen skills then waste scarce context tokens and can even degrade performance. In this paper, we present a theory-grounded and practical framework for budgeted skill selection. We give the first model of how skill sets shape execution outcomes, capturing complementary capability coverage and diminishing returns from redundancy through a monotone submodular benefit, while accounting for context degradation with a linear token penalty under a hard budget. Based on this model, we develop Best Prefix Selection (BPS), a polynomial-time algorithm, and prove, to our knowledge, the first performance guarantee for skill selection: a bicriteria $(1-1/e,1)$ approximation whose benefit coefficient is optimal in polynomial time. We construct a controlled testbed based on BigCodeBench to isolate the effect of skill selection on execution success. On it, BPS with a learned capability encoder reaches a success rate of 0.65, and the strongest baselines need at least 28% more tokens to reach 0.60.

cs.AI↗

Trace Anomaly and Effective Topological Sources in Neutron Stars

This work investigates whether the trace anomaly can diagnose the stellar response to a topological scalar field in tensor multi-scalar gravity. Eleven cold tabulated equations of state (EoSs) were examined, with six selected after conservative GR stability and speed of sound checks over the tabulated density range. Their fiducial topological configurations were compared with the corresponding GR models under complementary matching prescriptions. After controlling for stellar mass and EoS dependence, the GR trace source strength $S_T$ remained strongly correlated with the topological mass response, with a partial Spearman coefficient $ρ= 0.975$ within the sampled fiducial sequences. Near $1.4\,M_\odot$, the topological configurations were systematically more compact, with reductions in the Jordan frame radius of $6.7$--$8.7\%$ at fixed baryonic mass and $0.87$--$1.29\,\mathrm{km}$ at fixed gravitational mass. These shifts are comparable to current uncertainties from the Neutron Star Interior Composition Explorer (NICER) and their observational impact depends on the source. Despite the global deformation, the interior effective source remained dominated by matter, with a median topological contribution of about $0.8\%$. The GR matter trace therefore emerges as a useful diagnostic of the fiducial topological stellar response.

nucl-th↗

Learning Exact NVIDIA SASS Encoders with $\mathbb{F}_2$ Linear Algebra

NVIDIA provides a SASS disassembler but no public SASS assembler for recent data-center GPUs, limiting controlled machine-code rewriting. We present F2Asm, which learns exact 128-bit SASS encoders from paired disassembly and original CUBIN instruction words. To our knowledge, F2Asm is the first system to learn SASS instruction encoders as vector-valued affine maps over $\mathbb{F}_2$ and the first open-source NVIDIA SASS assembler to support Rubin SM107. F2Asm uses Gaussian elimination over $\mathbb{F}_2$ to incrementally build a compact basis, detect inconsistencies, and reject inputs outside the learned span. F2Asm separates target-specific control bits, relocation rules, and CUBIN metadata from its learning algorithm. We train encoders for Hopper SM90/SM90a, Blackwell SM100, and Rubin SM107 using 3,225 CUBINs from pinned NVIDIA and third-party production libraries, CUDA 13.3 packages, and CUDA 13.4 Developer Preview archives. In round-trip tests, F2Asm reassembles each CUBIN's disassembled SASS, and all compared executable text sections match the originals exactly. Joint training with F2Asm yields one shared encoder for five Blackwell SM targets and another for three Rubin SM targets, providing strong evidence of a common SASS encoding scheme for instructions shared within each family. Continual training extends the Rubin encoder to all 17,159 previously unsupported cuTile and GROMACS queries with 1,504 additional basis rows, matching the derived lower bound.

cs.LG↗

Logarithmic Brunn--Minkowski Inequality under $n-2$ Reflection Symmetries

For origin-symmetric convex bodies in $\R^n$ invariant under $n-2$ orthogonal hyperplanes reflections, we establish the logarithmic Brunn--Minkowski inequality and classify all equality cases without regularity assumptions. We derive the logarithmic Minkowski inequality and rigidity for cone-volume measures, including an explicit classification for bodies of revolution and uniqueness for simplicial polytopes in the prescribed symmetry class.

math.MG↗

Don't Solve, Just Compare: Tiny Advisors for Runtime Intervention in LLM Agents

LLM agents are emerging as an important paradigm for real-world tasks that require reasoning, tool use, and sequential decision-making. As these agents operate over longer horizons, runtime intervention offers a way to improve reliability without retraining the underlying actor. Effective intervention must provide a useful direction for recovery besides a warning. Existing approaches often rely on an expert solver or a critic that generates task-specific corrections, incurring either the cost of another capable solver or the capacity demands of a task-capable critic. We introduce Comparison-Only Tiny Advisor (COTA) for constructive runtime intervention, which reduces the learned intervention role to local action comparison. A lightweight comparator judges the actor's proposal against available alternatives, and preferred alternatives are returned as non-binding advice for replanning. The comparator is trained from same-prefix counterfactual branches. Across WebShop, ALFWorld, and tau^3-Retail with three LLM actors, COTA instantiated with a 0.5B comparator consistently improves the original actor and achieves the strongest overall performance--cost trade-off among the compared methods. These results suggest that effective runtime intervention need not itself be a task-solving problem: the intervention role can be separated from task solving and handled by a lightweight model specialized for local comparison.

cs.AI↗

Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance

The transport of dangerous goods by sea is a high-consequence activity governed by the International Maritime Dangerous Goods (IMDG) Code, a complex regulatory framework where errors in classification, packaging, stowage, or segregation can result in fire, explosion, toxic release, or loss of life or vessel. Correct compliance requires accurately interpreting hundreds of pages of interacting provisions, updated on a two-year amendment cycle. Practitioners increasingly use Large Language Models (LLMs) as decision-support tools, yet no systematic evaluation exists of whether they can reliably interpret IMDG requirements for safety-critical use. This paper introduces DGEval, the first benchmark for evaluating LLM knowledge of IMDG Amendment 42-24. Built from expert-written questions on a commercial e-learning platform and structured lookups from the Dangerous Goods List (DGL), it comprises 1,678 questions across multiple-choice, open-ended, DGL lookup, and regulatory identification tasks. We evaluate 13 models from six providers across multiple thinking configurations, including one maritime domain-specific fine-tuned model, and test the effect of web search. Although the best-performing model exceeds the human practitioner baseline on multiple-choice questions, all models are weakest in the operationally safety-critical areas of stowage, segregation, and regulatory recall. These results indicate that LLMs may support compliance tasks, particularly structured DGL lookups with web search, but unreliability in operational areas and regulatory-text recall means human oversight and authoritative source verification remain necessary before deployment in any safety-critical context. DGEval is designed as a safety assurance instrument to be applied continuously as models evolve, not as a settled characterisation of current capability.

cs.AI↗