Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,189 records · Page 66Linked to original sources

ATTUNER: Recomputation-Free KV Cache Reuse via Query-Side Adaptation

Large language model (LLM) agents repeatedly load reusable content, such as skills, documents, and memory entries, into the current context. Re-encoding this content for every request wastes computation. Position-independent caching (PIC) alleviates this by encoding each artifact independently and reusing its key-value (KV) states at arbitrary positions, but it incurs a quality loss relative to full-context prefill. Existing methods repair this loss by restoring global position IDs or recomputing selected tokens. In this work, we isolate the source of the loss, finding that the positional mismatch has minor effect, and independently cached artifacts retain faithful representations: reading a provided artifact stays largely accurate, and performance degrades only when the model must select among multiple artifacts. Moreover, replacing PIC's attention scores with full-prefill scores recovers performance with the cached KV unchanged, localizing the failure to the attention rather than KV recomputation. Motivated by this, we propose \textsc{Attuner}, a query-side adaptation method that learns to read a frozen artifact cache. \textsc{Attuner} inserts low-rank adapters into the query projections and is trained by distilling full-prefill distribution into the student. It trains fewer than 0.05\% of the model parameters and, at inference, requires neither cache recomputation nor a full-context reference. On Qwen3-4B and Qwen3-8B across seven benchmarks covering skills, documents, memory, and code, \textsc{Attuner} substantially outperforms prior PIC baselines in both in-domain and out-of-domain settings, matches full-context prefill quality while providing up to $3.73\times$ speedup.

cs.CL↗

The Critical Value for the Beurling-Type Theorem in Weighted Bergman Spaces

For every $α>1$ we construct a finite set $A\subset\mathbb{D}$ such that the zero-based invariant subspace $I_A$ fails the wandering subspace property. Consequently, the Beurling-type theorem holds on the weighted Bergman space $A^2_α$ if and only if $-1<α\le1$, confirming a conjecture of Shimorin. As a further application, we extend the prescribed-curvature construction of Hedenmalm and Perdomo to all $α>1$, thereby replacing the constant $α_0\approx1.04$ in their result by $1$.

math.FA↗

Routing in Gradient Space: Balanced Usage Is Not Expert Specialization

Sparse expert models can distribute traffic evenly while still grouping incompatible training signals within the same experts. We study routing as a gradient-partitioning problem and introduce gradient-aligned routing (GAR), whose load-normalized router objective rewards grouping observations with aligned gradients. On five multi-task text-classification mixtures, we compare GAR with task-loss-only routing, gradient-combination and gradient-conflict methods, and load-balancing losses. With a fully trainable RoBERTa backbone and classification-head experts, GAR has the highest aggregate validation accuracy, 1.07 percentage points above task-loss-only routing. With frozen DeBERTa and Qwen3-1.7B backbones and low-rank adapter experts, it again ranks first, 1.10 points above task-loss-only routing, with better-balanced expert load and higher gradient-mass purity, the share of each expert's gradient-norm mass from its dominant task; the load-balancing losses flatten load further but leave this purity near its task-loss-only level. Top-1 routing, trainable full-parameter feed-forward network (FFN) experts, and a larger backbone also show positive aggregate gains. The results distinguish expert-load balance from gradient-based routing organization and indicate the predictive value of gradient-informed routing in multi-task text classification.

cs.LG↗

Arithmetic Nonexistence Results for Tight Spherical $5$-Designs

We prove two arithmetic nonexistence criteria for tight spherical $5$-designs in dimension $(2m+1)^2-2$ with $m$ even. Together, they recover the applicable criteria of Bannai-Munemasa-Venkov and Nebe-Venkov while covering parameters excluded by neither earlier result. Examples include $m=16,100$ under the first criterion and $m=96,168$ under the second. The proofs combine lattice moment identities, finite Gauss sums, and a determinant calculation modulo $3$.

math.CO↗

Can AI Scientists Change Their Minds? Prior-Evidence Conflict in Synthetic Universes

Can a scientific agent distinguish a law it inferred from evidence from one it merely recognizes? We introduce Synthetic Universes, a controlled benchmark that pairs canonical famous worlds with matched twisted twins governed by nearby noncanonical mechanisms. We evaluate each reported law twice: by executing it on held-out continuations and transfer settings, and by independently checking whether it recovers the generating mechanism. In the current checkpoint of a pre-specified 60-cell study, 22 trials were graded and one additional run ended in infrastructure failure. Among 20 twin trials, 8 pass predictive verification while 5 recover the generator. The dissociation is bidirectional: six parsable outputs predict successfully while missing the mechanism, whereas three recover the mechanism but fail predictive rollout. Drag exhibits the first pattern (5/5 predictive pass, 1/5 mechanism recovery); Gravity exhibits the second (1/5 predictive pass, 4/5 mechanism recovery). Because matched famous controls, the corrected identifiability sweep, and the Evidence Ladder remain incomplete, we do not claim a confirmatory causal prior-conflict effect. Instead, the completed runs establish a narrower verification result: predictive adequacy and mechanism recovery are distinct scientific claims and require distinct tests.

cs.AI↗

HARS: LDPC Bit-Flipping Decoding With Initial-Syndrome-Conditioned Parameter Mapping and Local-Reliability Weighting

HARS (Hybrid Adaptive Reliability-Aware and Syndrome-Aware) is a bit-flipping decoder that combines local channel reliability with parameter selection from the initial syndrome. A bounded check reliability modifies the parity-check weight, while the initial syndrome weight selects the base weight, reliability coefficient, threshold decay factor, and perturbation amplitude for each frame. Computing these quantities at initialization supports synchronous multi-bit updates with a simple iterative datapath. For rate-one-half PEGReg and WiMAX codes over the binary-input additive white Gaussian noise channel, HARS reduces both bit-error rate (BER) and mean iteration count relative to a calibrated SM-NGDBF baseline. The PEGReg BER at 3.25 dB is approximately one-sixth of the comparison value; the WiMAX BER at 2.75 dB is reduced by 59%. A seven-fractional-bit FPGA implementation shares 132 variable-node lanes among four frame contexts. Precomputed check weights, encoded threshold states, and parallel perturbation generation support one group per clock during sustained decoding. On an XC7A200T at 60 MHz, for preloaded 128-frame batches with a ready receiver, measured coded throughputs are 49.7 Mb/s at 2.50 dB and 361.9 Mb/s at 4.00 dB, including initialization and output transfers.

cs.IT↗

Stochastic dominance of first return times for nearest-neighbor random walks on $\mathbb{Z}^d$

For a $d$-dimensional probability vector $\mathbf{h}=(h_1,\dots, h_d)$, let $(S^{\mathbf{h}}_n)_{n\geq 0}$ be a nearest-neighbor random walk on $\mathbb{Z}^d$ such that at each step, it moves to one of the two nearest neighbors in the $i$-th dimension with probability $\frac{1}{2} h_i$ ($i=1,\dots, d$). Let $T^{\mathbf{h}}=\inf\{n\geq 1: S^{\mathbf{h}}_n=(0,\dots,0)\}$, the first return time to the origin. For two $d$-dimensional probability vectors $\mathbf{h}'$ and $\mathbf{h}''$ with the former majorizing the latter, we show that $T^{\mathbf{h}'}$ is stochastically smaller than $T^{\mathbf{h}''}$. In particular, the first return time for the $d$-dimensional simple random walk is stochastically larger than $T^{\mathbf{h}}$ for all $d$-dimensional probability vectors $\mathbf{h}$.

math.PR↗

Population search for dark matter spikes in megamaser rotation curves

A black hole growing adiabatically inside a dark matter halo is expected to develop a dense spike, with a power-law slope between 1.5 and 2.5, depending on redshift and baryonic activity. Direct dynamical tests of this prediction are almost non-existent. Here we show that water megamaser disks offer one; their masers orbit within a parsec of the black hole and are mapped individually with very long baseline interferometry, so the enclosed mass at each radius can be read off the rotation curve. We develop a search method and apply it to eleven megamaser disks with public kinematics. Compared with a black hole alone, a spike is preferred in NGC 1194 and NGC 4258 and moderately improves the fit in several others. Allowing the disks to warp removes this preference in all but NGC 4258 and NGC 6264. No combination of disks, warped or not, allows a spike heavier than about ten per cent of the black hole mass, a sensitivity validated with mock injection tests. Mapping the warps with maser accelerations and detecting fainter inner masers, where a spike and a tilt differ most, can break the remaining degeneracy. Megamaser disks open a window onto dark matter around supermassive black holes, and, importantly, one that enables population statistics.

astro-ph.GA↗

Can Agents Design Libraries for Agents?

Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design. We evaluate the library through the correctness and simplicity of programs written by three user agents from different model families. The benchmark spans 242 expert-validated programming problems across 15 library-design tasks in four languages. On eleven of the fifteen tasks, agent designers reproduce the abstractions of the human-written production library. Downstream agents adopt agent- and human-written libraries alike but underuse them, reimplementing capabilities the library already provides. Our failure analysis finds that downstream agents write extra code mainly because agent-written libraries are rigid or hard to use, not because capabilities are missing. We also experiment with giving designers more prescriptive, agent-first guidance and having them test their library with subagents; this improves downstream scores and yields simpler programs. LibraryDesignBench provides both a testbed for evaluating library-design practices for agent users and an initial design baseline that improves downstream reuse.

cs.AI↗

Deep Learning Latency Attacks and Defenses: A Cross-Domain Survey of Availability Threats

Adversarial machine learning has focused mainly on integrity, but availability is an increasingly consequential complement. Latency attacks (also energy-latency attacks) increase inference-time work, energy, or response time, causing deadline misses, throughput collapse, or resource exhaustion in vehicle controllers, interactive services, or battery-powered sensors, sometimes while preserving the nominal prediction. This survey unifies a fragmented literature spanning perception pipelines (including physical attacks on autonomous-driving detection and tracking), input-adaptive neural inference (sponge examples, dynamic networks), and autoregressive and agentic systems (output-length, verbose-image, and reasoning denial-of-service attacks on LLMs, VLMs, mixture-of-experts models, and tool-using agents). We organize attacks by exploited computational bottleneck rather than formulation, separating what makes a computation expensive from how the attacker triggers it; the delivery channel (input, prompt or retrieved content, message, poisoning, or weight tampering) is an orthogonal attribute. Many attacks share one mechanism, intermediate-work amplification, motivating a work-budget defense abstraction; we distinguish caps on the work entering an expensive stage from caps on the results leaving it. We further analyze when a model-level cost increase becomes a system-level availability failure, which depends on critical-path share, slack, existing ceilings, accumulation, resource sharing, and fallback policy, not on the amplification factor alone. We also provide a threat-model taxonomy, consolidated quantitative comparisons, a defense review by control mechanism, and open challenges such as standardized evaluation, physical realizability, and whole-system availability. Companion website: https://github.com/guzonghua/awesome-latency-attacks.

cs.CR↗

Short-range correlated pairs from nucleon density profiles as a new probe of the nuclear equation of state

We investigate to what extent the observed nuclear systematics of Short-Range Correlations (SRCs) can be described by the geometry of the underlying proton and neutron density distributions. Building on previous connections between nuclear SRC contacts and one-body densities, we formulate an explicit density-overlap representation in which proton--proton, proton--neutron, and neutron--neutron SRC source terms are constructed from the full spatial density profiles obtained with Energy Density Functionals (EDFs). The framework contains two global parameters describing the overall SRC pair formation strength and the relative contribution of the spin-singlet channel, while the nucleus-dependent evolution is generated by the density-overlap integrals. The framework simultaneously reproduces several independent experimental SRC observables, including proton--proton to proton--neutron pair ratios, relative SRC pair abundances, and proton and neutron high-momentum double ratios. A comparison between relativistic and Skyrme-type EDF families demonstrates the sensitivity of these observables to the underlying density geometry and allows the effect of the complete density profiles to be distinguished from simple radius-based geometrical estimates. Applied to neutron-rich oxygen isotopes, the framework predicts a pronounced evolution of the neutron--neutron pair abundance. Comparison with an independent calculation based on occupied harmonic-oscillator wave functions and an explicit finite-distance criterion supports the predicted isotopic evolution in neutron-rich oxygen isotopes. This agreement identifies the proton--proton to neutron--neutron SRC pair ratio as a promising experimental probe of neutron-skin thicknesses in neutron-rich nuclei and, through their connection to the symmetry energy, of the nuclear equation of state.

nucl-th↗

Distilling What Matters: Confidence-Aware Selective Distillation for Large Language Models

Knowledge Distillation (KD) trains a smaller-capacity student model to imitate a larger-capacity teacher model by matching output distributions, implicitly assuming the teacher to be a reliable oracle. In large language models (LLMs), this assumption often fails: teacher predictions can exhibit high entropy and hallucinations, causing standard KD to degrade well-calibrated student priors. We propose CaRE-KD, a confidence-gated distillation framework that replaces static objectives with uncertainty-adaptive optimization. CaRE-KD has two components: a token-level loss (CaRE-Divergence) that adaptively switches between Forward and Reverse KL divergence based on teacher--student confidence, and a batch-level epistemic rejection mechanism (Revival) that suppresses updates when the teacher is more uncertain than the student. We provide a gradient-level analysis showing how this dual-granularity design induces a conditional calibration mechanism that prior static divergences cannot reproduce. Empirically, across eight teacher--student pairs and eleven benchmarks spanning instruction following, chat alignment, code generation, and mathematical reasoning, CaRE-KD delivers consistent gains over strong baselines (Skewed-KL, $α$--$β$ divergence). Highlights include up to $+3.2$ average ROUGE-L on instruction-following tasks, $+2.1$ pass@1 on MBPP, $+1.7$ accuracy on GSM8k, and $+1.8$ accuracy on CollegeMath over the strongest baseline, with consistent gains in LLM-as-a-judge factuality (up to $+2.5$ per task over Skewed-RKL). Revival further acts as a principled, loss-agnostic plug-in that systematically strengthens existing distillation objectives by filtering epistemically unreliable teacher supervision.

cs.CL↗

BiFE: Search-Efficient Discovery of CPU-Only Branching Policies via LLM-based Bi-Fidelity Evolution

In branch-and-bound (B&B) for mixed-integer linear programming (MILP), branching variable selection critically impacts efficiency. Existing neural branching policies often require GPU inference, while CPU-efficient symbolic expressions lack the representational capacity for complex logic. Large Language Model (LLM)-generated code provides a flexible search space for designing lightweight branching rules with diverse algorithmic logic. To discover effective rules within LLM-based evolutionary frameworks, a core challenge arises: full B&B evaluation on real instances is prohibitively expensive, whereas offline imitation learning suffers from distribution shift. To address this, we introduce a Bi-Fidelity Evolutionary framework (BiFE). It employs low-fidelity imitation scores as a rapid pre-screener and selectively applies high-fidelity on-instance evaluation only to elite candidates, effectively balancing search efficiency with performance reliability. Experiments validate both the search efficiency of BiFE and the competitiveness of its discovered rules, which outperform the SCIP solver and other baselines on CPUs, and even surpass certain GPU-based neural policies.

cs.AI↗

From Neurons to Conversation: Speech Brain-Computer Interfaces

Speech brain-computer interfaces (BCIs) aim to restore communication by transforming neural activity related to speech, language, or communicative intent into external outputs such as text, synthesized voice, or avatar control. Recent advances in intracortical and electrocorticographic recording, deep sequence models, and language-model-assisted decoding have enabled rapid progress, including high-performance attempted-speech decoding and increasingly naturalistic speech synthesis. Yet these achievements also reveal that speech BCIs are not simply neural-to-text decoders. They are adaptive clinical systems in which neural representations, recording hardware, decoding architectures, language priors, feedback, and user learning interact over time. Here, we synthesize speech BCI research from a system-level perspective. We first examine the neural substrates of speech and language, emphasizing their hierarchical, distributed, temporally structured, and non-stationary organization. We then examine recording and decoding choices, closed-loop adaptation, evaluation, clinical translation, and ethics. Across these domains, we highlight recurring trade-offs between signal resolution and invasiveness, low-level motor and high-level semantic targets, decoder accuracy and user agency, and language-model fluency and faithful neural evidence. We argue the next generation of speech BCIs should be evaluated not only by offline accuracy, but also by robustness across sessions, calibration burden, latency, uncertainty, usability, and safeguards against unintended decoding. By reframing speech BCIs as adaptive, user-centred systems, we outline the interdisciplinary priorities spanning speech neuroscience, neural engineering, machine learning, clinical practice, and neuroethics needed to move from proof-of-concept decoding toward reliable, expressive, and controllable communication neuroprostheses.

cs.HC↗

Reconstructing the Vocal Tract with Differentiable Acoustic Simulation

The vocal tract is the region of the human body responsible for filtering one's voice to create speech. In this paper, we present a differentiable and GPU accelerated acoustic simulator for the vocal tract. The differentiable simulator synthesizes speech by propagating sound along an acoustic tube model of the vocal tract, and via its gradients, can solve the inverse problem: reconstructing the shape of the vocal tract solely from the sound it produces. Although the inverse mapping between geometry and sound is notoriously non-convex, we discover that gradient descent succeeds with three technical contributions: (1) we design a frequency domain formulation of the vocal tract's fluid dynamics that is 70x more GPU parallelizable than finite differences in time, (2) we integrate a differentiable model for turbulence to synthesize consonants, and (3) similar to prior work in implicit neural representations (INRs) and neural fields, we find that parameterizing the geometry with a neural network accelerates convergence and escapes local minima that trap discrete representations. Because the simulator is differentiable, it is readily integrated with other deep learning pipelines to enable novel linguistics and medical imaging applications. (1) We demonstrate self-supervised autoencoding of vocal tract shapes across 11 languages, and (2) we couple our simulator with a generative model of MRI (magnetic resonance imaging) images to reconstruct one's moving vocal tract from only their speech without paired data.

cs.SD↗

Backpropagated Output Momentum: Relocating Optimizer History from Parameters to Task Space

Optimizer momentum is usually stored as a parameter-sized moving average of past gradients, which makes history costly and fixes each past signal in the coordinates in which it was computed. We introduce Backpropagated Output Momentum (BOM), which instead stores a compact moving average of prediction errors at the model output and reprojects that history through the current network at every step. A batch-level analysis characterizes the information retained and omitted by this relocation, while the implementation preserves the current supervised gradient and can replace the first-moment component of several adaptive optimizers. As a plug-in for momentum-based optimizers, including ones that already compress their state, BOM reduces parameter-shaped optimizer state by 49.7-99.8% in three compositions and, averaged over three language backbones, paired step time by 4.0%. It also improves mean validation performance across language and vision fine-tuning, by 1.42 points in the primary five-task comparison. Language and vision pretraining studies, together with matched mechanism controls, further test the construction across output spaces and model scales.

cs.LG↗

Frontier Autolab: Organizational Memory, Adversarial Dissent and Temporal Leakage in Multi-Agent LLM Firms Across Fifty Years of Technological Change

Multi-agent LLM systems are increasingly structured like organizations, with roles, critics and shared memory, yet they are evaluated on tasks that last minutes. We ask how such an organization behaves when the ground it stands on keeps moving. Frontier Autolab is a long-horizon testbed in which one simulated firm, voiced by sixteen role personas and a dedicated Red Team, must re-found itself in nine technology eras from 1990 to 2040. Each era is temporally gated: the firm decides from a dated briefing, a historian-judge then reveals what happened and scores the decision on a five-dimension rubric, and lessons enter a persistent Playbook. Six eras are scored against history, one against the live market and two are open forecasts. Across four trajectories (36 era decisions, 180 subscores) we find a consistent foresight-commitment gap: in all 24 historically scored eras the judge rated the firm's recognition of the coming shift above its choice of where to build (mean gap 1.9 points on a 10-point scale), because boards chose the layer their existing assets could reach. Organizational design shaped long-run character. A Red Team armed with numeric kill gates produced fifty years of gated pilots and no product, and the rubric rated this firm highest; firms whose memory stored market-structure lessons pivoted every era, while a firm whose memory stored only validation procedure kept one method throughout. We also show why such results are hard to trust. Scores rise across eras in every run while the judge's own hindsight subscore falls (within-run r = -0.58), so apparent learning is confounded with recall of history, and we trace further distortions to self-judging, briefing selection and score aggregation. We release all records and an API harness, and specify fictional and post-cutoff eras that would turn the testbed into a benchmark.

cs.CE↗

Efficient Offline Learning of Ranking Policies via Top-$k$ Policy Decomposition

Many recommender systems such as for e-commerce and news platforms aim to provide users with rankings they are likely to interact with. Off-Policy Learning (OPL) of ranking policies enables us to learn new ranking policies using only historical logged data. However, ranking settings make OPL remarkably challenging because their action spaces consist of permutations of unique items, being extremely large. Existing methods primarily use either policy- or regression-based approaches. The policy-based approach, which typically uses importance-weighted policy gradients, can suffer from high variance due to large action spaces. The regression-based approach, on the other hand, estimates the expected reward using conventional machine learning methods, avoiding variance issues but potentially suffering from severe bias. To circumvent these issues of existing methods, we propose a new OPL method for ranking, named Ranking Policy Optimization via Top-$k$ Policy Decomposition (R-POD), which combines the policy- and regression-based approaches in an effective fashion. Specifically, R-POD decomposes a ranking policy into a first-stage policy for selecting top-$k$ actions and a second-stage policy for choosing the bottom actions given the top-$k$ actions. It learns the first-stage policy using a new policy gradient estimator and the second-stage policy via the regression-based approach. This method can substantially reduce variance, since it applies importance weighting only to the top-$k$ actions. We also demonstrate that our policy-gradient estimator for the first-stage policy is unbiased under a conditional pairwise correctness condition, which only requires that the expected reward differences of pairs of rankings sharing the same top-$k$ actions can be estimated correctly.

cs.LG↗