Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,279 records · Page 71Linked to original sources

HXI-DLA2: A Physics-Constrained Deep Learning Algorithm for the ASO-S Hard X-ray Imager

Solar flare hard X-ray imaging is a key diagnostic of flare energy release and electron acceleration. The Hard X-ray Imager (HXI) aboard ASO-S compresses the two-dimensional source distribution into counts measured by 91 sub-collimators, making image reconstruction an inherently underdetermined inverse problem. Conventional algorithms such as CLEAN rely on point-source priors and manual tuning, whereas recent deep-learning methods offer no guarantee that their reconstructions obey the instrument's modulation-sampling forward equation. In this work we show that the counts decompose into two nearly decoupled quantities---the counts average energy, which tracks the total source flux, and the normalized counts distribution, which encodes the source spatial structure---and we exploit this property to construct a physics-constrained network, the Hard X-ray Imager Deep Learning Algorithm 2 (HXI-DLA2). Non-negativity and exact counts-average-energy closure are enforced at the network output, while a distribution-consistency loss aligns the re-projected counts with the measurement, so that the reconstruction satisfies the forward equation by construction. Tests on simulated Gaussian sources, observed soft X-ray morphologies, and a real HXI flare event show two main improvements over existing methods: the limiting resolvable dynamic range of double sources is pushed well beyond that of conventional imaging algorithms and our previous method; and complex morphologies on which prior reconstructions degrade, such as ring-like and diffuse structures, are reliably reconstructed, with the real-event result consistent with contemporaneous SDO/AIA imaging. Embedding the instrumental forward equation as a hard constraint while learning source priors from data offers a general inversion framework for modulation imaging.

astro-ph.SR↗

SR-OPSD: Self-Referenced On-Policy Self-Distillation

On-policy self-distillation (OPSD) converts feedback into dense token-level supervision on student-generated trajectories, complementing reinforcement learning with sparse outcome rewards. Its self-teacher, derived from the student's current or exponentially averaged parameters and conditioned on additional context, evolves alongside the student and its rollout context distribution. The benefit of modifying this moving target depends on how target--student probability mismatches translate into updates. We propose \emph{Self-Referenced On-Policy Self-Distillation (SR-OPSD)}, which constructs a normalized geometric target from the self-teacher and a frozen initial policy, then minimizes the forward Rényi divergence from this target to the student. The interpolation coefficient controls the self-teacher's contribution, while the Rényi order controls the power weighting of target-to-student probability ratios in the gradient. For fixed contexts and target components, we establish a conditional variational characterization and derive the exact token-logit gradient, revealing how anchoring and projection jointly shape the effective update target. Experiments across scientific reasoning, tool use, mathematical reasoning, and code generation demonstrate strong performance across multiple model families and scales. Ablations further show that reference anchoring can improve or degrade performance depending on the projection objective, supporting the joint design of target construction and projection geometry.

cs.LG↗

RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance

General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from large-scale heterogeneous corpora remains underexplored. Existing approaches tie supervision to task-internal anchors such as preferences or normalized progress, none of which transfer cleanly across embodiments and data sources. We introduce RynnValue, an open-source value foundation model for robotic manipulation that replaces these anchors with temporal distance, the directed cost-to-go from an observation to the language-specified goal. Because temporal-distance labels can be derived directly from timestamps, RynnValue scales to over 7,000 hours and roughly 3M instruction-conditioned clips without preference or progress annotations. To make temporal-value learning reliable at scale, we combine random temporal sampling, temporal-order shuffling, and value-isolation attention, suppressing shortcuts that would leave predictions insensitive to failures and regressions. Trained without preference labels, RynnValue attains an average Kendall's $τ_a$ of 0.704 on RBM-EVAL-OOD, surpassing the fully preference-supervised state of the art (0.655) and more than doubling a progress-only counterpart (0.292), while generalizing zero-shot to unseen tasks, embodiments, and viewpoints. As a zero-shot reward model, RynnValue serves a range of downstream applications. Converted into dense rewards via potential-based shaping, it raises real-world policy success from 52.5% to 72.5% online and from 63.8% to 82.5% offline; used for data filtering, it improves multi-task behavior cloning success from 35.0% to 42.5%; and applied as inference-time value guidance, it lifts a frozen policy's success from 67.5% to 80.0%. These results establish temporal distance as a scalable supervision target and practical reward interface for generalist robot policies.

cs.RO↗

RAISE: Diagnosing Acquisition Collapse in Costly LLM Signals

Large language models (LLMs) are increasingly used as costly, on-demand components in real systems, but calling them indiscriminately can waste substantial compute, latency, and serving budget. The key deployment question is therefore not only whether an LLM helps on average, but when it is worth calling. We identify a common failure mode, which we call acquisition collapse: an LLM signal can appear useful in aggregate or post hoc, yet still provide too little before-call information to support reliable selective use. We introduce RAISE (Reward-SNR Actionability in Signal Evaluation), a pre-routing diagnostic framework for testing whether available evidence supports selective use before committing to a routing strategy. We instantiate RAISE with Structured Hypothesis Embeddings (SHE), a frozen-LLM intent signal for recommendation using one LLM call per user, and evaluate it through controlled, retrospective, and fresh-cohort studies and a prospective offline pilot whose audit decisions are frozen before independent outcomes are revealed. Across these settings, predictable incremental benefit, not average lift alone, distinguishes settings with recoverable selective value; deployment additionally depends on cost and operational constraints. Seemingly strong oracle or subgroup gains can disappear under independent evaluation. More broadly, RAISE reframes costly inference as an information-acquisition problem: before paying for an expensive model, tool, sensor, or measurement, first test whether its value is predictable at decision time. This principle motivates cost-aware acquisition in settings ranging from agent tool use and stronger-model consultation to robotic sensing and clinical decision pipelines.

cs.LG↗

Three Tokens Force Exponential Feature Rank in Nonnegative Kernel Attention

How much feature rank does comparison require in kernel attention? On Min-IP over $m$-bit tokens, rank one solves every sequence of length at most two exactly. At length three, the minimum feature rank of one normalized nonnegative kernel-attention head is $2^{Θ(m)}$ for error strictly below $1/2$ on every input, even with arbitrary finite-dimensional tokenwise values and query-dependent affine readouts. Dense softmax solves this three-token task with $m$-dimensional scores and temperature constant in $m$. For every fixed number of heads $H$, the minimum total feature rank is $2^{Θ_H(m)}$ for the same error guarantee at exact length $H+2$ in one attention layer with affine mixing. These bounds also hold with position-dependent maps and a final causal query. For one head with polynomial readout of fixed degree at most $D$, rank one suffices at exact length $D+1$, while exact length $D+2$ requires exponential feature rank. With unrestricted exact-real decoding, a scalar rank-one construction solves the task at every finite length. This motivates a separate bound on total communication for deterministic models with finite-alphabet cross-token channels and any number of heads and layers. In this setting, correctness up to length $n$ requires $Ω(n\log m)$ bits over a range of lengths that grows exponentially with $m$.

cs.LG↗

Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

We present the first systematic study of massive activations (MAs) in layer-interleaved Hybrid linear attention large language models (HLA LLMs), examining their architectural organization, training-time emergence, underlying mechanisms, and functional significance. Across five linear attention architectures, six hybridization configurations, and five input domains, we identify two architecture-aligned morphologies: pre-attention spikes (PAS) immediately before full attention and inter-spike plateaus (ISP) persisting through intervening linear attention layers. Denser full attention increasingly connects PAS through ISP, approaching the persistent MAs of conventional Transformers. This organization also recurs across 12 public checkpoints spanning 1.2B-397B parameters, covering linear attention and state-space hybrids. Controlled pretraining of Gated DeltaNet (GDN) hybrids up to 1.3B reveals early emergence and consolidation of both morphologies, alongside asymmetric gating effects. Specifically, full attention output gates strongly attenuate MA magnitudes without eliminating their organization, whereas removing GDN output gates yields modest amplification. Mechanistically, we develop a shared systematic-outlier account: PAS follows a localized write-sink-cancel process, while ISP is consistent with delayed cancellation. Functionally, our interventions show that deleting only the four largest-magnitude PAS coordinates at each full attention input reduces mean downstream accuracy by 21.9%-63.6% relative to normal inference. Moreover, reference-conditioned spike-to-plateau connection consistently improves mean real-world retrieval accuracy, yielding relative gains of 1.1%-12.6% without retraining. Our code is available at https://github.com/StartLuxLabs/Massive-Activations-HLA.

cs.CL↗

Single-axis high-energy X-ray diffraction tomography for elastic residual strain: uniqueness and stability of solutions in the presence of equilibrium constraints

It is well-established that single-axis Transverse Ray Transform tomography data from high-energy X-ray diffraction is insufficient for general reconstruction of three-dimensional elastic strain. In this paper we show that, when combined with the constraint of mechanical equilibrium, this problem becomes uniquely solvable for isotropic elastic samples with known elastic constants and non-zero Poisson's ratio. Building on the reconstruction framework of Desai and Lionheart [N.M. Desai, W.R.B. Lionheart, An explicit reconstruction algorithm for the transverse ray transform of a second rank tensor field from three axis data, Inverse Problems 32 (11) (2016) 115009], we formulate this constrained single-axis tomography problem in the Fourier domain and show that the resulting system is invertible everywhere outside a finite collection of measure zero characteristic planes. In the case of bounded samples, this is sufficient to establish uniqueness for reconstructions. We further derive conditional stability estimates that quantify the regularity requirements associated with reconstruction of the various strain components.

math.AP↗

Decode-Branch Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation

As large language models serve ever more requests, cumulative inference cost is growing relative to the one-time cost of training. In typical serving, prompt prefill runs in parallel and is compute-bound, whereas autoregressive decode is sequential and memory-traffic-bound. Conventional width or depth scaling raises both costs together, since every added layer is evaluated in both phases and enlarges the weights read at each decode step. We instead ask whether additional learned computation can be allocated to continuation prediction while preserving prompt-wide primary computation and a single KV cache. We realize this with the Decode-Branch Transformer. Its primary path alone processes the prompt and writes the KV cache; the decode branch is omitted during prefill and activated only from the final prompt position onward, adding continuation computation without writing state or affecting the primary path. The paths share attention, MLP, and output matrices, using separate token embeddings with lightweight coupling. Grouped decode reuses loaded weight tiles and the primary KV cache across both paths, so the added arithmetic does not proportionally increase dominant memory traffic or decode latency. Across matched-token comparisons, Decode-Branch achieves lower validation loss across architectures and data settings. In MoE models, the primary and branch expert fan-outs become independent knobs for trading prompt cost, decode cost, and predictive quality. We study two expert-allocation regimes, holding prefill or decode computation fixed, and expose a prefill-decode-quality trade-off enabled by phase-specific expert allocation.

cs.AI↗

Spatial Memory Agent: Experience-Grounded Procedural Memory for Spatial Intelligence

Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLMs, existing work has mainly followed two lines. One line uses post-training methods, such as supervised fine-tuning and reinforcement learning. Another line adopts an agentic paradigm in which the model calls external spatial tools to gather intermediate spatial evidence. We study a complementary and underexplored route: Can a frozen VLM agent improve its spatial reasoning through \textbf{parameter-update-free self-evolution}, without depending on external expert spatial tools at inference time? We present \textbf{Spatial Memory Agent (SMA)}, an experience-grounded runtime memory framework that converts verified spatial experience into reusable transferable lessons. Specifically, SMA first queries the frozen VLM in a verifiable spatial environment, obtains a predicted answer and reward, and uses verifier-guided reflection to distill compact transferable lessons stored in memory cards. SMA further assigns each memory card a \textbf{Transfer Reliability Score (TRS)}, which is initialized uniformly and calibrated from later retrieval outcomes as visit evidence of future transfer reliability. During read-only deployment, SMA retrieves memory cards through semantic filtering and combined similarity--TRS ranking, allowing the retrieved memory to guide frozen model inference. Experiments across five representative spatial benchmarks show that SMA achieves the best macro-average accuracy for all four base VLMs and the best accuracy in most individual evaluations, establishing a practical parameter-update-free path for spatial self-evolution through reusable experience.

cs.AI↗

Chebyshev polynomials on a Jordan arc

We describe the asymptotics of Chebyshev polynomials on an analytic Jordan arc in the plane. This gives an affirmative answer to a conjecture of Christiansen-Simon-Zinchenko, based on predictions of Widom from 1969. The proof combines weighted Faber polynomials with extremal signatures, discrete orthogonal polynomials and a Marcinkiewicz-Zygmund sampling inequality, and yields Szegő-Widom asymptotics for the Chebyshev polynomials themselves.

math.CA↗

Spectral Localization in Cavity-Mediated Entanglement Harvesting

We investigate the relation between spectral localization and entanglement harvesting in an analytically solvable model of two qubits coupled to a leaky single-mode cavity, which in turn couples to a continuous electromagnetic bath. We derive the no-jump concurrence at a prescribed reference gate time in closed form, $\mathcal C_g(Q)=2e^{-π/(2Q)}(1+e^{-π/(2Q)})/(1+3e^{-π/Q})$, where $Q\equiv|Δ|/κ$ is the ratio of the qubit-cavity detuning $Δ$ to the cavity linewidth $κ$. This reference time is inherited from the first maximally entangling unitary gate. At finite loss, $\mathcal C_g$ is the conditional concurrence at this prescribed readout time. In the high-$Q$ limit, $\mathcal C_g\simeq1-π^2/(16Q^2)$, while in the low-$Q$ limit it decays exponentially. For the Lorentzian spectral family, $Q$ is proportional to a detuning-normalized inverse participation ratio and simultaneously measures the balance between coherent exchange and collective decay. This yields a compact, model-specific description of the crossover from nearly unit conditional entanglement to overdamped suppression. The predicted $\mathcal C_g(Q)$ curve can, in principle, be examined in circuit QED with ideal monitoring and no-jump post-selection.

quant-ph↗

Enhanced Third-Harmonic Generation in a Bound State in the Continuum Assisted Multiband All-Dielectric Metasurface

Multiband Fano resonances are demonstrated in the near-infrared (near-IR) using an all-dielectric metasurface whose unit cell consists of four silicon nanoblocks on a glass substrate. An in-plane asymmetry triggers symmetry-protected quasi-bound states in the continuum (QBICs), producing multiple high-Q resonances. Their origin is identified through multipolar decomposition of the scattering cross section and field distributions at the resonances. The strong field localization at these resonances enables efficient multiband third-harmonic (TH) generation in the ultraviolet (UV), with a maximum simulated conversion efficiency of $8.5 \times 10^{-3}$ at a peak pump intensity of $1.6\,\mathrm{GW/cm^{2}}$. The metasurface is fabricated in symmetric and asymmetric configurations, and its linear and nonlinear responses are measured under normal incidence. A TH conversion efficiency of $1.2 \times 10^{-6}$ is obtained at a peak pump intensity of $3.25\,\mathrm{GW/cm^{2}}$. These results establish a route to multiband photonic devices, including multiwavelength lasers, multiband harmonic generation, and single-photon sources for quantum photonics.

physics.optics↗

Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions

Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant studies, yet the quality of retrieved evidence and factors influencing study selection remain unclear. We evaluated three general-purpose LLM chatbots (Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5) using 20 clinical questions adapted from 2026 Cochrane reviews. We simulated patient, clinician, and evidence-synthesis researcher roles and obtained four independent responses for each chatbot-role-question combination, yielding 720 responses (3 chatbots $\times$ 3 user roles $\times$ 4 repetitions $\times$ 20 review questions). Chatbots were asked to support their answers with primary clinical citations, which were benchmarked against the included and excluded study sets of the corresponding Cochrane reviews. On average, a single response retrieved 39.2% $\pm$ 29.8% of the corresponding Cochrane included-study set and 5.0% $\pm$ 9.4% of the excluded-study set. Recall of included studies varied significantly by model and user role. ChatGPT achieved higher recall than Claude or Gemini (63.1% $\pm$ 29.5% vs. 37.0% $\pm$ 23.8% vs. 17.3% $\pm$ 13.1%; blocked permutation test, $p=2.0\times10^{-5}$), and the researcher role yielded higher recall than the clinician or patient roles (42.8% $\pm$ 30.8% vs. 38.6% $\pm$ 28.9% vs. 36.1% $\pm$ 29.3%; $p=2.0\times10^{-5}$). Controlling for publication year, citations per year, and open-access status, sample size was the only significant predictor of retrieval: each doubling of sample size was associated with 50% higher odds of retrieval (odds ratio 1.50, 95% CI 1.24-1.81). These findings show that LLM chatbots can retrieve studies identified by expert reviewers, but retrieval varies substantially across models and user roles and favors larger clinical trials.

cs.IR↗

Autonomous Fashion Outfit Composition via Unified Aesthetic Foresight Model

Fashion Outfit Composition (FOC) requires sequentially assembling fashion items into a stylistically cohesive ensemble. Existing works struggle to model this step-by-step process effectively, primarily because they fail to jointly optimize the two critical capabilities required for FOC: intermediate outfit value evaluation and complementary item prediction. This structural disconnect, compounded by the severe sparsity of step-wise aesthetic signals, leaves them without a mechanism to autonomously determine when to stop the composition process. To address these challenges, we introduce the Unified Aesthetic Foresight Model (UAFM), which seamlessly unifies both capabilities into a dual-head architecture over a shared backbone. Crucially, we formulate FOC as a deterministic Markov Decision Process and optimize a value head via a post-decision state temporal difference (TD) objective to recursively backpropagate sparse terminal aesthetic rewards to intermediate states. This enables the model to accurately estimate the potential of a partial outfit evolving into a compatible ensemble. Consequently, UAFM can compute step-wise marginal aesthetic gains for dynamic termination. Furthermore, this unified design strictly aligns aesthetic evaluation and complementary item prediction within a shared representation manifold, intrinsically regularizing the combinatorial search space. Extensive experiments on the Polyvore-Outfits dataset demonstrate that UAFM establishes a new state-of-the-art across both FOC and conventional fashion tasks. Extensive ablation studies further confirm that our post-decision state TD formulation provides the aesthetic value prediction necessary for autonomous dynamic termination, while validating the synergistic benefits of our unified dual-head architecture.

cs.LG↗

Spectral nonassociative $\mathrm{L}^p$-spaces for $\mathrm{JBW}^*$-algebras

We complete the construction of tracial spectral nonassociative $\mathrm{L}^p$-spaces for general $\mathrm{JBW}^*$-algebras. More precisely, if $\mathcal{M}$ is a $\mathrm{JBW}^*$-algebra equipped with a normal finite faithful trace $τ$ and $1 \leq p < \infty$, we prove that $\|x\|_{\mathrm{L}^p(\mathcal{M})} \overset{\mathrm{def}}{=} (τ[(x^* \circ x)^{\frac p2}])^{\frac1p}$, where $x \in \mathcal{M}$, defines a norm on $\mathcal{M}$. This resolves the remaining exceptional case left open by the corresponding result for $\mathrm{JW}^*$-algebras. The main difficulty is the complexified Albert algebra $\mathrm{H}_3(\mathbb{O}_{\mathbb{C}})$, which admits no embedding into an associative operator algebra. To treat this case, we establish a Jordan analogue of the joint convexity of the Kiefer map $\mathrm{M}_n \times \mathrm{H}_n^{++} \to \mathrm{H}_n^{+}$, $(a,h) \mapsto a^*h^{-1}a$, where $\mathrm{H}_n$ is the space of Hermitian matrices, and a Jordan version of a variational formula of Carlen and Lieb. This provides a complex Banach space framework for Jordan-algebraic models arising in some generalized probabilistic theories.

math.FA↗

Statistical validation of calorimeter inpainting with generative diffusion priors

Localized detector inefficiencies produce incomplete calorimeter data that limit the ability to perform precision measurements. We address this problem in relativistic heavy-ion collisions from a Bayesian perspective using pretrained calorimeter diffusion models as priors to reconstruct the missing signal conditioned on surrounding measurements. In this work, we conduct a systematic comparison of several diffusion-based inpainting algorithms, whose performance is evaluated using Bayesian posterior diagnostics of energy response, spatial bias, and uncertainty calibration. The reconstruction fidelity is also analyzed across collision centralities and masked region sizes. This study establishes a general validation strategy for probabilistic reconstruction of missing detector information.

physics.data-an↗

Critical molecular theory and the maximal admissible class for Goldberg-type splittings of $h^p(\mathbb{R}^n)$, $0<p\le 1$

We extend the classical Taibleson-Weiss molecular theory to arbitrary radii, reaching the critical decay |x|^{-n/p}, with the resulting embedding shown sharp. This molecular theory governs the largest class of functions certifying membership in h^p(R^n), 0 < p <= 1, through Goldberg's convolution splitting: a critical Hölder modulus and an \ell^p-Dini decay majorant, sharp on the regularity, decay, and cancellation axes.

math.CA↗

First measurement of the one-point charge correlator in $e^+e^-$ collisions at $\sqrt{s} = 91.2$ GeV with DELPHI Open Data

The chiral structure of the $Z$ couplings imprints a parity-odd flow of electric charge on hadronic $Z$ decays. The related forward-backward asymmetries, a key set of observables in the electroweak precision program, were measured at LEP and SLC using the jet charge. The one-point charge correlator offers a complementary route, measuring the hadronic charge flow directly as a function of polar angle relative to the incoming electron-beam axis, without reference to jets or a reconstructed quark direction, following the formalism developed in a companion paper. We report its first measurement, using $61~\mathrm{pb}^{-1}$ of archival DELPHI Open Data recorded at $\sqrt{s} = 91.2$~GeV in 1994 and 1995. Detector effects are corrected in two stages. The first is derived from fully simulated samples, and the second bounds the residual charge-misreconstruction difference between data and simulation using a measurement in $e^+e^-\toτ^+τ^-$ events. The measured charge correlator exhibits the characteristic parity-odd $\sin(2θ)$ modulation and agrees with the \textsc{PYTHIA}~8.3 prediction. The measurement demonstrates that the parity-odd charge flow of hadronic $Z$ decays is directly accessible in data, differentially in polar angle, and the paper discusses strategies for controlling the associated detector effects, paving the way for a new program of charge-flux measurements, both in archival $e^+e^-$ data and at future colliders.

hep-ex↗