Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 991 records · Page 55Linked to original sources

Willmore Energy Estimates for Klein Bottles

In this note, we provide an estimate for the Willmore energy of Klein bottles in $\mathbb R^N$, building on the ideas of Li-Yau and Montiel-Ros for $2$-tori. In particular, we prove that the Willmore energy $\mathcal{W}(ϕ)> 6π$ for any conformal branched immersion $ϕ:\mathbb{K}^2_b\rightarrow \mathbb R^N$, where $\mathbb{K}^2_b$ is a flat Klein bottle with $0.350\lesssim b \lesssim 0.755$. This confirms partially a conjecture of Kusner. Moreover, for a flat immersion $ψ:\mathbb{K}^2_b\rightarrow \mathbb R^N$, we derive a lower bound of $\mathcal{W}(ψ)$, which confirms a conjecture of Hirsch-Mäder-Baumdicker for flat Klein bottles in $\mathbb R^N$. We also construct a smooth family of flat Klein bottles in $\mathbb R^5$, with Willmore energy between $6.912π$ and $8π$.

math.DG↗

Distilling Privileged Control Barrier Functions into RGB-Only Safety Filters for Dynamic Visual Navigation

RGB-only end-to-end visual navigation policies remain vulnerable to collisions in real-world dynamic environments, motivating a dedicated safety layer. Existing visual Control Barrier Function (CBF) approaches seek to provide safety from RGB observations, but often rely on real-time rendering or explicit scene reconstruction and are primarily designed for static scenes, limiting their practicality for onboard deployment. We propose a teacher-student visual distillation framework that transfers the safety behavior of a privileged CBF teacher to an RGB-only student filter for dynamic environments. The student maps a short RGB history, robot velocity, and a nominal control action directly to a safe action, while the teacher uses ground-truth robot and obstacle states in a real-to-sim dynamic Gaussian Splatting environment. To reduce the teacher-student information gap, the teacher constructs safety constraints only from obstacles observable within the student's RGB history. It also accounts for obstacle-velocity uncertainty to improve robustness to motion variations, while action augmentation exposes the student to diverse safe and unsafe nominal actions to better capture the safety boundary. At deployment, the student requires only RGB observations and robot velocity, without explicit 3D reconstruction or online rendering. Experiments show that the proposed method outperforms visual CBF baselines and improves the safety of RGB-based navigation policies under dynamic obstacle motion. Project page: https://syeon-yoo.github.io/distill-cbf-site/.

cs.RO↗

PDE-OBS: Controlled Evaluation Across Observation Patterns

Physical-field reconstruction and forecasting depend on both measurement density and spatial layout, yet evaluation under a single observation pattern does not characterize performance when that pattern changes. We introduce PDE-OBS, an integrated benchmarking platform spanning numerical data generation, model training, and inference and evaluation under varying observation conditions. It combines 560,000 fields and trajectories from seven partial differential equation families with configurable observation operators and seven adapted baseline methods for stationary reconstruction and short-horizon forecasting. Separating observation construction from physical records allows users to specify parameterized patterns and deterministic mixtures for training and testing while preserving prediction targets and data splits. The evaluation protocol uses references trained for each test pattern to compare models on identical test observations and targets, alongside equal-count groups for spatial-layout comparisons. On a 14,000-record subset, we evaluate 441 trained models under nine test patterns, yielding 3,969 evaluations. Mean cross-pattern error exceeds mean matched-pattern error in all 49 PDE-method pairs, and this finding persists in a configuration-matched subset of 117 models. Denser test observations do not consistently reduce error for a fixed model. Mixed-pattern training on five completed pairs reduces large single-pattern transfer errors, although destination-trained references usually remain more accurate. Together, the benchmark and findings support systematic evaluation of observation-pattern sensitivity and provide a reusable workflow for developing methods under changing measurement conditions. Code: https://github.com/ru1ch3n/PDE-OBS.

cs.LG↗

ParaAnya: Accelerating Parallel Diffusion Sampling with Plug-and-Play Output Caching

Diffusion models have achieved remarkable success in generative tasks, but their inherently sequential sampling process introduces a severe computational bottleneck. Recent Parallel-in-Time (PinT) solvers attempt to mitigate this by parallelizing generation across a sliding window of timesteps, advancing the window only when step-wise changes stabilize. However, this overlapping window mechanism forces the network to repeatedly evaluate the same timesteps. When the input variations between iterations are minimal, these redundant evaluations lead to significant computational waste. To address this inefficiency, we propose ParaAnya, an output cache mechanism agnostic to the parallel sampling algorithm that can reduce the number of function evaluations (NFE). ParaAnya caches input-output pairs of diffusion models and reuses the cached output at overlapping timesteps. By dispatching only cache-miss timesteps to GPU workers, our approach eliminates redundant computation while preserving the structure of the underlying algorithms' update rules. We integrate ParaAnya into four representative parallel sampling algorithms and evaluate its performance on Stable Diffusion v1.5. Across four parallel samplers evaluated with DDIM on eight GPUs, ParaAnya provides $1.30$--$2.43\times$ speedups over their uncached counterparts and reduces NFE by up to 70.1\%, reaching up to a $5.62\times$ speedup over single-GPU serial sampling while maintaining comparable CLIP scores.

cs.DC↗

Non-global logarithms and fiducial transverse-momentum-dependent observables in deep-inelastic scattering

Extracting the intrinsic transverse-momentum structure of quarks from deep-inelastic scattering data requires separating nonperturbative effects from perturbative radiation, which broadens the measured transverse-momentum distributions. We address this problem in electron--proton scattering, $ep\to eX$, by introducing the fiducial imbalance $\boldsymbol{q}_T$, defined as the vector sum of the scattered electron's transverse momentum and all hadronic transverse momenta within a specified rapidity window. This construction reduces radiative recoil without requiring jet reconstruction or an explicit jet veto. We account for non-global logarithms (NGLs) arising from correlated soft emissions across the acceptance boundary. The azimuthally averaged NGL contribution is resummed to all orders at leading-logarithmic accuracy in the large-$N_c$ limit using Banfi--Marchesini--Smye evolution, while the leading NGL correction to the first azimuthal harmonic is included at $\mathcal{O}(α_s^2)$. For the EIC and EicC kinematics considered, widening the rapidity window increases the cross section at low $q_T$, a region particularly sensitive to nonperturbative transverse dynamics, and reduces the dilution of spin asymmetries by azimuthally symmetric soft recoil. Both the Sivers single-spin asymmetry $A_{UT}$ and the worm-gear double-spin asymmetry $A_{LT}$ increase in magnitude, while the $\boldsymbol{q}_T$ integrated cross section remains unchanged. The unpolarized $\cosϕ$ moment generated by soft gluon radiation is sensitive to NGL effects, with its ratio definition reducing common normalization uncertainties. The absence of jet reconstruction makes this observable particularly relevant at the lower collision energies of EicC, providing a way to study nonperturbative spin--momentum correlations by varying the fiducial acceptance.

hep-ph↗

Square matrices with integer eigenvalues under entry permutations

We investigate the problem of which multisets of integer entries yield integer eigenvalues in every square matrix arrangement? Combining module reduction with a dyadic criterion for complete splitting of cubic polynomials, we first show that if the multiset has at least one zero entry, the only possibilities have either at most one nonzero entry or exactly two whose product is a perfect square. We then apply block embedding to extend it to higher dimensions, by obtaining a linear threshold on the number of zeros. For any multiset close to a nonzero constant, we use rank reduction to derive an exact criterion for one exceptional entry, thus yielding nonconstant examples in infinitely many dimensions, and to also derive an exact reduction for two distinct exceptional entries in dimension three. The general fully nonzero three-dimensional case remains open.

math.RA↗

Reliability Testing of Medical Model Performance under Distributed Deployment

Distributed inference has become an indispensable part of deploying medical models under practical latency, memory, and throughput constraints. Although modern frameworks improve serving efficiency through tensor parallelism, mixed precision, kernel fusion, and multi-device communication, they are generally assumed to preserve the behavior observed during centralized HuggingFace evaluation. This assumption creates an evaluation-deployment mismatch: a model may pass offline evaluation but produce a different output after the execution stack changes. To address this mismatch, we propose a testing framework and an improved, distributed-execution-sensitive medical-model benchmark that evaluates the same checkpoint and input under a centralized HuggingFace reference and matched distributed deployments. Extensive experiments across language, vision, and multimodal medical models show that execution changes can produce measurable output disagreements. Across supported visual settings, the test success rate ranges from 0.21 to 0.43 for single-modality models and from 0.32 to 0.98 for multimodal models. The benchmark is aimed at extending medical-model evaluation from capability and security to evaluation-deployment consistency.

eess.IV↗

Adapting Context Compression for Long-Horizon Agents with Counterfactual Continuations

Long-horizon agents require context compression to manage growing interaction histories. Compression quality, however, is ultimately determined by downstream execution. Existing prompt-adaptation methods infer compression errors by comparing full-context and compressed trajectories. Such comparisons cannot isolate individual compressions and are confounded by agent stochasticity. We first find that compression degrades reliability before solvability. Using matched counterfactual continuations that compare execution from the same agent state with versus without compression, we further show that severe degradation concentrates at isolated compression events. Motivated by this finding, we propose PAIR (Prompt Adaptation using Interventional Rollouts) for adapting structured compression prompts. PAIR identifies individual compressions that degrade subsequent execution, diagnoses their effects, and revises the relevant sections of a fixed compression template. PAIR achieves the strongest cross-run reliability among compressed methods in every main benchmark-scope combination, consistently exceeding the competing prompt-adaptation baseline. Without modifying the downstream agent, PAIR brings compressed execution close to the no-compression baseline and sometimes numerically exceeds it.

cs.LG↗

SCOPE: Observation-Conditioned Full-Target Prediction for Sparse PDE Inference

Recovering complete physical fields from sparse observations is challenging because the measurements may not uniquely determine the underlying state. Diffusion-based PDE solvers address this problem through iterative sampling whereas neural operators provide deterministic one-pass predictions. We propose SCOPE (Sparse-Context Observability-aware Predictive Embeddings) to recover complete PDE fields from sparse observations by coupling full-field latent prediction with physical reconstruction. A shared decoder reconstructs fields from both predicted and complete-view representations so that representation learning is guided by both physical recovery and latent matching. We derive a quadratic risk decomposition at fixed teacher-decoder pairs showing why optimal latent prediction need not yield optimal field reconstruction. We also establish sufficient conditions for decoder improvements on complete inputs to transfer to recovery from partial observations. Experiments across five PDE settings show that SCOPE outperforms mask-aware neural operators on all ten forward and inverse tasks and achieves lower errors than those reported for diffusion-based solvers including DiffusionPDE and FunDPS. Decoder-only adaptation further improves recovery without retraining the backbone while retaining deterministic single-pass inference.

cs.LG↗

Experimental demonstration of broadband-laser suppression of cross-beam energy transfer

In direct-drive inertial confinement fusion (ICF), cross-beam energy transfer (CBET) redirects a significant fraction of the incident laser energy out of the plasma. We report the experimental demonstration that broadband lasers reduce CBET-enhanced reflected-light return, performed at the low-coherence Kunwu laser facility with two crossed beams of 0.6% bandwidth and up to 550J at $\sim$2.6$\times$10$^{14}$ W/cm$^{2}$. A coupled ray-tracing model shows that the strong CBET amplification of narrowband reflected light is much weaker under broadband illumination. In symmetric incidence condition of two orthogonal beams, the total fractional scattered energy decreases from 7.57% to 4.64%; in asymmetric incidence condition, which separates stimulated Brillouin scattering (SBS) from specular reflection, the SBS fraction decreases from 3.75% to 0.46% and the specular-reflection fraction decreases from approximately 2.2% to 1.3%. These results establish broadband lasers as an effective, experimentally validated approach to reducing energy escape through CBET-enhanced reflection in direct-drive ICF.

physics.plasm-ph↗

Triadic Linear Attention: Three-Dimensional Recurrent States for Long-Context Sequence Modeling

Recurrent neural networks (RNNs) compress the historical context into a memory state of fixed size, thus allowing for constant-time inference. The memory state size is a crucial factor in their performance, as exemplified by the strong performance and resurgence of linear attention, which extends the vector-valued hidden states of ordinary RNNs to matrix-valued hidden states. Crucially, linear attention does so in a parameter-efficient way, in particular by using an outer product of the key and value vectors to write to the matrix-valued hidden state. We generalize this construction and propose triadic linear attention, which writes the triadic outer product of a key, a second key, and a value, into a third-order (i.e., 3D) tensor state, and reads from it by contracting both key axes with two queries. An $E$-dimensional second key thus yields an $E$-fold increase in state size while adding only two projections. Triadic linear attention is compatible with data-dependent forgetting, the delta rule, and chunkwise-parallel training. Applied to Gated DeltaNet and scalar-gated linear attention, triadic linear attention substantially improves long-context language modeling and recall, outperforming alternatives that enlarge the state.

cs.LG↗

Trajectory-Level Mode Guidance for Controllable Diffusion-Based Multi-Robot Motion Planning

Motion planning often admits multiple feasible solutions, making multimodal generation valuable, particularly for flexible multi-robot coordination. Diffusion models naturally learn such trajectory distributions, yet incorporating coarse and partial trajectory priors without restricting generation remains challenging. Such priors indicate a desirable region of the solution space rather than a single solution, motivating conditioned generation that preserves multimodality. In this paper, we guide trajectory generation in the clean trajectory space and progressively incorporate trajectory priors with a timestep-dependent guidance strength. At each reverse diffusion step, the reconstructed clean trajectory provides a unified space for integrating planning costs and partial trajectory priors. Planning costs are incorporated through gradient-based refinement, while the partial prior is progressively injected at the corresponding noise levels with decreasing guidance strength. This guides generation toward the prior in early stages while gradually releasing the constraint to preserve the inherent multimodality of the diffusion model. The framework naturally extends to multi-robot planning by incorporating inter-robot collision costs. Experiments on single- and multi-robot planning tasks demonstrate controllable trajectory synthesis, diverse feasible solutions, and safe multi-agent coordination.

cs.RO↗

Foresight at the Event Boundary: Evaluating Physical Prediction in Video World Models

Video world models are largely regarded as predictive models of the physical world and are therefore expected to anticipate the consequences of observed events. However, evaluation has mainly focused on reference similarity, physical-law consistency, or judgment plausibility, estimating anticipation only indirectly. We address this directly: when a release or impact has just occurred but its consequence is withheld, can a world model anticipate what should happen next? We introduce an event-anchored evaluation based on 62 controlled real-world free-fall recordings and 124 clips spanning three object types, with fine-grained release and impact annotations and ground-truth trajectories. The protocol separates consequence production, temporal placement, and physical realization. Across six contemporary video generation and world models, Runway and Veo produce release and subsequent impact events at rates above 93% but often initiate them substantially late, whereas Cosmos-Predict-2.5 and MAGI-1 frequently preserve the pre-event state and produce little or no measurable consequence. Among measurable falls, plausible timing does not necessarily imply physically consistent motion. We further conduct a 15-participant, 20-condition human study in which participants describe the expected consequence from a single event-anchored frame and draw its trajectory. Human predictions favor the recorded future in aggregate while revealing genuine ambiguity among plausible continuations. Overall, physical foresight emerges as a sequence of distinct challenges: initiating a consequence, anchoring it in time, and realizing its motion.

cs.CV↗

Hierarchical Utility Calibration for Structured Multiclass Decisions

In multiclass probabilistic prediction, Utility Calibration (UC), which focuses auditing on specified utilities, has recently received attention as a way to guarantee downstream decisions while controlling computational and sample requirements. At the same time, some multiclass problems have meaningful label hierarchies that play important roles in medicine and image classification, yet how UC evaluates utility within a hierarchy remains insufficiently understood. We show that the difference between realized utility and predicted mean utility admits an exact decomposition into a sum of contributions from the internal nodes of the label tree. This decomposition shows that positive and negative contributions from different nodes can cancel, and that even when UC is small, the utility errors remaining in parts of the hierarchy need not be small. To address this problem, we propose Hierarchical Utility Calibration (HUC), which evaluates each node contribution before summation while retaining the same target utility, subgroup, and predicted-utility interval. We further provide finite-sample evaluation over all predicted-utility intervals and propose HUC-Boost, which updates only violated internal nodes, with theoretical guarantees for both.

stat.ML↗

Calculation of the temporal structure for a pulsed slow positron beam based on superconducting accelerator

The development of superconducting accelerator (SCA) provides a novel approach to generating pulsed slow positron beams with high time resolution and high intensity. SCAs can produce electron beams with repetition frequencies on the order of MHz and pulse widths of less than 100 ps. A pulsed positron beam based on SCA can, to a certain extent, preserve the excellent temporal structure of the primary electron beam. However, thermalization, diffusion, and surface re-emission of positrons in the moderator will inevitably lead to time broadening of positron pulses. In this paper, a calculation model coupling Geant4 Monte Carlo simulations with positron diffusion theory is established to evaluate the effect of the positron moderation process on time broadening. After being moderated by a tungsten foil, the initial time broadening of the pulsed slow positron beam is approximately 320 ps. By combining the time broadening calculation of pulsed beam during its transport, it is expected that the pulsed slow positron beam generated by the SCA, can be directly applied to the measurement of positron annihilation lifetime.

physics.acc-ph↗

Retrieval Sensitivity to Identity Signals in Queries

Dense retrievers decide which documents reach users and the language models that use them, yet they are typically evaluated with neutral queries. We ask whether the identity signals that real users express in their queries---political ideology and dialect---bias what a retriever returns. We design evaluations in two domains, political news and consumer-health questions, each pairing a controlled synthetic set that varies only the identity signal with naturalistic queries. Across five dense retrievers and a sparse baseline, every retriever (i) retrieves articles that align with the query's own political lean and (ii) performs worse for questions written in African American Language (AAL) than in White Mainstream English (WME). Two analyses tie these gaps to queries' identity signals beyond surface vocabulary: partialling out an aggregate lexical-asymmetry score leaves the synthetic gaps largely intact, and linear probes recover lean and dialect from the retrievers' query embeddings beyond token-level features. Left unaddressed, such retrieval biases risk contributing to polarization and reinforcing the health disparities already faced by AAL speakers. Code is available at https://github.com/Andrewtcr/bias-ret.

cs.CL↗

When Updating Stops Being Learning: Rethinking LLM Self-Evolution via learnable information gain

Self-evolution lets large language models (LLMs) improve iteratively using their own generated data, but often suffers from self-evolution degeneration: performance improves, plateaus, then declines. Existing methods address this issue at the component level, targeting either the Questioner or the Solver, and overlook that self-evolution is a tightly coupled system. We propose a holistic framework based on learnable information gain, which measures how much novel, parameterizable information a round provides relative to the previous round. Theoretically, this gain equals the Kullback-Leibler divergence between the two rounds' data distributions plus their entropy change. Practically, it is estimated by fitting a small language model to the previous round and scoring new data via negative log-likelihood. Based on this diagnostic, we propose ATRI (Adaptive Training Regulation via Information-gain), which reweights samples within a round and halts training across rounds when information gain remains low. Experiments on popular datasets demonstrate the superiority of our proposal.

cs.CL↗

Diffusion prior for KSTAR equilibrium reconstruction under sensor dropout

We study equilibrium reconstruction for the Korea Superconducting Tokamak Advanced Research (KSTAR) device under magnetic sensor dropout, with a diffusion model as the prior. Magnetic measurements leave most components of the toroidal current density $J_ϕ$ undetermined, and sensor loss leaves more of them to the prior. The diffusion model learns only $J_ϕ$, conditioned on the coil currents, the plasma current and the product of major radius and toroidal field, which do not depend on the dropout. The physics enters through a linear forward operator built from sensor response matrices of the LIUQE code. Because the observations are linear, a variant of decoupled annealing posterior sampling fits them with a closed-form linear correction. On measured signals of 37 KSTAR shots, we switch off a random fraction (0.0 to 0.9, ten settings) of the 124 magnetic channels in use and compare with LIUQE under the same masks, taking the full-sensor LIUQE reconstruction as the label. Both methods fit the remaining channels to a similar level. The median relative distance of $J_ϕ$ from the label stays at 3.22--4.07\% for the diffusion reconstruction in all settings, whereas that of LIUQE reaches 11.89\% at dropout 0.9. At dropout 0.5--0.9, the diffusion reconstruction is closer to the label on 27--37 of the 37 shots. At low dropout, LIUQE is closer on most shots. For the derived flux, the diffusion reconstruction is closer on 21--36 shots at dropout 0.5--0.9, and its error at high dropout comes from the vessel current estimate, not from the prior.

physics.plasm-ph↗