Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,783 records · Page 99Linked to original sources

A Fixed-Offset Transition for Random Stackability on Paths

We study a support-collapse version of graph pebbling on paths. A configuration is stackable if a sequence of legal pebbling moves can produce a nonzero configuration supported on a single vertex. On the path P_n, we choose a configuration uniformly from all weak compositions of total n times mu_n, where mu_n is a positive integer. We prove a two-sided fixed-offset transition for the logarithmic density. The transition is centred at sqrt(log_2 n) - (1/2) log_2 log_2 n + log_2(3e). For every fixed epsilon greater than zero, the stackability probability tends to zero when log_2 mu_n is eventually at most the centre minus epsilon, and tends to one when it is eventually at least the centre plus epsilon. No assertion is made at zero offset. The proof uses an exact recursive stackability score on trees, a one-dimensional path-message recurrence, binary-partition asymptotics for rare dyadic deficit excursions, a constant-cost regeneration argument, and an exact deep-message necessity theorem. Conditioning independent geometric occupancies on their sum returns the uniform fixed-total model. The finite deterministic necessity theorem and its exact fixed-total corollary are formalised in Lean and registered with Palomar; the full probabilistic asymptotic theorem is not part of that registration.

math.CO↗

Free Everywhere, Exact on Trees: PPO's Dropped Correction Buys Sample Efficiency Under Aggressive Reuse

Common policy improvement methods, including TRPO, PPO, and GRPO, estimate policy improvement under the behavioral policy's state-visitation distribution rather than the improved policy's own. The substitution makes the objective estimable from the behavioral policy's rollouts but adds a bias growing with policy divergence, hence the trust region or clip, and hence no reuse of a batch far off-policy. We show that under history-injective dynamics, where each state is reached by exactly one history, the dropped state-visitation ratio equals the product of per-step policy ratios along the sampled prefix, on every trajectory and not only in expectation. The ratio is therefore restored exactly, from log-probabilities PPO already computes. Autoregressive generation and canonical-order constructive optimization are both history-injective. The exact correction pays importance-sampling variance that grows with the horizon, so we generalize it to a one-parameter family with PPO ($α{=}0$) and the full correction ($α{=}1$) as endpoints: a single bias--variance knob. A gradient-level analysis of the unclipped surrogate identifies two channels the correction acts through and three conditions under which it carries signal; an enumerable testbed confirms the conditions' predictions. On hard credit-assignment scheduling tasks, a short corrected warmup with aggressive early sample reuse learns faster than PPO and than the same reuse uncorrected; the marginal gain grows with task difficulty ($+0.02$ to $+0.09$ learning-curve AUC), and the early win over PPO tracks the prefix bias that reuse incurs. A correction held throughout, or applied where clipping already contains the reuse bias, is null to harmful.

cs.LG↗

SAGE: Salient Factor Discovery and Generation with Visual Foundation Representations

Given a target dataset, such as faces with eyeglasses, and a background dataset, such as faces without, contrastive analysis separates \textit{salient} factors specific to the target from \textit{common} content shared by both. We aim for salient representations that capture target-specific detail in each image, such as the shape, color, and position of the glasses, so that they reveal subtypes without subtype labels and guide the generation of new examples of a discovered subtype, even one with no name or text description. We introduce SAGE, which learns both factors directly in the high-dimensional spatial latent of a frozen representation autoencoder and conditions a diffusion transformer on the learned salient representation of a reference image. On Digits-ImageNet and FFHQ eyeglasses, SAGE combines high-fidelity \textit{reconstruction} (rFID below $2$) with unsupervised \textit{subtype discovery}, recovering the digits better than baselines (probe accuracy $0.950$ vs.\ at most $0.281$) and revealing eyewear types, finer sunglasses styles, and mislabeled images; salient-conditioned \textit{generation} raises Digits-ImageNet subtype accuracy over the unfactorized latent ($90.5\%$ vs.\ $27.7\%$) and diversity on both datasets. On retinal OCT, SAGE's salient space separates three diseases using only normal/disease labels.

cs.CV↗

The exact number of nonconstant positive solutions to the planar isotropic $L_p$ dual Minkowski problem

We determine the exact number, up to rotations, of nonconstant positive $C^2$ solutions to the planar isotropic $L_p$ dual Minkowski problem for every $(p, q) \in \mathbb{R}^2$. We obtain a new parametrization of the associated period integral, characterize the regions where the period is strictly increasing or strictly decreasing, and prove that in the remaining nonmonotone region it has a unique nondegenerate maximum. As a consequence, we have a complete classification.

math.DG↗

Prediction is Better than Detection: Traffic Congestion Control using Drones

A central question in deploying teams of mobile robots for persistent monitoring is how task performance scales with fleet size, and whether this scaling holds once sensing drives downstream action rather than mere observation. We study this question for a team of drones performing traffic-jam detection and prediction in a simulated road network, whose reports drive an adaptive traffic-signal controller in closed loop. We build a multi-agent simulation, with vehicles following Nagel-Schreckenberg cellular-automaton dynamics and drones patrolling junctions via a round-robin policy, and sweep fleet size, traffic level, and network size to evaluate detection rate, detection delay, and prediction rate. We show how performance plateaus for fleet size approximating the number of junctions being monitored, and offer a general fleet-provisioning rule for persistent-monitoring deployments. More significantly, adapting the signal on a predicted jam, rather than a detected one, roughly doubles the resulting reduction in jam duration, showing that the value of onboard prediction in a sensing-to-action pipeline can exceed the value of adding more robots. Prediction accuracy, not sensing coverage, is now the binding constraint on further improvement, pointing to onboard inference, not fleet size, as the more promising direction for future work.

cs.RO↗

Neutrino Oscillations without Mass: A Re-analysis of KamLAND Data following the Dirac Equation in Curved Spacetime

The Dirac equation in stationary curved spacetime implies that describing two-flavor neutrino oscillations requires treating the invariant mass and conserved energy of each propagating state distinctly, as gravity affects energy differently than mass. By re-analysing the publicly released KamLAND dataset, we show that modeling flavor states of massless neutrinos as two-level quantum states with energy-dependent level splitting, analogous to modified Jaynes-Cummings models, results in a higher-likelihood fit to the observed data than the standard massive-neutrino paradigm.

gr-qc↗

Marginal Response Surface Elicitation for Zero-Label Tabular Learning

Tabular learning uses structured data to predict target outcomes. Traditionally, this process has relied on labeled data. However, large language models (LLMs) can be used to elicit domain priors based on the task description and feature semantics, thereby enabling predictions without labeled data. We propose Marginal Response Surface Elicitation (MARS), a method that transforms feature-level LLM priors into a reusable, zero-shot tabular classifier. To construct this classifier, MARS selects representative values for each feature from unlabeled data and prompts the LLM to provide corresponding class support scores and feature weights. It then aggregates multiple responses using the median to construct feature response functions, and makes predictions through their weighted sum without further LLM queries. Across eight tabular benchmark tasks, MARS achieves the highest average AUC and AP, outperforming direct prompting by 1.97 and 6.21 percentage points respectively, while substantially reducing end-to-end costs. Evaluations with LLMs of different sizes further demonstrate its predictive advantage over direct prompting.

cs.CL↗

Zero-Compute Cross-Lingual Transferability Estimation Using Typological Feature Proxies

Cross-lingual transfer describes how knowledge in a source language benefits a target language. Measuring it quantitatively requires broad multilingual pre-training, as prior work has done with cross-lingual transfer matrices. We ask whether transfer is predictable from freely available typological features, and whether the prominence of high-resource source languages reflects typology or data quality and quantity. We show that typological databases contain cheap and dense signals about cross-lingual transfer. Our typology-only random forest on a 24-language prior-work transfer matrix scores leave-one-language-out $ρ{=}0.705$ and $R^2{=}0.49$, beating a non-typological control at $ρ{=}0.62$, which verifies the ability of typology-only predictions to reconstruct costly measured cross-lingual transfer. The signal survives leave-one-script-out and leave-one-family-out protocols, so script and family confounding do not explain the effect. By decomposing the transfer into a typology term and a resource-and-script bias term, we find the best-source ranking sensitive to this bias. In contrast, typology is not affected by this bias, which makes it a zero-compute screening tool that replaces hundreds of training runs with a model fit. Our code is available \href{https://github.com/dharmsen/typo-x-ling-transfer}{here}.

cs.CL↗

A Math Circle in Third Grade

Problems developed for a mathematics circle conducted throughout the school year in a primary school class, involving all pupils without any selection, thus working with a heterogeneous group of children.

math.HO↗

On the linearised force balance condition in the analysis of atomistic dislocation models

A technical condition arising in the analysis of atomistic models for dislocations is studied. Key estimates for the far-field strain behaviour proved in Ehrlacher, Ortner, Shapeev (2016) rely on a summation-by-parts argument for linearised forces, under conditions of sufficient decay and of vanishing net force in an infinite system. In particular, the latter condition is required for an application of this theory to the standard far-field dislocation predictor, but the vanishing of the associated force sum has not been fully justified. Here, the missing verification of this condition is provided and its role in determining the far-field decay of the corrector is clarified. It is moreover shown that for more general physically compatible predictors, the linearised force sum need not vanish, leading to slower decay and potentially divergent finite-domain approximations.

math.NA↗

Explicit Robin Green's Functions, Resonance Spectra, and Impedance Recovery on Annuli and Spherical Shells

This paper constructs explicit closed-form Green's functions for the Helmholtz equation on annular and spherical-shell domains with two independent Robin impedances on the inner and outer boundaries. Graf's and the hyperspherical addition theorems reduce each angular mode to an explicit $2\times2$ linear system, and the resonance spectrum is governed by a characteristic determinant bilinear in the two impedances. This bilinearity has a direct inverse-problem consequence: two resonant frequencies of one non-radial angular mode generate at most two candidate impedance pairs via an explicit quadratic equation, a third resonance selects the physical pair, and an exact reflection-symmetry obstruction identifies where the radial spherical mode cannot recover the impedances. Spectrally, we prove the branches positive, simple and strictly increasing in both impedances; derive first-order asymptotics at the four corners of the impedance plane, with coefficients given by boundary masses and normal derivatives of the limiting eigenfunctions, and an explicit mixed second-order coefficient at the Dirichlet--Dirichlet corner, which for the radial mode evaluates in closed form to $2π/(R_2-R_1)^3$; establish a low-frequency resonance-free band, with a second-order threshold expansion explicit in dimension three and a rational approximation accurate over the whole impedance range; and prove the universal high-frequency spacing law with a shell-curvature correction. Jacobian-based sensitivity and conditioning criteria are included. The kernels and spectra are computable to machine precision, all asymptotic regimes are confirmed numerically, and the kernels provide reference solutions for finite-element and boundary-element validation.

math.AP↗

RiboUnmix: Learning Shared Translational Dynamics from Biased and Noisy Ribo-seq Measurements

Ribosome profiling (Ribo-seq) measures ribosome distributions along mRNAs, but observed occupancy profiles also contain experiment-specific distortions and stochastic variability. Consequently, models that accurately predict measured profiles may reproduce technical effects rather than recover the underlying biology. We ask whether jointly modeling datasets collected under different experimental conditions can reveal shared, sequence-dependent patterns of ribosome occupancy. We introduce RiboUnmix, a probabilistic multi-dataset framework in which each expected measured profile is represented as a shared sequence-dependent signal modulated by a dataset-specific multiplicative factor. A negative-binomial observation model captures variability across replicates. We evaluate RiboUnmix on a controlled synthetic benchmark combining programmed translation kinetics, ribosome traffic, stochastic count sampling, and sequence-dependent experimental distortions. Because the underlying kinetics and distortions are known, recovery of the shared profile and dataset-specific effects can be assessed separately. Both inferred components correlate strongly with their targets, demonstrating that RiboUnmix can disentangle shared kinetic patterns from experimental effects. Across four organism-specific real-data benchmarks, RiboUnmix outperforms sequence-to-profile baselines in predicting measured profiles. Models trained independently on subsets of 114 HEK-derived datasets recover concordant shared profiles for held-out transcripts, and experiments varying the number and composition of training datasets show that the learned representation remains stable. RiboUnmix thus converts variation across experiments into evidence for reproducible sequence-dependent patterns of ribosome occupancy, supporting biological hypothesis generation from diverse Ribo-seq datasets.

cs.LG↗

SEPAL: Separated Expert Pairs with Answer-Level Fusion for Reliable LLM Collaboration

Multi-agent collaboration lets large language models (LLMs) improve question answering through deliberation and feedback. Yet shared discussion couples correction with exposure to the same mistakes, which can erode the diversity needed for voting. Self-consistency offers sampling diversity without feedback, while single-pair Actor-Critic collaboration refines only one candidate. We introduce SEPAL, which assigns three private Actor-Critic teams to direct reasoning, evidence grounding, and verification. Role-specific training gives the teams different reasoning objectives beyond sampling variation. Each Critic guides revisions within its own team, preventing feedback from carrying errors across candidates. Once revision ends, majority voting combines only the final answers, keeping the reasoning histories separate until the decision. Across five open-weight backbones and five question-answering benchmarks, SEPAL improves mean accuracy by 1.81 percentage points over a matched single Actor-Critic pair, with improvements across all five backbones. Code is available at https://github.com/zhansan114514/SEPAL.

cs.CL↗

Beyond Uniform Compression: Budgeted Transmission Allocation for Extreme Federated Learning

Federated learning faces severe communication bottlenecks when clients upload high-dimensional model updates. Existing methods often compress these updates uniformly across all layers. This uniform approach ignores the heterogeneous value of different parameter blocks and wastes limited bandwidth on insensitive layers. To address this issue, we propose Layer-wise Budgeted Adaptive Transmission (LBAT). LBAT reframes federated communication under extreme uplink budgets as a resource allocation problem. Our framework dynamically estimates the transmission value of different layers utilising local training signals. It then employs an exact byte dynamic programming allocator to determine optimal rank and bit configurations under strict budgets. We validate LBAT on highly heterogeneous federated tabular prediction and data generation tasks. Extensive experiments demonstrate that LBAT consistently outperforms uniform rank, uniform quantisation, and fixed compression baselines across various extreme budget regimes. Furthermore, it achieves significantly better communication and utility tradeoffs while preserving essential distributional fidelity.

cs.LG↗

Diffractive Vector-Meson Production from the Light-Front Quark Model

We study exclusive diffractive electroproduction of the $ρ^0$, $ϕ$, $J/ψ$, and $ψ(2S)$ mesons in the color-dipole picture using two spin-orbit light-front wave functions, S-1 and S-2. S-1 follows from the Melosh--Wigner rotation in the light-front quark model (LFQM), while S-2 is the spin-improved form commonly used in holographic and boosted-Gaussian approaches. Both share the same radial wave function, with their parameters determined from the LFQM Hamiltonian. The color-glass-condensate dipole parameters are fixed from inclusive structure-function data, with no fit to diffractive data. We compare the predictions with measurements of cross sections, $R=σ_L/σ_T$, $t$ and $W$ dependences, and $σ_ϕ/σ_ρ$. S-1 gives a better description of the light-meson observables, while S-2 provides a better description of the charmonium ratio $σ_{ψ(2S)}/σ_{J/ψ}$ and the $J/ψ$ decay constant. We also present the charge, magnetic, and quadrupole form factors of the four mesons and compare them with available lattice-QCD results.

hep-ph↗

From Modes to Memories: Characterizing the Scale-Space Dynamics of Diffusion Models

Diffusion models are typically viewed as stochastic processes that transform noise into data. We take a complementary perspective: a diffusion model defines a family of deterministic dynamical systems indexed by noise scale. At each fixed scale $σ$, we treat the denoiser as a self-map and study its dynamics. For an exact denoiser, fixed points correspond to critical points of the smoothed data density, while attractors correspond to its modes; as $σ$ increases, sample-level modes merge into progressively coarser ones. This suggests a geometric view of memorization: examples that receive excess probability mass due to duplication or overfitting, as well as outliers, should remain distinguishable under stronger smoothing than ordinary examples. We quantify this persistence by the critical scale $σ_c$, the largest noise scale at which an example is retained by the fixed-scale dynamics. In conditional models, the same construction extends naturally to image--caption pairs. Experiments in controlled settings and on large-scale models show that $σ_c$ tracks memorization arising from duplication, overfitting, and outliers, and identifies both memorized and partially memorized examples in Stable Diffusion. Moreover, $σ_c$ yields interpretable measures of the image spatial distribution and caption dependence of memorization.

cs.LG↗

FANVIDv2: Evaluating Video Super-Resolution by Face and Licence-Plate Recognition Under Compound Degradation

Video super-resolution (VSR) is normally judged by PSNR and SSIM on clips that were downsampled bicubically, although in surveillance its purpose is to make faces and licence plates \emph{recognisable}. We present FANVIDv2, a benchmark that scores VSR by what a recognition pipeline can do with its output. FANVIDv2 provides $320\times180$ low-resolution (LR) clips with high-resolution (HR) references for 48 public figures (with one HR gallery image each) and 375 licence-plate clips covering 360 distinct plate strings. LR clips are generated with a randomised compound degradation (blur, resize jitter, sensor noise, JPEG compression, final downsampling) rather than bicubic downsampling alone. Two metrics score recognition \emph{inside} detections: FaceRecBox rewards a face only if it is localised and correctly identified, and TextRecBox scores plate transcriptions by normalised edit distance weighted by localisation quality. With a 2.3\,M-parameter VSR baseline (RCDM), FaceRecBox rises from 0.6864 to 0.7222, identity accuracy on matched faces from 84.35\% to 86.93\%, and TextRecBox from 0.3088 to 0.3667; a residual-map gated variant (RCDM-RMGF) reaches 0.3801 on plates. We describe the degradation model, the baseline architectures and the scorers in detail, and release annotations, metadata, download and degradation scripts and evaluation code.

cs.CV↗

Phase-sensitive avalanche quantum sensing of sub-shot-noise fields

Avalanche-based detection of small perturbations is commonplace in precision measurement devices from Geiger counters to single-photon avalanche detectors. Here, we expand this principle to sensing of light waves below the quantum shot noise limit, demonstrating numerically how atomic clouds trapped in photonic cavities can exhibit non-perturbative sensitivity to changes in the cavity field when the cavity mode is tuned to a parity-prohibited transition. We find that, when pumped by a continuous-wave laser, the atomic cloud creates emissions which amplify the initially sub-shot-noise fluctuation by two orders of magnitude. Crucially, this amplification is broadly tunable with respect to frequency and independent of atomic energy structure. Even more strikingly, our amplification mechanism preserves the phase imparted on the atoms by the original ultra-weak light wave, marking a qualitative improvement over conventional protocols. Our finding has implications for both quantum state characterization and detection of weak classical signals.

quant-ph↗