Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,351 records · Page 75Linked to original sources

Oort's conjecture for split unitary Shimura varieties

We prove that generically on the basic stratum of split unitary Shimura varieties, the universal abelian variety with endomorphism structure and polarization has minimal automorphism group in a precise sense, except for a few degenerate cases. This is a direct analogue of Oort's conjecture on automorphisms of supersingular abelian varieties. On the way, we explicitly compute the generic automorphism group of the universal $p$-divisible group in a given basic isogeny class.

math.AG↗

Last Translation Benchmark

For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for machine translation are approaching saturation. Further, automatic translation metrics are unreliable, opaque, and vulnerable to reward-hacking. Even gold human evaluation is not problem-free, because it often lacks reproducibility, objectivity, and scalability. Overall, this prevents us from tracking progress in the field and identifying pathways for improvement. We introduce the Last Translation Benchmark, a collection of human-authored and peer-reviewed examples (texts, images, audio, videos) that break leading machine translation models. We also present a new evaluation approach: each example comes with handcrafted verification rules describing concrete failure cases on that example, therefore allowing reliable and actionable future evaluation. The Last Translation Benchmark is a live dataset that accepts ongoing contributions. The latest version is LTBv1, containing accepted contributions prior to September 1st 2026, with future releases planned as new data is continuously collected.

cs.CL↗

SMILE: Bridging Continuous Optimization and Discrete Symbolic Recovery

Symbolic regression (SR) discovers closed-form mathematical expressions from data, offering interpretability beyond black-box models. Existing methods suffer from slow convergence in combinatorial search spaces and lack mechanisms to exploit compositional structure in the data. We introduce SMILE (Sine, Multiplication, Identity, Logarithm, Exponential), a hybrid framework that unifies continuous gradient-based optimization with discrete symbolic recovery through three stages: structural analysis of the data to identify the compositional hierarchy of the target expression, continuous optimization to learn parameters of a network that encodes the target expression using interpretable activations, and symbolic recovery through structured pruning, coefficient optimization, and rounding. This final stage distills the learned network into a compact expression with exact symbolic constants. We evaluate SMILE on SRBench across ground-truth and black-box datasets, with ablation studies validating each component. SMILE achieves the highest symbolic solution rate at the largest noise levels, demonstrating strong robustness where competing methods degrade substantially. It consistently lies on the Pareto front of accuracy versus complexity, recovering significantly simpler expressions in a fraction of the time required by the competing methods.

cs.LG↗

Linguistic Trajectory Encoding for Efficient Long-Horizon Spatial Memory in Embodied Agents

Embodied agents performing long-horizon tasks require a memory representation in which the state transitions of dynamic objects remain queryable in natural language across hours-to-days observation horizons. Existing systems either drop fine-grained motion (clip-level video-language embeddings), keep it only as raw coordinates (geometric SLAM), or organise it around immediate task context (agent working memories). None of them gives the agent a per-object timeline whose state transitions are themselves queryable in language. Our key contribution is \textbf{Linguistic Trajectory Encoding} (LTE), which compresses dynamic object motion histories via a hybrid representation combining natural language descriptions, sparse spatial anchors, and visual anchors. LTE adapts compression to motion complexity by anchoring periods without reliable observations to the last seen location, while representing motion with geometric waypoints and linguistic descriptions to preserve accuracy. To evaluate these capabilities across extended time horizons, we construct the \textbf{Spatial Memory Benchmark} (SMB) from EgoLife multi-day recordings, targeting capabilities absent in existing benchmarks: semantic trajectory retrieval and long-horizon object retrieval. On SMB, the LTE-based system achieves $45.3\%$ success in semantic trajectory retrieval and $48.7\%$ in long-horizon object retrieval, outperforming structured-memory and VLM baselines (best prior: $31.9\%$ and $34.4\%$). LTE achieves trajectory compression by factors of $8.7\times$ to $26.1\times$ with sub-second query latency on $24$\,h video. On Ego4D natural-language queries, the system reaches $28.75\%$ / $55.10\%$ R@1/R@5, $+15.80$ / $+31.30$ pts over EgoVLPv2.

cs.CV↗

SimFuse3D: Source-Guided Target Simulation and Confidence-Guided Multi-Stage Localization Reweighting for Cross-Platform 3D Object Detection

Changes in sensor height and viewpoint alter object-level point distributions, making cross-platform LiDAR unsupervised domain adaptation (UDA) difficult. Self-training uses labeled source scans and unlabeled target scans, yet a retained prediction may provide a useful target location while enclosing sparse foreground returns, background clutter, or points inconsistent with the predicted box. We refer to this mismatch as box-point inconsistency. We introduce SimFuse3D, which preserves the target placement and repairs the associated pseudo object using measured geometry from labeled source scans. Object Memory retrieves a similar labeled source instance. Target Simulation places the retrieved source geometry at the target location, aligns its points with the target viewing geometry, and filters the aligned crop to approximate the target observation. Confidence-Guided Multi-Stage Localization Reweighting (CMLR) maps each target pseudo-object confidence score to a bounded weight shared by RPN localization and R-CNN box regression. All components operate only during adaptation, leaving the detector architecture and inference graph unchanged. Across six cross-platform transfers, SimFuse3D consistently outperforms Pi3DET-Net and achieves the best performance among the compared adaptation methods on nearly all metrics. On nuScenes-to-KITTI, it ranks first among the compared adaptation methods with both evaluated detectors.

cs.CV↗

Molecular Déjà Vu: Digit-Level Retrieval of Molecular Properties in Frontier Language Models

Large language models (LLMs) are increasingly employed to predict molecular properties. However, prediction error alone cannot distinguish prediction from retrieval of published values. We audit 22 frontier models on 12 molecular regression datasets in a zero-shot setting, assessed against a molecule-blind reference derived from each dataset's labels. Significant retrieval is concentrated on 5 datasets, with isolated flagged LLMs elsewhere. Increasing the reasoning setting raises the number of flagged model--dataset combinations from 47 to 89 of 264. An in-context blinding experiment reduces retrieval but leaves a quarter of the combinations flagged. Blinding changes model rankings and increases errors. Because blinding also removes chemically interpretable structure, the error increase can only be partially attributed to reduced retrieval.

cs.AI↗

Counting sets with given doubling via dimension

We determine, up to a factor of $2^{o(k)}$, the number of $k$-sets $A \subset \{1, \ldots, n\}$ such that $|A + A| \leq m$, where $k = Θ(\log n)$ and $m \leq k^{1 + α}$, for small $α> 0$, answering a question of Green and Morris.

math.CO↗

FP-DeErr: Application-Oriented Error Decomposition for Foundation Potentials

Foundation potentials (FPs) have emerged as a new basis for atomistic modeling. While their evaluation using average energy and force errors often indicates near-DFT accuracy, their performance in practical computational studies remains inconsistent. While recent benchmarks evaluate FPs on downstream computational tasks, the underlying errors that determine task success or failure are not always clear. Here, we develop FP-DeErr, an application-oriented error-decomposition framework that resolves FP errors according to the physically meaningful quantities, configurations, and computational stages governing specific tasks. Rather than averaging errors over an entire dataset, FP-DeErr uses decomposed error metrics devised for physically meaningful quantities, focusing on the configurations where errors arise and matter most. We demonstrate FP-DeErr by evaluating multiple state-of-the-art FPs on three fundamental tasks: force prediction for atomistic simulations, energy ranking for substitutional and vacancy orderings, and ion/vacancy migration. These error-decomposition metrics resolve force errors among highly accurate, large-error, and far-from-equilibrium atoms; relative-energy errors among competing orderings, phases, and compositions; and ion migration errors among endpoints and along-path errors. By identifying where FP errors arise, this error decomposition provides targeted guidance for FP development. FP-DeErr also provides an open benchmark, evaluation code, and a public leaderboard for rigorous FP assessment.

cond-mat.mtrl-sci↗

Safety Monitors Mostly Catch What the Model Already Refuses

Safety monitors are evaluated by recall on harmful prompts, regardless of whether the target model would answer them. Yet a monitor matters most on the prompts the model does answer. We measure recall on exactly those prompts, defined by sampling the target model and judging its responses. Across four text guards, two activation probes, and Latent Guard, recall at a 1% false positive rate falls sharply on this subset: at a common threshold, every monitor catches the requests the model refuses 1.1 to 6.4 times as often as the requests it answers. Standard metrics hide this; AUROC stays above 0.85 for most monitors. Rewriting each request to be less explicit, with intent held fixed and verified, raises compliance 28-fold and lowers every monitor's flag rate. Of the requests newly answered after rewriting, 44 to 93% slip past the monitor, depending on which is used, and most of their completions are graded harmful. We trace the gap to explicitness itself. As wording softens with intent fixed, the target model's harm and refusal readings fall and it answers; the guards' harm readings fall too, and steering a guard along explicitness alone flips its verdict. Model and monitors miss the same prompts, and stacking monitors does not recover them. Fine-tuning a guard on the rewrites at every level of explicitness, on both sides of the label, raises recall on answered requests from .24 to .89 while transferring to unseen benchmarks; hard negatives, the natural alternative, teach the guard to discount indirect phrasing instead.

cs.CL↗

Beyond Retraining-Free MoE Compression: A Cost-Normalized Study of Post-Compression Adjustment

Retraining-free MoE compression reduces deployment memory by pruning or merging experts, but often treats the compressed checkpoint as the final artifact. We argue that this view is incomplete: compressed MoE checkpoints are better understood as compressed initializations that benefit from a tiny post-compression adjustment stage. Across two MoE LLM backbones, four pruning/merging methods, three expert-retention ratios, and 28 benchmarks, we compare LM fine-tuning and teacher-based KD under matched small-data budgets and measured GPU costs. Using only 3,000 C4 examples and a single epoch of adjustment, Full FT recovers 37.3% of the original-to-compressed performance gap on average. Moreover, LM fine-tuning is more cost-effective than standard token-level KD, and full-parameter adjustment gives the strongest cost--recovery trade-off among the tested scopes. These results suggest that retraining-free compression should be paired with small post-compression adjustment to recover a substantial portion of the performance lost during compression.

cs.LG↗

CST-WM: A Causally Structured World Model for Embodied Visual Tracking

Embodied visual tracking requires a robot to choose actions that keep a moving target observable at a suitable distance, and to recover it after occlusion, out-of-view drift, or distractor crossings. We cast the task as planning over future target evidence with an action-conditioned world model. In logged tracking data, however, the behavior policy's actions are correlated with where the target is, so a generic predictor can learn a shortcut: it writes the current action directly into its prediction of target evidence, instead of letting the action affect that evidence only by moving the robot and changing what it observes. We call this failure causal hallucination; the resulting rollouts look plausible but rank candidate actions for the wrong reason. We propose CST-WM, a causally structured world model whose state is split into target-evidence, robot, and observation branches. Its transition removes the same-step edge from action to target evidence but keeps the path through robot motion and the resulting views, so candidate actions are still distinguished by their predicted ego-motion. With rollout-based model-predictive control, a single model handles both steady following and re-acquisition after target loss. On EVT-Bench and Habitat 3.0, covering standard tracking, target-loss recovery, and cross-dataset transfer, CST-WM improves following, distance-range control, safety, and re-acquisition over reactive trackers and world-model baselines, and removing the action mask causes the largest drop in re-acquisition among our ablations. Offline, CST-WM has lower multi-step rollout error, and its ranking of candidate actions agrees better with the simulator's. On a Unitree Go2 quadruped, CST-WM succeeds in 20 of 30 real-world trials under occlusion, distractor crossing, and fast motion, against 14 for TrackVLA.

cs.CV↗

A finite set with noncompact closed convex hull in a Hadamard space

We show that there is a Hadamard space in which the closed convex hull of a set containing merely three points is noncompact, by providing an AI-generated construction. The seemingly innocuous problem of finding a compact or finite set with this property, or showing that none exists, was recorded by Gromov and periodically revisited by researchers in metric geometry, functional analysis, fixed-point theory, and geometric group theory. Nevertheless, it remained unanswered for over three decades. The construction itself is based on building an inductive sequence of CAT(1) metric graphs and taking the Euclidean cone over the limiting space. We discuss the history of this problem, the main ideas and details of the construction, and its relationship with existing techniques.

math.MG↗

Smoothed Picard Hamiltonian Monte Carlo

We develop a new low-accuracy sampler, called smoothed Picard Hamiltonian Monte Carlo, which combines Gaussian smoothing, Picard iteration, and higher-order discretization. For a log-concave target $π\propto \exp(-V)$ in dimension $d$ satisfying $0 \prec αI \preceq \nabla^2 V \preceq βI$, with condition number $κ:= β/α$, smoothed Picard HMC returns a sample with $\sqrt α\,W_2(\cdot,π) \le \varepsilon$ using $\widetilde O(κ^2 + κ^{7/6} d^{1/6}/\varepsilon^{1/3})$ gradient queries. We also prove stronger $W_q$ bounds, and then develop an algorithmic framework, the recursive warm start generator, to upgrade these $W_q$ bounds to stronger divergence guarantees. This produces a warm start for the proximal bouncy particle sampler, introduced in a companion work, leading to a high-accuracy log-concave sampler with complexity $\widetilde O((κ^{7/6} d^{1/6} + κ^{1/2} d^{1/4})\mathrm{polylog}(1/\varepsilon))$.

math.ST↗

Complexity Amplification from Compression in Quantum Random Access Optimization

Compressed quantum encodings aim to overcome hardware limitations towards tackling challenging problems at scale, with many classical variables mapped onto noncommuting observables of fewer qubits. Classically, relaxations such as the semidefinite program formulation of MaxCut trade solution quality for computational efficiency. By contrast, quantum relaxations based on compression can amplify the worst-case complexity of the problem being solved. We study quantum random access optimization (QRAO), a special case of the Pauli correlation encoding (PCE) framework that assigns up to three binary variables to the Pauli $X$, $Y$, and $Z$ observables of each qubit, with the packing choices determining the compressed Hamiltonian to be optimized. We identify explicit QRAO optimal energy promise problems complete for NP, StoqMA, and QMA, with inverse-polynomial promise gaps for the latter two. Our problem reductions preserve inverse-polynomial promise gaps without requiring gadgets or ancillas. For any prescribed packing, we show that weighted MaxCut instances compress, up to a known shift and rescaling, to arbitrary nonnegative-weight pairwise Pauli couplings allowed by the packing. For QRAO, using one aligned axis gives an NP-complete energy problem. Using two or three positive aligned Pauli axes generally gives QMA-complete problems, with bipartite restrictions in BQP $\cap$ StoqMA. We show that this computational hardness survives compilation and is practically relevant. Notably, this result applies directly to the current QRAO compiler implementation in Qiskit Optimization 0.7.0, confirming our hardness results are not artifacts of artificial or contrived packing rules. Altogether our results identify worst-case complexity barriers arising from quantum compression, while making no broad claims about typical cases or the performance and trainability of algorithm pipelines that use it.

quant-ph↗

Boundedness of Multilinear Hilbert Transforms Along Moment Curves

We prove the boundedness of multilinear Hilbert transforms along moment curves by establishing uniform estimates for both translated and untranslated multilinear paraproducts and by applying a Sobolev smoothing inequality. As a byproduct of our method, we obtain the same bounds for multilinear maximal functions along moment curves.

math.CA↗

X-DigCheck: Co-Evolving Application Profiles and Knowledge Graphs, Demonstrated on the RTI Documentation of Rupe Magna

We demonstrate X-DigCheck, a domain-independent environment for building and maintaining application profiles as they co-evolve with the data they describe. Profiles developed against a fixed ontology quickly drift from the schema they were meant to capture. X-DigCheck treats profile construction as a continuous ontology-data co-evolution loop: data are lifted into RDF against the profile, checked through competency questions and SHACL, and the resulting reports jointly drive revisions of the ontology, mappings, constraints, and graph. The loop is agnostic to the domain and to the pipeline that produces the graph. We validate and demonstrate the tool in the cultural heritage domain, on the construction of RupeMagna-RTI, the first Reflectance Transformation Imaging (RTI) specialisation of the Cultural Heritage Survey ODP (CHS-ODP), aligned with CIDOC-CRM/CRMdig, ArCo, CHAD-KG, and Getty AAT, with semRTI as the lifting pipeline of this use case. The demonstration lets visitors run one full turn of the loop -on the shipped Rupe Magna (Grosio, Italy) RTI survey, or on a profile and graph of their own -executing the competency-question and SHACL checks live and reading the bidirectional coverage report that flags modelling gaps and stale assumptions. The result is a portable co-evolution environment for profile engineering, together with a reusable RTI application profile produced through it. A screencast of the demonstration is available at https://zenodo.org/records/22210609.

cs.DB↗

Kalman Delta Networks: Uncertainty-aware Associative Memory

Linear attention enables efficient long-context inference by compressing token history into a fixed-size recurrent memory. This compression makes each update a trade-off between incorporating new information and preserving useful associations. Models such as DeltaNet, Gated DeltaNet, and KDA predict write strength from the current token representation, without explicitly tracking uncertainty in the stored memory. Yet this uncertainty matters: a new observation should have greater influence when the existing association is uncertain and less when it is already well supported. We introduce Kalman Delta Networks (KDNs), a family of linear-attention models that explicitly track memory uncertainty to guide each update. By formulating associative memory as a linear-Gaussian state-space model, KDNs propagate both the memory estimate and its uncertainty, using the Kalman gain to balance accumulated evidence against the reliability of new observations. This formulation also recovers standard delta-rule updates by replacing tracked covariance with a token-predicted isotropic surrogate. To support hardware-efficient training and inference, we derive Diagonal KDN and Isotropic KDN, which retain one uncertainty value per key channel and per head, respectively. Their uncertainty updates admit associative scans with logarithmic parallel depth, requiring only $O(d_k)$ and $O(1)$ auxiliary state per head. Across controlled pretraining at 750M and 1.3B parameters, both variants consistently improve perplexity and mean downstream accuracy over the evaluated state-of-the-art linear-attention baselines.

cs.LG↗

Bogomol'nyi equations for Kruglov strings

We construct the Bogomol'nyi equations for Abelian gauge--Higgs vortices in which the Maxwell gauge sector is replaced by Kruglov nonlinear electrodynamics, a power-law family that interpolates between Maxwell theory, Born--Infeld electrodynamics, and exponential electrodynamics, characterized by a dimensionless exponent $σ$. Using the stressless method, we derive a pair of first-order equations directly from the vanishing of the spatial stress tensor, without assuming the Higgs potential \textit{a priori}. For generic $σ$, the gauge and Higgs sectors are coupled through an implicit algebraic relation. We therefore introduce a constitutive map $Φ(Y;σ)$ and analyze its monotonicity and range to determine the conditions for a smooth admissible Bogomol'nyi branch. For $σ>1/2$, the constitutive map is strictly monotonic and unbounded, whereas for $0<σ<1/2$ it possesses a finite maximum; the marginal case $σ=1/2$ is bounded. These properties yield explicit bounds on the nonlinear parameter $β$ for the latter cases. We further obtain closed-form constitutive relations, BPS potentials, and gauge-field equations for six representative values of $σ$, spanning linear, quadratic, and cubic algebraic structures. The corresponding vortex profiles are then computed numerically. The resulting BPS string tension is purely topological, $μ_{\rm BPS}=2πn$, independent of both $σ$ and $β$.

hep-th↗