Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 667 records · Page 37Linked to original sources

From Semantic Decisions to Feasible Trajectories: Self-Evolving LLM-Guided Optimal Control for Narrow-Space Parking

Autonomous parking in nonconvex and narrow environments remains challenging. Although optimal-control methods can explicitly enforce vehicle dynamics and collision constraints, nonconvexity compromises solver robustness and can cause failures. Large language models (LLMs) exhibit strong semantic reasoning capabilities, but directly generating dense trajectories makes it difficult to guarantee physical feasibility. We introduce SE-LLM-OCP, a unified framework in which LLMs make high-level discrete maneuver decisions, while an optimal-control module enforces low-level vehicle dynamics and collision constraints. Online, the LLM proposes sparse maneuver plans, decomposing the parking task into a sequence of short-horizon trajectory-optimization problems. A low-level solver then sequentially solves optimal-control problems. If the solver fails, the LLM aggregates failure evidence from the solver and validation stages to guide replanning. Offline, SE-LLM-OCP automatically evolves a structured decision-making knowledge base from scratch, driven by accumulated online failures. We validate our proposed framework in simulation on a car-like vehicle model and on a differential-drive robot. Our experimental results show that SE-LLM-OCP enables safer autonomous parking in narrow scenarios and demonstrates transfer of the same maneuver representation to a different kinematic platform.

cs.RO↗

Superconducting qubit based on altermagnets

Altermagnets, characterized by vanishing net magnetization and momentum-dependent spin splitting, provide a promising platform for next-generation Josephson devices. Here, we exploit the Josephson effect in superconductor-altermagnet-superconductor junctions and show how to engineer prescribed current-phase relations by device design. Based on these programmable Josephson potentials utilizing altermagnetism, we propose a new class of superconducting qubits that combine large anharmonicity with enhanced robustness against decoherence via coherent two-Cooper-pair tunneling. We show that in the $2ϕ$-junction regime, this kind of qubit is intrinsically protected against both charge and flux noise due to parity protection. Magnetic flux can be used to precisely control the qubit and, under appropriate bias, this architecture further suppresses charge and flux noise. Our results establish altermagnets as a versatile platform for Josephson-potential engineering and open a new route toward high-performance superconducting qubits combining high coherence, large anharmonicity, and broad tunability.

quant-ph↗

Anomalously enhanced lifetimes of low angular momentum Rydberg states in singly charged alkaline-earth metal ions

Trapped ions excited to high-lying electronic states, so-called Rydberg states, open new opportunities for quantum simulation and quantum computing. Generally, the fidelity of quantum coherent operations critically depends on the longevity of Rydberg states. However, scaling laws predict that the lifetimes of Rydberg states in singly charged alkaline-earth metal ions are 16 times shorter, compared to their neutral atom counterparts. Here, we show that this is not generally the case. We report an anomalous lifetime enhancement of certain low angular momentum ionic Rydberg series by factors larger than eight. The anomaly is present at both zero and finite temperature, although it is caused by different mechanisms. At zero temperature, the anomalously enhanced lifetimes are caused by accidental cancellations of the relevant dipole transition matrix elements, while at room temperature the anomaly originates from the enlarged energetic separation of ionic Rydberg levels with respect to neutral-atom levels.

physics.atom-ph↗

The Copy Ceiling: An Input-Exposure Control for Ontology-Grounded Generation over Curated Corpora

We built a node that grounds a replaceable language model in a maintained ontology corpus, then asked what its successful-looking evaluation could support. Across ten models, grounding raised target-name recall from 0.265 unaided to about 0.92. A copy baseline, the recall a verbatim copy of the shown context already achieves, scores 0.964, and every model sits 0.022 to 0.067 below it. Copying therefore scores higher on this limited recall measure, which does not assess whether answers are better. The comparison tests what a recall score establishes; it does not test whether reasoning occurred, because a reasoned answer and a copy score alike when the answer name is already in context. We report exposure accounting (four counts classifying each gold item by whether the context exposed it and the answer recovered it) and a model-judged audit of 423 sampled item observations. A separate paired production study found a model-judged quality gain of +0.27 [+0.11, +0.45] on a 0-5 scale. Operational studies found failures that recall alone would not show: rephrasing questions out of the graph's vocabulary cut exposure from 0.964 to 0.328, yet the absence-keyed fallback would have fired on only 2 of 506; and inserting extracted facts degraded judged pages in every arm, so that step was disabled. Five-arm controls show that any well-formed on-corpus block beats no context but do not establish that the specific content matters, and no matched comparison against flat-text retrieval was run. The corpus is public and largely LLM-generated, which establishes neither training exposure nor novelty. Each study has its own outcome measure. Where gold derives from the injected corpus, we recommend reporting the accounting beside quality judgements, not in place of them.

cs.CL↗

Improving the Loss Tolerance of Heralded Photonic GHZ States for Long-Distance Device-Independent Conference Key Agreement

Heralded multipartite entanglement distribution is a key requirement for device-independent conference key agreement (DI-CKA) over lossy quantum networks. Although locally equivalent in the absence of loss, different single-rail photon-number encodings of Greenberger--Horne--Zeilinger (GHZ) states respond differently to photon loss. Here, we investigate the critical detection efficiencies for detection-loophole-free parity--CHSH violations of computational-basis GHZ states---a coherent superposition of the vacuum and an $n$-photon component---and of fixed-photon-number GHZ states, deriving exact analytical conditions for both. We show that for states that are not permutation symmetric, such as the latter, the assignment of measurement roles to physical modes affects loss tolerance. We introduce a star-network protocol employing heterogeneous sources to directly herald the loss-tolerant vacuum-$n$-photon GHZ states while retaining the favourable long-distance scaling $O(η_{\text{c}}^{n/2})$, where $η_{\text{c}}$ is the channel transmittance. For four users, we characterize the heralded state under photon loss and show that tunable source parameters allow genuine multipartite entanglement to persist at any finite channel distance. With ideal Pauli and displacement-based measurements, our protocol achieves positive DI-CKA key rates at lower detection efficiencies than previous schemes, while retaining comparable or greater rates and communication distances at high efficiency. Overall, our work improves the loss tolerance of heralded photonic GHZ states for DI-CKA both by directly heralding a more loss-tolerant encoding and by optimizing existing schemes. These results identify photon-number encoding, source architecture, and measurement-role assignment as key design parameters for loss-resilient multipartite quantum networks, offering a practical route toward near-term DI-CKA.

quant-ph↗

VideoGen-Agent: Reinforcing Video Generation Agents

Recent advances in video generative models have enabled high-fidelity, temporally coherent video generation. However, these models often struggle to satisfy prompts requiring specialized knowledge, specific identities, physical consistency, or ordered events. In this paper, we present VideoGen-Agent, a multimodal agent trained through multitask agentic reinforcement learning to use external tools for video generation. The agent coordinates augmentation, generation, and verification tools through multi-turn interactions, using the prompt and intermediate observations to guide its decisions. We train a shared policy on a category-balanced dataset spanning six tasks. Supervised fine-tuning on teacher-generated trajectories establishes tool-use behavior, which is then refined through reinforcement learning. A category-aware hybrid reward evaluates tool-call validity, task-appropriate tool use, and generated video quality. We further introduce VABench, a held-out benchmark of 600 prompts covering procedural knowledge, single- and multi-entity identity preservation, physical consistency, scene composition, and multi-shot temporal structure. On VABench, VideoGen-Agent improves over its base text-to-video generator by 19.1 points, from 56.5 to 75.6. Upgrading the generation tools further raises the score to 86.1 without additional agent training. Human raters prefer the upgraded configuration over the strongest standalone baseline in 84.3% of comparisons. These results support learning tool use across video-generation tasks and show that the trained agent can benefit from subsequent advances in generation tools. Project page: https://andyca111.github.io/VideoGen_Agent/

cs.CV↗

Bistationary Traces, Wide Levels, and Branch-Cover Rigidity for an Unrestricted Typed Variant of the Hayut-Magidor Forcing

For every uncountable regular cardinal $α$, $\mathbb S^{\ast}(α)$ is an explicitly typed four-coordinate forcing motivated by the ladder-system construction of Hayut and Magidor. The forcing is $σ$-closed and, after adjoining a formal maximum, $α$-strategically closed. For $α\geqω_2$, every nonempty countable family of designated generic branches has a stationary and costationary common trace on the generic ladder-coordinate set $L_α$, while no countable family of cofinal branches generates $L_α$. These conclusions persist under a Kurepa-style level-size bound. In the unrestricted forcing, for every infinite cardinal $μ<α$ in the ground model, some level of the generic tree contains a copy of $({}^μ2)^V$. Consequently, the endpoint-corrected restriction family indexed by $\mathcal P_{ω_2}α$ is too wide, whereas the scaled restriction system indexed by $\mathcal P_αα$ has all levels of size less than $α$ exactly when $α$ is strongly inaccessible in the ground model. When these equivalent conditions hold, the branch-covering number of $L_α$ relative to the scaled system is at least $ω_1$. The low-cofinality empty-value convention also ensures that the set of domains of $L_α$ contains no club in $\mathcal P_αα$. The unrestricted tree clause of the motivating presentation is retained, without asserting forcing equivalence.

math.LO↗

Sub-quorum colorings of graphs

A sub-quorum coloring is a partial vertex coloring in which every colored vertex sees at least half of its colored closed neighborhood in its own color. Hedetniemi, Hedetniemi, Laskar and Mulder introduced its maximum number of colors, $\psq(G)$, as an open direction in their foundational work on quorum colorings. We establish general bounds, relate $\psq$ to $2$-independence, discuss computational complexity, and determine exact values for several classical families. For rectangular grids $G_{m,n}=P_m\square P_n$, we give a new profile proof of the known dissociation-number formula, equivalent to earlier exact $3$-path vertex-cover results. The proof supplies equality and rigidity information used to establish the same formula for the auxiliary parameter when the representative matching is restricted to one direction. We also obtain a five-sixths inequality for mixed-direction matchings on even-by-even rectangles. Exact transfer certificates establish the sub-quorum coloring formula for all fixed strip widths $2\le m\le11$. For hypercubes, we prove the dimension-free identity $\psq(Q_n)=\bii(Q_n)=2^{n-1}$ for every $n\ge2$. The upper bound for the sub-quorum coloring number follows from Huang's signed adjacency matrix through a restricted energy estimate and an injective linear map. The computer-assisted grid claims use integer arithmetic and are independently reproducible by the accompanying verifier.

math.CO↗

Entropy Can Flow, or It Can Guide. Be Entropy. LEDFlow: Introducing Entropy-guided Generation Order into Uniform Discrete Flow

Uniform discrete flow permits repeated updates at every generation position. While continued revision supports correction of wrong tokens, it also exposes correct intermediate predictions to later errors. An experiment on Sudoku puzzles shows that 9.4% of generated cells are correct at an intermediate step but incorrect in the final output. We introduce generation order into uniform discrete flow through selective absorption, which fixes chosen predictions while preserving the uniform flow velocity at active positions. To prioritize reliable predictions for absorption, we propose Low-Entropy Discrete Flow (LEDFlow), a training-free sampler that adaptively orders absorption by local entropy. By decomposing absorption error into joint dependence and conditional prediction terms, we show that, under entropy-error regularity, selecting the lowest-entropy positions under a fixed absorption count minimizes an upper bound on the conditional term. We further analyze sensitivity of global lookahead, whose worst-case decision-error bound grows with lookahead window under an imperfect denoiser. Across reasoning benchmarks, LEDFlow attains 0.845 Nikoli Sudoku solve accuracy, with largest gains on strongly constrained tasks. On a text-to-image generation benchmark it attains the best overall score among decode-time samplers, and on multimodal understanding it improves over the default sampler on all six benchmarks, at an inference cost comparable to standard flow sampling.

cs.LG↗

Sex Estimation from Footwear Outsole Impressions Using CNN Transfer Learning and Interpretable Image Statistics

Footwear outsole impressions are a common form of forensic pattern evidence, yet quantitative methods for estimating wearer attributes from these images remain relatively underdeveloped. We investigate binary sex estimation from footwear outsole impressions by comparing convolutional neural network (CNN) transfer learning with traditional feature-based classification. Using a publicly available outsole-impression dataset, we adopt a shoe-level training and test partition that keeps replicate scans of the same physical shoe together to reduce data leakage. We evaluate pretrained CNNs through end-to-end fine-tuning, frozen feature extraction followed by support vector machine classification, and hybrid feature fusion incorporating handcrafted, geometric, and metadata-derived descriptors. Fine-tuned CNNs achieve the strongest overall predictive performance and substantially outperform traditional classifiers trained on the manually specified descriptors alone, while frozen-feature approaches offer a less computationally demanding alternative. Exploratory analysis of low-dimensional CNN representations reveals associations with frequency threshold ratio, image contrast, and wavelet-based summaries, providing a connection between learned representations and measurable properties of outsole impressions. These findings suggest that CNN transfer learning captures discriminative information beyond the descriptors considered and offers a promising approach to footwear-based forensic screening. Further validation on independently collected and casework-like impressions is needed before operational use.

cs.CV↗

A Systematic Study of Resonance-Driven Flux Modifications in Extreme-Mass-Ratio Inspirals

Transient orbital resonances can introduce phase-dependent corrections to the evolution of extreme-mass-ratio inspirals (EMRIs), potentially altering their long-term dynamics and emitted gravitational-wave signals. In this work, we quantify the resonance-induced modifications to the energy, axial angular momentum, and Carter constant fluxes and compute the corresponding resonance coefficients across a broad region of the orbital parameter space. Using the publicly available $\texttt{pybhpt}$ code, we solve the Teukolsky equation in the frequency domain to coherently combine the degenerate radial and polar harmonics that arise at resonance. We analytically derive a selection rule governing the relative radial-polar phase dependence of the resonant flux modifications. We argue that the relative strength of the resonant flux modifications reflects a balance between symmetry-induced cancellations and the degree to which the resonant orbit samples the underlying two-dimensional orbital phase space. For the dynamically important $3{:}2$ and $2{:}1$ resonances, we also characterize how the resonance coefficients vary with the primary black-hole spin, orbital eccentricity and inclination. Our results constitute the largest set of Teukolsky-based resonance coefficients calculated to date and provide essential input for future studies of transient orbital resonances in EMRIs.

gr-qc↗

FAST-ML: A Hybrid Physics-Machine Learning Framework for Tropical Cyclone Intensity Forecasting

Rapid intensification (RI) remains one of the most consequential and difficult aspects of tropical cyclone (TC) forecasting. Although full-physics numerical weather prediction models can represent the processes governing RI, resolving storm-environment interactions remains computationally expensive, while purely data-driven approaches often lack physical interpretability. We present FAST-ML, a hybrid framework that bridges data-driven efficiency with physical constraints. A physically informed dual-stream neural parameterization ingests 3D ERA5 fields to diagnose ventilation controls---environmental wind shear and mid-level entropy deficit. By optimizing these parameters end-to-end through a differentiable FAST intensity model, this architecture establishes a robust new paradigm for observation-driven parameter optimization, ensuring storm evolution remains strictly governed by thermodynamic principles. By better capturing the storm's continuous intensity evolution, FAST-ML improves upon its physical baseline, reducing ensemble CRPS across forecast lead times, with a reduction of approximately 31% at 60 h and nearly halving the RI false alarm ratio without sacrificing detection skill. In a 100-member ensemble configuration, FAST-ML produces intensity forecasts comparable to FNV3 for selected storms under the evaluated input configurations. Furthermore, zero-shot tests on selected Eastern Pacific storms provide encouraging evidence of cross-basin transferability. FAST-ML provides a modular intensity forecasting framework that can be coupled with externally supplied storm tracks and environmental fields. It demonstrates that observation-driven parameter learning within physically constrained dynamics simultaneously enhances accuracy, interpretability, and computational efficiency.

physics.ao-ph↗

Isolated Sign Language Recognition for Icelandic Sign Language: Experiments in a Low-resource Setting

We present the first experiments on isolated sign language recognition (ISLR) for Icelandic Sign Language (ÍTM). We use ÍTM SignWiki, a dataset derived from a bilingual Icelandic--ÍTM online dictionary. It is genuinely low-resource: 1,845 videos cover 849 classes, 86% of which have only two examples, making the full task effectively one-shot recognition across signers. We compare two open-source ISLR frameworks, OpenHands and SPOTER, on three tasks of increasing vocabulary size (22, 117 and 849 classes), and evaluate three pose estimators and two forms of cross-lingual transfer. With ÍTM data alone, SPOTER outperforms OpenHands on all three tasks, and MediaPipe poses give better results than AlphaPose or SDPose. Cross-lingual transfer brings the largest gains: pretraining SPOTER on American Sign Language data before finetuning on ÍTM raises accuracy by 14--24 percentage points, to 72.7%, 47.9% and 22.6% on the three tasks, and multilingual training with data from six other sign languages lifts OpenHands from 1.41% to 28.86% on the full task. Although far from practical use, the results suggest that transfer from better-resourced sign languages is promising for very low-resource ones. We release our adapted versions of both frameworks.

cs.CL↗

Control Barrier Functions for Safe Free-Flying Robotic Spacecraft Operations in Tumbling Target Capture

This paper presents a modular control barrier function (CBF) framework for safe free-flying robotic spacecraft operations during tumbling target capture. Motivated by latest ESA guidelines for safe close proximity operations, safety zones and requirements are translated into dedicated CBFs. The 13-DoF system is decomposed into translational, attitude, and robotic subsystems, each equipped with a safety filter that minimally modifies nominal control inputs in a lightweight quadratic program. The filters enforce a conical approach corridor, collision avoidance zone, attitude line-of-sight pointing, angular velocity limits, robotic joint limits, link-base collision avoidance, and actuator constraints. Dynamic coupling between subsystems is handled by treating upstream safe control commands as known interconnection inputs in the downstream safety filters, preserving modularity while supporting system-level safety. The framework is validated in an on-orbit servicing scenario, including final approach, angular rate synchronization, and tumbling target grasping, using the high-fidelity astrodynamics simulator Basilisk. Monte Carlo simulation results demonstrate runtime efficiency and operational safety for various tumbling rates.

cs.RO↗

KwaiMind Technical Report

Commercial image editing requires product identity preservation, accurate text rendering, and user appeal alongside general editing quality. We present KwaiMind, an image editing system combining general capabilities with e-commerce specialization. An agent-based data engine maintains approximately 1.8 million high-quality editing pairs. Built on a multimodal diffusion transformer, KwaiMind undergoes continued pre-training and supervised fine-tuning, followed by preference optimization and online reinforcement learning. A general-purpose vision-language judge and specialized rewards for click-through rate (CTR), text rendering, and product consistency guide specialized policies, which are consolidated through on-policy distillation. We introduce Ecom-Bench, covering 11 commercial editing tasks with task-specific visual evaluation and CTR-based ranking. KwaiMind achieves the strongest overall scores among evaluated open-source editors on ImgEdit, GEdit, both language splits of REDEdit, and Ecom-Bench visual quality, and the highest aggregate CTR ranking score among compared systems. Offline, CTR-guided optimization increases the proportion of generated images whose predicted CTR exceeds that of the original product image from 12.16% to 37.41%. In an online A/B experiment, CTR-based selection of product main images yields an approximately 2.44% relative increase in actual CTR. These results demonstrate the value of domain-specific data and reward-driven alignment for commercial image editing.

cs.CV↗

Double Descent and Malign Overfitting in Diffusion Models

Conventional wisdom in deep learning holds that overparameterization---having more parameters $p$ than training samples $n$---is benign: larger models generalize better and, even without regularization, interpolating models generalize well, the test error following a double-descent curve. One might expect the same benign overfitting for diffusion models, whose training reduces to regression, i.e. to minimizing a quadratic score-matching loss. Yet the opposite is observed: overfitting here is catastrophic, driving the model into a memorization regime. We resolve this paradox by combining experiments on U-Nets trained on CelebA with a random-features model for which we derive closed-form learning curves. We show that with a fixed number $m$ of noise realizations per training sample, an interpolation peak does occur, but at $p\sim nm$ rather than at $p\sim n$ as in standard regression. The rise of the test loss, however, sets in much earlier, at $p\sim n$, independently of $m$. This overfitting is malign because, although the implicit regularization of training is fully at work, it drives the model toward the empirical score, which memorizes the training set, rather than toward the true score. A bias-variance decomposition pinpoints the mechanism: the bias of the score estimator starts to grow at $p\sim n$; past the peak the variance decays, as in regression, whereas the bias keeps growing and both saturate at a large value. Since diffusion models are trained with $m\gg1$, the peak is pushed to very large model sizes, and therefore sit on the rising branch that precedes it, where malign overfitting is already in play. Nevertheless, overparameterization remains beneficial when paired with regularization: in the random-features theory and in U-Net experiments, optimally regularized large models---via a ridge penalty or early stopping, respectively---outperform any unregularized models.

cs.LG↗

Metric foundations of geometry

A metric space is called all-set-homogeneous if every isometry between two of its subsets extends to an isometry of the whole space. We classify all-set-homogeneous geodesic spaces: besides the classical examples, they include the universal metric trees of finite valence. We also prove that every complete all-set-homogeneous length space is geodesic, and hence the same classification holds in this setting.

math.MG↗

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge that returns label probabilities instead of text, and whose confidence decides whether to accept its verdict or escalate to a reasoning judge. Against sixteen generative and reward-model judges, with blinded human adjudication, JEV comes within three points of GPT-6 wherever a verdict can be read off the text, at 0.36% of its fee and a 0.15-second median latency, and falls behind where the verdict must be derived, as in math, code, and logic. Its confidence marks this boundary. With a threshold frozen in advance, accepting confident verdicts and escalating the rest is 0.9 points more accurate than GPT-6 on 1,610 held-out pairs at 41% of its fee, and in a pre-specified live test on two new workloads the cascade matches GPT-6's accuracy exactly. Confidence routing weakens on style-adversarial pairs and reference-free prose; we close with a simple recipe for validating thresholds locally.

cs.AI↗