Searcharxiv⌕ Search

SEARCH · Searcharxiv

Search Searcharxiv

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,477 records · Page 82Linked to original sources

VideoGen-Agent: Reinforcing Video Generation Agents

Recent advances in video generative models have enabled high-fidelity, temporally coherent video generation. However, these models often struggle to satisfy prompts requiring specialized knowledge, specific identities, physical consistency, or ordered events. In this paper, we present VideoGen-Agent, a multimodal agent trained through multitask agentic reinforcement learning to use external tools for video generation. The agent coordinates augmentation, generation, and verification tools through multi-turn interactions, using the prompt and intermediate observations to guide its decisions. We train a shared policy on a category-balanced dataset spanning six tasks. Supervised fine-tuning on teacher-generated trajectories establishes tool-use behavior, which is then refined through reinforcement learning. A category-aware hybrid reward evaluates tool-call validity, task-appropriate tool use, and generated video quality. We further introduce VABench, a held-out benchmark of 600 prompts covering procedural knowledge, single- and multi-entity identity preservation, physical consistency, scene composition, and multi-shot temporal structure. On VABench, VideoGen-Agent improves over its base text-to-video generator by 19.1 points, from 56.5 to 75.6. Upgrading the generation tools further raises the score to 86.1 without additional agent training. Human raters prefer the upgraded configuration over the strongest standalone baseline in 84.3% of comparisons. These results support learning tool use across video-generation tasks and show that the trained agent can benefit from subsequent advances in generation tools. Project page: https://andyca111.github.io/VideoGen_Agent/

cs.CV↗

Bistationary Traces, Wide Levels, and Branch-Cover Rigidity for an Unrestricted Typed Variant of the Hayut-Magidor Forcing

For every uncountable regular cardinal $α$, $\mathbb S^{\ast}(α)$ is an explicitly typed four-coordinate forcing motivated by the ladder-system construction of Hayut and Magidor. The forcing is $σ$-closed and, after adjoining a formal maximum, $α$-strategically closed. For $α\geqω_2$, every nonempty countable family of designated generic branches has a stationary and costationary common trace on the generic ladder-coordinate set $L_α$, while no countable family of cofinal branches generates $L_α$. These conclusions persist under a Kurepa-style level-size bound. In the unrestricted forcing, for every infinite cardinal $μ<α$ in the ground model, some level of the generic tree contains a copy of $({}^μ2)^V$. Consequently, the endpoint-corrected restriction family indexed by $\mathcal P_{ω_2}α$ is too wide, whereas the scaled restriction system indexed by $\mathcal P_αα$ has all levels of size less than $α$ exactly when $α$ is strongly inaccessible in the ground model. When these equivalent conditions hold, the branch-covering number of $L_α$ relative to the scaled system is at least $ω_1$. The low-cofinality empty-value convention also ensures that the set of domains of $L_α$ contains no club in $\mathcal P_αα$. The unrestricted tree clause of the motivating presentation is retained, without asserting forcing equivalence.

math.LO↗

Sub-quorum colorings of graphs

A sub-quorum coloring is a partial vertex coloring in which every colored vertex sees at least half of its colored closed neighborhood in its own color. Hedetniemi, Hedetniemi, Laskar and Mulder introduced its maximum number of colors, $\psq(G)$, as an open direction in their foundational work on quorum colorings. We establish general bounds, relate $\psq$ to $2$-independence, discuss computational complexity, and determine exact values for several classical families. For rectangular grids $G_{m,n}=P_m\square P_n$, we give a new profile proof of the known dissociation-number formula, equivalent to earlier exact $3$-path vertex-cover results. The proof supplies equality and rigidity information used to establish the same formula for the auxiliary parameter when the representative matching is restricted to one direction. We also obtain a five-sixths inequality for mixed-direction matchings on even-by-even rectangles. Exact transfer certificates establish the sub-quorum coloring formula for all fixed strip widths $2\le m\le11$. For hypercubes, we prove the dimension-free identity $\psq(Q_n)=\bii(Q_n)=2^{n-1}$ for every $n\ge2$. The upper bound for the sub-quorum coloring number follows from Huang's signed adjacency matrix through a restricted energy estimate and an injective linear map. The computer-assisted grid claims use integer arithmetic and are independently reproducible by the accompanying verifier.

math.CO↗

Entropy Can Flow, or It Can Guide. Be Entropy. LEDFlow: Introducing Entropy-guided Generation Order into Uniform Discrete Flow

Uniform discrete flow permits repeated updates at every generation position. While continued revision supports correction of wrong tokens, it also exposes correct intermediate predictions to later errors. An experiment on Sudoku puzzles shows that 9.4% of generated cells are correct at an intermediate step but incorrect in the final output. We introduce generation order into uniform discrete flow through selective absorption, which fixes chosen predictions while preserving the uniform flow velocity at active positions. To prioritize reliable predictions for absorption, we propose Low-Entropy Discrete Flow (LEDFlow), a training-free sampler that adaptively orders absorption by local entropy. By decomposing absorption error into joint dependence and conditional prediction terms, we show that, under entropy-error regularity, selecting the lowest-entropy positions under a fixed absorption count minimizes an upper bound on the conditional term. We further analyze sensitivity of global lookahead, whose worst-case decision-error bound grows with lookahead window under an imperfect denoiser. Across reasoning benchmarks, LEDFlow attains 0.845 Nikoli Sudoku solve accuracy, with largest gains on strongly constrained tasks. On a text-to-image generation benchmark it attains the best overall score among decode-time samplers, and on multimodal understanding it improves over the default sampler on all six benchmarks, at an inference cost comparable to standard flow sampling.

cs.LG↗

X-Planner: Event-Structured Task Planning for Embodied Intelligence

Task planning bridges high-level instructions and executable behavior in long-horizon manipulation, yet modern Vision-Language-Action (VLA) systems often leave this intermediate structure implicit. Existing chain-of-thought (CoT) planners also tend to rely on coarse task-level annotations or serialize long reasoning traces token by token. We present X-Planner, a planning front-end that addresses both the supervision and representation of embodied reasoning. Our planning data combine Ego, UMI, and teleoperation under a hierarchy granularity with source-dependent annotation depth. Takeover-time annotations and human-designed failures supervise ongoing error recognition. On the model side, a shared VLM backbone exposes two event-structured plan forms: a discrete interface that emits interpretable event states and a latent interface that relays continuous CoT states across staggered Transformer depths through Staircase Decoding. A frozen latent-to-text reconstruction objective provides a semantic anchor for the latent representation. Offline two-step planning evaluation places X-Planner second among four evaluated models on both BERTScore-F1 and a judge-based Overall score. In real-robot experiments, respectively, outperforming the evaluated baselines. These results characterize planning-text quality and downstream execution.

cs.AI↗

Sex Estimation from Footwear Outsole Impressions Using CNN Transfer Learning and Interpretable Image Statistics

Footwear outsole impressions are a common form of forensic pattern evidence, yet quantitative methods for estimating wearer attributes from these images remain relatively underdeveloped. We investigate binary sex estimation from footwear outsole impressions by comparing convolutional neural network (CNN) transfer learning with traditional feature-based classification. Using a publicly available outsole-impression dataset, we adopt a shoe-level training and test partition that keeps replicate scans of the same physical shoe together to reduce data leakage. We evaluate pretrained CNNs through end-to-end fine-tuning, frozen feature extraction followed by support vector machine classification, and hybrid feature fusion incorporating handcrafted, geometric, and metadata-derived descriptors. Fine-tuned CNNs achieve the strongest overall predictive performance and substantially outperform traditional classifiers trained on the manually specified descriptors alone, while frozen-feature approaches offer a less computationally demanding alternative. Exploratory analysis of low-dimensional CNN representations reveals associations with frequency threshold ratio, image contrast, and wavelet-based summaries, providing a connection between learned representations and measurable properties of outsole impressions. These findings suggest that CNN transfer learning captures discriminative information beyond the descriptors considered and offers a promising approach to footwear-based forensic screening. Further validation on independently collected and casework-like impressions is needed before operational use.

cs.CV↗

A Systematic Study of Resonance-Driven Flux Modifications in Extreme-Mass-Ratio Inspirals

Transient orbital resonances can introduce phase-dependent corrections to the evolution of extreme-mass-ratio inspirals (EMRIs), potentially altering their long-term dynamics and emitted gravitational-wave signals. In this work, we quantify the resonance-induced modifications to the energy, axial angular momentum, and Carter constant fluxes and compute the corresponding resonance coefficients across a broad region of the orbital parameter space. Using the publicly available $\texttt{pybhpt}$ code, we solve the Teukolsky equation in the frequency domain to coherently combine the degenerate radial and polar harmonics that arise at resonance. We analytically derive a selection rule governing the relative radial-polar phase dependence of the resonant flux modifications. We argue that the relative strength of the resonant flux modifications reflects a balance between symmetry-induced cancellations and the degree to which the resonant orbit samples the underlying two-dimensional orbital phase space. For the dynamically important $3{:}2$ and $2{:}1$ resonances, we also characterize how the resonance coefficients vary with the primary black-hole spin, orbital eccentricity and inclination. Our results constitute the largest set of Teukolsky-based resonance coefficients calculated to date and provide essential input for future studies of transient orbital resonances in EMRIs.

gr-qc↗

FAST-ML: A Hybrid Physics-Machine Learning Framework for Tropical Cyclone Intensity Forecasting

Rapid intensification (RI) remains one of the most consequential and difficult aspects of tropical cyclone (TC) forecasting. Although full-physics numerical weather prediction models can represent the processes governing RI, resolving storm-environment interactions remains computationally expensive, while purely data-driven approaches often lack physical interpretability. We present FAST-ML, a hybrid framework that bridges data-driven efficiency with physical constraints. A physically informed dual-stream neural parameterization ingests 3D ERA5 fields to diagnose ventilation controls---environmental wind shear and mid-level entropy deficit. By optimizing these parameters end-to-end through a differentiable FAST intensity model, this architecture establishes a robust new paradigm for observation-driven parameter optimization, ensuring storm evolution remains strictly governed by thermodynamic principles. By better capturing the storm's continuous intensity evolution, FAST-ML improves upon its physical baseline, reducing ensemble CRPS across forecast lead times, with a reduction of approximately 31% at 60 h and nearly halving the RI false alarm ratio without sacrificing detection skill. In a 100-member ensemble configuration, FAST-ML produces intensity forecasts comparable to FNV3 for selected storms under the evaluated input configurations. Furthermore, zero-shot tests on selected Eastern Pacific storms provide encouraging evidence of cross-basin transferability. FAST-ML provides a modular intensity forecasting framework that can be coupled with externally supplied storm tracks and environmental fields. It demonstrates that observation-driven parameter learning within physically constrained dynamics simultaneously enhances accuracy, interpretability, and computational efficiency.

physics.ao-ph↗

Isolated Sign Language Recognition for Icelandic Sign Language: Experiments in a Low-resource Setting

We present the first experiments on isolated sign language recognition (ISLR) for Icelandic Sign Language (ÍTM). We use ÍTM SignWiki, a dataset derived from a bilingual Icelandic--ÍTM online dictionary. It is genuinely low-resource: 1,845 videos cover 849 classes, 86% of which have only two examples, making the full task effectively one-shot recognition across signers. We compare two open-source ISLR frameworks, OpenHands and SPOTER, on three tasks of increasing vocabulary size (22, 117 and 849 classes), and evaluate three pose estimators and two forms of cross-lingual transfer. With ÍTM data alone, SPOTER outperforms OpenHands on all three tasks, and MediaPipe poses give better results than AlphaPose or SDPose. Cross-lingual transfer brings the largest gains: pretraining SPOTER on American Sign Language data before finetuning on ÍTM raises accuracy by 14--24 percentage points, to 72.7%, 47.9% and 22.6% on the three tasks, and multilingual training with data from six other sign languages lifts OpenHands from 1.41% to 28.86% on the full task. Although far from practical use, the results suggest that transfer from better-resourced sign languages is promising for very low-resource ones. We release our adapted versions of both frameworks.

cs.CL↗

Control Barrier Functions for Safe Free-Flying Robotic Spacecraft Operations in Tumbling Target Capture

This paper presents a modular control barrier function (CBF) framework for safe free-flying robotic spacecraft operations during tumbling target capture. Motivated by latest ESA guidelines for safe close proximity operations, safety zones and requirements are translated into dedicated CBFs. The 13-DoF system is decomposed into translational, attitude, and robotic subsystems, each equipped with a safety filter that minimally modifies nominal control inputs in a lightweight quadratic program. The filters enforce a conical approach corridor, collision avoidance zone, attitude line-of-sight pointing, angular velocity limits, robotic joint limits, link-base collision avoidance, and actuator constraints. Dynamic coupling between subsystems is handled by treating upstream safe control commands as known interconnection inputs in the downstream safety filters, preserving modularity while supporting system-level safety. The framework is validated in an on-orbit servicing scenario, including final approach, angular rate synchronization, and tumbling target grasping, using the high-fidelity astrodynamics simulator Basilisk. Monte Carlo simulation results demonstrate runtime efficiency and operational safety for various tumbling rates.

cs.RO↗

A Hybrid Classical-Learning Framework for Adaptive Decision Directed Speech Enhancement

Speech enhancement aims to recover clean speech signals from noisy observations while preserving speech quality and intelligibility. Classical methods such as Spectral Subtraction and Decision-Directed (DD) enhancement remain widely used because of their interpretability and low computational complexity, but they may suffer from musical-noise artifacts or excessive attenuation of weak speech components under low signal-to-noise ratio (SNR) conditions. This paper proposes an Adaptive Beta-Constrained Decision-Directed (ABCDD) speech enhancement framework that extends the conventional DD method through a frame-dependent lower gain bound. The introduced beta parameter controls the tradeoff between noise suppression and speech preservation. To automate parameter selection for large and diverse datasets, a lightweight multilayer perceptron (MLP) model is further developed to predict frame-level beta values directly from noisy-speech features. The proposed framework is evaluated using both a representative speech example and large-scale testing on the VoiceBank-DEMAND dataset. In the representative example, ABCDD outperformed conventional Spectral Subtraction and classical DD across multiple objective metrics, including SNR, Log-Spectral Distance (LSD), Root-Mean-Square Error (RMSE), correlation, and Scale-Invariant Signal-to-Distortion Ratio (SI-SDR). On 100 unseen VoiceBank-DEMAND test files, the proposed MLP-beta ABCDD method improved average scale-aligned SNR from 9.41 dB to 13.82 dB, corresponding to an average gain of 4.41 dB. The results indicate that combining interpretable classical enhancement structure with lightweight machine-learning-based parameter adaptation provides an effective and practical direction for robust speech enhancement.

eess.AS↗

KwaiMind Technical Report

Commercial image editing requires product identity preservation, accurate text rendering, and user appeal alongside general editing quality. We present KwaiMind, an image editing system combining general capabilities with e-commerce specialization. An agent-based data engine maintains approximately 1.8 million high-quality editing pairs. Built on a multimodal diffusion transformer, KwaiMind undergoes continued pre-training and supervised fine-tuning, followed by preference optimization and online reinforcement learning. A general-purpose vision-language judge and specialized rewards for click-through rate (CTR), text rendering, and product consistency guide specialized policies, which are consolidated through on-policy distillation. We introduce Ecom-Bench, covering 11 commercial editing tasks with task-specific visual evaluation and CTR-based ranking. KwaiMind achieves the strongest overall scores among evaluated open-source editors on ImgEdit, GEdit, both language splits of REDEdit, and Ecom-Bench visual quality, and the highest aggregate CTR ranking score among compared systems. Offline, CTR-guided optimization increases the proportion of generated images whose predicted CTR exceeds that of the original product image from 12.16% to 37.41%. In an online A/B experiment, CTR-based selection of product main images yields an approximately 2.44% relative increase in actual CTR. These results demonstrate the value of domain-specific data and reward-driven alignment for commercial image editing.

cs.CV↗

Double Descent and Malign Overfitting in Diffusion Models

Conventional wisdom in deep learning holds that overparameterization---having more parameters $p$ than training samples $n$---is benign: larger models generalize better and, even without regularization, interpolating models generalize well, the test error following a double-descent curve. One might expect the same benign overfitting for diffusion models, whose training reduces to regression, i.e. to minimizing a quadratic score-matching loss. Yet the opposite is observed: overfitting here is catastrophic, driving the model into a memorization regime. We resolve this paradox by combining experiments on U-Nets trained on CelebA with a random-features model for which we derive closed-form learning curves. We show that with a fixed number $m$ of noise realizations per training sample, an interpolation peak does occur, but at $p\sim nm$ rather than at $p\sim n$ as in standard regression. The rise of the test loss, however, sets in much earlier, at $p\sim n$, independently of $m$. This overfitting is malign because, although the implicit regularization of training is fully at work, it drives the model toward the empirical score, which memorizes the training set, rather than toward the true score. A bias-variance decomposition pinpoints the mechanism: the bias of the score estimator starts to grow at $p\sim n$; past the peak the variance decays, as in regression, whereas the bias keeps growing and both saturate at a large value. Since diffusion models are trained with $m\gg1$, the peak is pushed to very large model sizes, and therefore sit on the rising branch that precedes it, where malign overfitting is already in play. Nevertheless, overparameterization remains beneficial when paired with regularization: in the random-features theory and in U-Net experiments, optimally regularized large models---via a ridge penalty or early stopping, respectively---outperform any unregularized models.

cs.LG↗

Metric foundations of geometry

A metric space is called all-set-homogeneous if every isometry between two of its subsets extends to an isometry of the whole space. We classify all-set-homogeneous geodesic spaces: besides the classical examples, they include the universal metric trees of finite valence. We also prove that every complete all-set-homogeneous length space is geodesic, and hence the same classification holds in this setting.

math.MG↗

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

LLM-as-a-judge scales evaluation, but reasoning judges are slow and costly. We study JEV-as-a-Judge: evaluation with JEV, a decision-only judge that returns label probabilities instead of text, and whose confidence decides whether to accept its verdict or escalate to a reasoning judge. Against sixteen generative and reward-model judges, with blinded human adjudication, JEV comes within three points of GPT-6 wherever a verdict can be read off the text, at 0.36% of its fee and a 0.15-second median latency, and falls behind where the verdict must be derived, as in math, code, and logic. Its confidence marks this boundary. With a threshold frozen in advance, accepting confident verdicts and escalating the rest is 0.9 points more accurate than GPT-6 on 1,610 held-out pairs at 41% of its fee, and in a pre-specified live test on two new workloads the cascade matches GPT-6's accuracy exactly. Confidence routing weakens on style-adversarial pairs and reference-free prose; we close with a simple recipe for validating thresholds locally.

cs.AI↗

When Post-Processing Fairness Constraints Help and When They Harm: Evidence from Eight Cross-Domain Evaluations

Fairness audits in production ML typically occur once, at deployment, on a single domain. Both fail in practice: fairness can shift after retraining or a changing user base, and interventions validated on one dataset are rarely tested across the heterogeneous domains an organization deploys. We present FAPE (Fairness Auditing for Production Environments), a four-stage framework evaluating a single post-processing intervention, Fairlearn's ThresholdOptimizer, across eight domain evaluations: criminal justice, income prediction, legal admissions, credit lending, agricultural lending, a multi-domain benchmark corpus, healthcare, and education. Each is scored on demographic parity and equalized odds difference, plus disparate impact ratio and accuracy cost where computable. Intervention effectiveness tracks baseline disparity magnitude: across model-domain pairs the constraint improved disparity in 9 of 14 high-disparity cases and worsened it in 3 of 4 near-fair ones. Each of the five high-disparity exceptions reverses under one of two measurement checks, a minimum group size or thresholds fit on held-out data. A CUSUM monitor started at deployment, tested on a simulated shift, separates constrained models that never met a 0.1 parity convention from those that met it and later regressed. A single deployment-time audit is therefore an unreliable guide, which argues for baseline-disparity screening and continuous monitoring

cs.LG↗

First-Principles Nonadiabatic Dynamics via the Multi-Orbital Anderson-Newns Model

We develop a first-principles theory for nonadiabatic surface dynamics, providing a fit-free connection between density functional theory (DFT) calculations and the effective multi-orbital Anderson-Newns (AN) theory. Our theory contains two main advances. First, we outline the multi-orbital AN theory with orbital overlap and derive closed-form expressions for the hybridization energy, electronic dynamics, and electronic friction. Second, we describe a procedure to map the DFT Hamiltonian into the effective AN Hamiltonian with nuclear-position dependence. We obtain the electronic part of the AN Hamiltonian solely from the adsorbate-projected density-of-states matrix. We then define the bare nuclear potential as the difference between the total energy and the hybridization energy. The theory is applied to H and CO on the Cu surface, where widely used assumptions about the hybridization function, including the wide-band limit, semi-elliptical forms, and separability in energy and nuclear coordinate, are found to fail, and the single-orbital description breaks down qualitatively for CO. We expect this work to be broadly useful for first-principles modeling of coupled nuclear-electronic dynamics and chemical reactions at metallic surfaces.

physics.chem-ph↗

A Szemerédi-Trotter Theorem in Arbitrary Fields

Let $k$ be a field of characteristic $p\ge0$. We prove that $m$ points and $n$ lines in $k^2$ determine $O((mn)^{2/3}+m+n+mn/p)$ incidences, the last term being omitted in characteristic zero. The proof uses the polynomial method, and for $m=n$ the bound is sharp over prime fields. As applications, over prime fields in which $-1$ is not a square we obtain the $L^2\to L^r$ extension estimate for the paraboloid in $\F_p^3$ for $r>10/3$. Over every odd prime field, we show that a two-source extractor construction of Bourgain has exponentially small error at every min-entropy rate greater than $1/3$. We also improve sum-product estimates for small sets in positive characteristic and obtain projection and Furstenberg estimates over prime fields.

math.CO↗