Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,135 records · Page 63Linked to original sources

On the nonnegativity of monomial immanants for hook partitions

Let $A=(a_{ij})$ be an $n\times n$ real matrix and let $λ$ be a partition of $n$. Let $ϕ^λ$ be the class function dual to the Young permutation character, and let $$ ϕ^λ[A] = \sum_{σ\in\mathfrak{S}_n}ϕ^λ(σ) \prod_{i=1}^{n}a_{iσ(i)} $$ be the corresponding monomial immanant. Stembridge [Canad. J. Math. 44 (1992), pp. 1079-1099] posed the following open problem: If all minors of $A$ of order at most $r$ are nonnegative, and the partition $λ$ has length at most $r$, is it true that $ϕ^λ[A]\ge 0$? This paper solves the problem for all hook partitions.

math.CO↗

A Robust Hertz-Linewidth Quantum-Dot Coherent Swept Source with Adaptive Self-Linearization

Frequency-modulated continuous-wave (FMCW) techniques underpin both coherent optical ranging and microwave radar, creating a common demand for highly coherent and highly linear frequency-chirped sources across the optical and microwave domains. Conventional semiconductor lasers, however, are fundamentally constrained by trade-offs among linewidth, linear tunability, and operational robustness. Here, we present a robust self-linearized quantum-dot (QD) coherent swept-source architecture through the co-design of source physics and system control. The chaos-free characteristics of QD lasers enable stable low-quality-factor external-cavity locking, yielding a Lorentzian linewidth of 12.6 Hz. Unlike self-injection-locked lasers, broadband external optical feedback allows the laser to maintain high optical coherence and stable operation over a wide current-tuning range, thereby enabling robust turnkey operation together with a chirp bandwidth of 23.2 GHz. Furthermore, a statistically gated real-time iterative learning control (ILC) strategy reduces the chirp nonlinearity $(1-R^2)$ to as low as $8.96\times10^{-8}$, while maintaining excellent environmental stability under laser-temperature variations. To demonstrate practicality, we demonstrate isolator-free coherent LiDAR and photonic generation of frequency-agile, linearly chirped microwave waveforms, establishing a common FMCW source platform for optical ranging and radar waveform synthesis. To the best of our knowledge, we demonstrate for the first time a QD-based coherent swept source that simultaneously combines hertz-level linewidth and ultrahigh chirp linearity, which we envision as a unified source platform for FMCW signal generation and coherent sensing across the optical and microwave domains.

physics.optics↗

When One Leak Pays Forever: Context Binding and the Price of Deterring Collusion

A coalition that deviates once can profit many times when what it sells keeps working. In a threshold-encrypted mempool, a leading defense against maximal extractable value (MEV), a quorum of the decryption committee that sells its decryption capability to a front-runner exposes every later block that the capability still decrypts. We ask how large a penalty, such as slashable stake, deters this kind of collusion. In our repeated game, a single leak by any coalition in a monotone family of authorized coalitions (for example, any $k$ of the $n$ committee members) unlocks a set of future rounds, costs a one-time penalty, and ends the coalition's participation. We show that every dynamic deviation reduces to choosing a leak time, so deterrence holds if and only if each coalition's penalty covers the largest discounted value that a single leak reaches. Without discounting, over $T$ rounds of unit value full reuse needs a penalty of $T$ while binding each leak to its own round needs $1$, so no penalty that is constant in the horizon deters unbounded reuse; a reuse window of $w$ rounds costs at most $w$ times the largest per-round value. The cheapest profile of per-party stakes that deters every coalition solves a covering linear program. For blockchain design, per-epoch keys cut the required stake from the value of a key's lifetime to the value of one epoch; we calibrate the gap on Ethereum front-running data and place Ferveo and Shutter in the model. The analysis extends to sealed-bid auctions, multi-authority voting, and federated learning under a shared key.

cs.GT↗

Stochastic Heavy Ball with Polyak Step Size and Armijo Line Search: A General Convergence Analysis

Polyak step size (PS) and Armijo line search (ALS) have received increasing attention in stochastic optimization, with encouraging empirical performance and theoretical guarantees. However, their convergence theory for stochastic heavy ball (SHB) methods remains limited. In this work, we develop a unified convergence analysis for SHB equipped with PS and ALS. To this end, we introduce a modified Armijo rule that closely parallels the Polyak step size, together with a decoupling analysis that isolates the historical dependence induced by momentum. For SHB with standard PS and ALS, we establish expected convergence for strongly convex, convex, and non-convex objectives without interpolation or restrictive conditions on the momentum parameter. Under interpolation or strong growth, we further strengthen the results to almost sure rates and last-iterate convergence. Moreover, for general settings beyond interpolation, we prove almost sure convergence to the exact optimum or to stationarity for SHB with diminishing variants of PS and ALS. These results provide a more comprehensive theoretical view of Polyak step size and Armijo line search for stochastic heavy ball methods.

cs.LG↗

Leading-edge noise reduction by perforated and porous inserts: a compressible finite-chord prediction without fitted constants

Porous and perforated leading edges reduce turbulence-interaction noise, but predictions of the reduction have relied on material impedances adjusted to the acoustic data. This paper removes that adjustment. The finite-chord scattering problem is solved in compressible subsonic flow for an arbitrary chordwise admittance, with the shed wake, the Kutta condition and permeability-dependent edge singularities, and closed to the far field with the finite span and the array aperture represented. The admittance of a perforated insert is derived from the Rayleigh conductivity of one aperture with three computed corrections: plate thickness, interaction between neighbouring apertures, and Howe's grazing-flow blockage evaluated at the aperture-averaged velocity of the boundary layer. A bulk porous insert is represented by its measured permeability and pore-fluid inertia. For three perforates at four flow conditions the prediction agrees with measurements to 2.5 dB rms over 84 comparisons, against 4.4 dB with the unscaled impedance used previously. Although the Mach number is 0.05, the chord is acoustically non-compact, and an incompressible solution rises where the measurement falls. Two of three bulk porous inserts are predicted to within 1.7 dB rms where interaction noise dominates. Narrow, equally spaced peaks measured on perforated plates are reproduced by none of the admittances tested; their spacing implies a source convecting at 0.62 of the free-stream speed, absent from a frozen-gust model. A parameter study shows that the perforate geometry acts almost entirely through one low-frequency admittance group, in which the hole radius cancels for holes buried in the laminar boundary layer.

physics.flu-dyn↗

FineSID: Scalable and Efficient Semantic Identifier Learning for Generative Recommendation

A critical prerequisite of generative recommendation is designing semantic identifiers (SIDs) that are both scalable to large item sets and efficiently learnable. Existing SID learning methods fundamentally rely on Top-1 hard assignment during vector quantization. While heuristic strategies -- such as clustering-based initialization or forced post-hoc collision resolution -- can artificially inflate codebook coverage, they often disrupt end-to-end semantic alignment and fail to address the underlying optimization bottleneck: sparse gradient propagation. In standard Top-1 assignment, gradients concentrate on a narrow subset of frequently selected codewords, leaving the majority inherently under-trained and causing severe SID collisions. To overcome this limitation natively without relying on complex initialization priors, we propose FineSID, a unified quantization framework that moves beyond Top-1 assignment by enabling fine-grained gradient propagation across the entire codebook. Instead of updating only a single selected codeword, FineSID distributes learning signals to all codewords in a soft, differentiable manner. This design promotes globally balanced codebook optimization while strictly preserving semantic consistency, effectively alleviating SID collisions and stabilizing training in large, high-dimensional codebooks. Extensive experiments on multiple public benchmarks demonstrate that FineSID is robust to initialization configurations and consistently improves both codebook utilization and recommendation accuracy. Our work provides a principled, initialization-agnostic solution for semantic identifier learning, advancing the practicality of generative recommendation.

cs.AI↗

FairDiff: Mitigating the Self-Reinforcing Matthew Effect in Diffusion Recommender Models

While the "Matthew Effect" and filter bubbles are widely recognized outcome-level biases in recommender systems, we reveal that Diffusion Recommender Models (DRMs) uniquely compound this issue through their generative dynamics. Rather than merely inheriting data imbalances, DRMs trigger a self-reinforcing amplification of popularity bias. We identify that this phenomenon is driven by two compounding mechanisms. First, while optimization loss is universally dominated by high-frequency items across recommenders, DRMs suffer from a unique structural prior mismatch during generation. Because the forward terminal distribution of long-tailed data deviates significantly from the standard Gaussian prior, reverse sampling trajectories inherently collapse toward high-density popular items, fundamentally suppressing niche item generation. To dismantle this self-reinforcing loop, we propose FairDiff, a plug-and-play fairness-aware diffusion framework. To overcome the popularity-dominated loss, we introduce Popularity Condition Guidance (PCG). Rather than altering the training objective, PCG acts as an inference-time distributional reweighting mechanism, mathematically reshaping the score-based gradient field to penalize high-popularity regions and guide trajectories toward niche semantics. Furthermore, we design a Semantic Calibration (SC) Module to bridge the prior mismatch, aligning the forward and reverse distributions via one-step optimal transport. Comprehensive evaluations demonstrate that FairDiff achieves state-of-the-art performance while effectively mitigating the self-reinforcing Matthew Effect, highlighting its value as a general framework for DRMs.

cs.AI↗

Human-inspired, Task-Dimension-Guided Exploration for Efficient Learning in High Dimensions

Efficient exploration in high-dimensional decision spaces remains a central challenge for decision-making systems. Humans, in contrast, can navigate large decision spaces with remarkable efficiency. Recent behavioral studies suggest that humans reduce dimensionality in large decision spaces by probing candidate feature dimensions, identifying reward-relevant ones, and restricting the effective decision space. Inspired by this mechanism, we propose TDGE (Task-Dimension-Guided Exploration), a human-inspired, model-agnostic algorithm with an automatically constructed task-dimension--feature--item hierarchy. TDGE follows a top-down exploration strategy: it first selects task-relevant feature dimensions, then identifies informative features within those dimensions, and finally recommends concrete items based on the selected features. Experiments on MovieLens-20M, Last.fm, and Amazon recommendation datasets show that TDGE substantially improves exploration efficiency and cold-start adaptation over baseline algorithms. Comparisons with other structured algorithms and ablation studies attribute these gains to TDGE's hierarchical structure and semantic feature-space exploration, with robust results across clustering methods and hierarchy depths. Recommendation-trajectory visualizations also show exploration patterns similar to human dimension-guided behavior.

cs.LG↗

A Bayesian Inference Framework for Binary Lens Events Fully Incorporating Higher-Order Effects

Higher-order effects, such as microlensing parallax and lens orbital motion, are essential for characterizing the physical properties of microlensing planetary systems including their masses and distances. In practice, higher-order effects are often included in light-curve modeling only when their signals are strong, because their inclusion in weakly constrained cases tends to drive the inferred parameters toward regions of parameter space that are disfavored by standard Galactic models. This binary choice over physically continuous effects introduces an event-dependent selection function that complicates population-level interpretation and prevents strong and weak detections from being combined within a single uniform framework. In this work, we investigate the origin of instabilities associated with higher-order effects, and show that such instabilities can arise from a mismatch between commonly adopted uninformative priors in the light-curve parameter space and the distributions predicted by Galactic models. Motivated by this, we develop a new Bayesian framework that enables higher-order effects to be incorporated in a stable and uniform manner. As part of this framework, we introduce Galactic Prior Modeling Engine (gapmoe), a dedicated tool that enables efficient evaluation of the Galactic prior density and allows it to be incorporated directly into the light-curve inference. Using simulated events, we demonstrate that our framework robustly recovers the true physical parameters while fully accounting for microlensing parallax and lens orbital motion. This framework eliminates the need for ad-hoc model selection and provides a scalable and statistically consistent pathway for analyzing the large samples expected from next-generation surveys such as the Roman Space Telescope.

astro-ph.EP↗

Weak and strong Lefschetz properties for vertex cover Artinian algebras associated to graphs

Let $G$ be a finite simple graph and let $A_c(G)$ be the Artinian algebra associated with its cover ideal. We prove that $A_c(G)$ has the WLP when $τ(G)>|V(G)|/2$, where $τ(G)$ denotes the size of a minimum vertex cover of $G$. As a consequence, $A_c(G)$ has the WLP with high probability when the Erdős-Rényi random graph model is considered. Moreover, we study the borderline case $τ(G)=|V(G)|/2$ and as a result, classify the WLP for paths, cycles, Ferrers graphs, and well-covered trees.

math.AC↗

Gödel Forest: Balancing Search Depth and Breadth for Data-Centric Recursive Self-Improvement

Recursive self-improvement (RSI) aims to achieve compounding gains by having models improve themselves. While most existing RSI systems optimize external agent harnesses or prompts around a frozen base model, data-centric RSI directly updates the model's own parameters by training on agent-generated data. However, because validating data strategies requires expensive model training, existing methods face a fundamental dilemma: a single agent gets trapped in narrow directions and lacks exploration breadth, while naive parallel search or heavy trace sharing sacrifices long-horizon search depth. To address this challenge, we introduce G"odel Forest, a multi-agent framework that organizes recursive self-improvement as an ensemble of co-evolving search trees. In G"odel Forest, each agent autonomously grows a persistent tree, deepening, branching, or pruning data strategies based on model feedback to secure depth, while parallel trees explore distinct regions of the data space to expand breadth. Crucially, rather than leaving trees isolated or flooding them with heavy execution logs, a dynamically co-evolving memory connects the forest: agents continuously distill their successes and failures into compact procedural lessons anchored to a global leaderboard. Through this forest ecosystem, a dead-end in one tree instantly warns the whole forest against unpromising paths, while an empirical breakthrough quickly seeds new exploration branches in neighboring trees. Evaluated on RSIBench-Data across six diverse domains, G"odel Forest outperforms the single-agent baseline by an average of 10.70% while reducing wall-clock time on five tasks. Ablations confirm that co-evolving shared memory yields a +7.00% gain over independent parallel search, demonstrating that collective distillation is key to scalable self-improvement. The code is available at https://github.com/evolvent-ai/Godel-Forest.

cs.CL↗

Kinematic Nonlinear Spatio-Temporal Trajectory Warping for Contact-Rich Dexterous Manipulation Demonstrations

We present a straightforward but effective method for repurposing existing contact-rich dexterous manipulation demonstrations. Starting from inputs of hand and object trajectories, our method outputs high-quality nonlinear trajectory warps that account for intermediate waypoints, environmental barriers, temporal shifts, and varied start/end configurations. Foundational to our method is the utilization of contact distributions, which we show allows us to reliably compute complex and high-dimensional dexterous hand trajectories following a simple object-centric warp specification pipeline. We evaluate our method across 12 variations sourced from 4 demonstrations in a publicly available dataset of human hand motion data, perform baseline comparisons, and demonstrate generalization of our approach to different manipulators. Results and code will be made available on publication.

cs.RO↗

ReWorld-Track: A Recursive Event World Model for Language-Guided Multi-Camera Tracking

Language-guided multi-camera tracking must preserve a target identity across unobserved gaps, where similar candidates and uncertain returns can make early associations unreliable. A wrong match can corrupt the history used to predict later observations and propagate identity errors across subsequent camera handoffs. We propose ReWorld-Track, a recursive event world model that carries association uncertainty into future predictions. Candidate matches and continued waiting define alternative target states, whose posterior probabilities are used to update a persistent recurrent belief. This representation preserves uncertainty about alternative trajectories through successive observations. This belief predicts the next camera, arrival time, and entry region, while appearance and language evidence guide association. By training across successive handoffs, the model learns to retain uncertainty that remains useful for later predictions and identity decisions. ReWorld-Track achieves HOTA scores of 65.19 on CityFlowV2 and 45.36 on MTMMC, with improved identity continuity across repeated handoffs. On MTMMC, its structured posterior update gains 0.50 HOTA points over a similarly sized generic updater and 0.94 points over fixed-moment soft association, raising next-camera accuracy from 86.03% to 87.41% and reducing median arrival-time error from 0.78 s to 0.71 s for subsequent target returns.

cs.CV↗

Understanding Private Evolution as Learning-Augmented Clustering

Private Evolution (PE) is a differentially private algorithm for synthetic data generation. While it can be viewed as a Wasserstein learning algorithm, it performs much better in practice than worst-case Wasserstein analyses would predict. We recast PE as generative model-augmented Wasserstein learning. We show theoretically that when we take into account the use of a generative model that is able to capture something about the true distribution, then we can obtain much better performance bounds. For example, if the generator gives samples in the same low-dimensional space as the distribution, then sample complexity depends on intrinsic, not ambient, dimension. We also show that standard variants of PE can fail to converge on simple well-clustered instances, and propose a new geometry-aware version of PE with provable convergence on such instances. Experimentally, we show that our new algorithm is competitive with standard baselines and can improve recall.

cs.LG↗

MLToolBench: Learning Tool-Augmented Agents for Machine Learning Development

Machine learning engineering (MLE) agents have made substantial progress, but learning through ML experimentation remains costly in time and computation. Synthetic environments reduce these costs while introducing variations in data and experimental settings that require task-specific diagnosis. Access to diagnostic tools alone does not ensure that agents learn when to use them or how to act on their findings. We introduce ToolMLBench, a suite of executable tools for data inspection, code verification, and experiment diagnosis, together with an SFT and RL pipeline for learning their use. Diagnostic calls acquire evidence whose value depends on subsequent decisions, so final outcomes provide limited guidance on which calls to reinforce. We address this challenge with SPICE, which measures how privileged context changes the likelihood of a sampled tool action and uses this difference as a turn-level reward alongside the final outcome. We train on 80 synthetic tasks and evaluate on 25 in-domain and 10 out-of-domain tasks. Providing tool interfaces and descriptions alone yields inconsistent gains across unadapted models. With the same diagnostic interface, our training pipeline raises in-domain success from 24.8% to 52.4% for Qwen3-8B and from 35.6% to 69.2% for Qwen3.5-35B-A3B. The latter also improves from 31% to 48% out-of-domain, supporting learned diagnostic tool use on held-out sources and targets.

cs.AI↗

Reprogramming Vision-Language Models via Structured Prompt Reparameterization

Visual reprogramming adapts pretrained models to downstream tasks by modifying their input and output interfaces while keeping the backbone fixed. In vision-language models, existing methods mainly rely on intra-class prompt aggregation and do not explicitly model relationships among classes. However, fine-grained categories often exhibit highly overlapping attribute descriptions and strong inter-class correlation in the text embedding space, where discriminative cues lie in subtle low-variance components. We propose Reparameterized Inter-Class Visual Reprogramming (RVP), a structured framework that aggregates multiple text prompts within each class and applies residual correction across classes. We also show that CLIP-based visual reprogramming with input-independent linear output aggregation can be expressed as a linear mapping from frozen image embeddings to downstream logits, and use this view to design a structured reparameterization that models shared semantic components and class-specific differences. RVP uses only a single visual prompt and can be reparameterized at inference into a frozen backbone followed by a linear classifier, incurring nearly zero computational overhead. Across 11 few-shot classification benchmarks and four CLIP backbones, RVP consistently improves over prior visual reprogramming methods with comparable or better inference efficiency.

cs.CV↗

Averaging in thermodynamic dislocation theory: general macroscopically uniform stress and strain states

The averaging procedure, developed for polycrystalline bars under axially symmetric tension or compression, is extended to arbitrary macroscopically uniform stress and strain states. Starting from the equal probability hypothesis for grain orientations, the mean resolved shear stress and the mean resolved elastic and plastic shear strains are defined as root-mean-square averages over all slip-system orientations. Two exact identities for isotropic orientation averages show that the mean resolved shear stress is proportional to the von Mises equivalent stress, and that the direction of macroscopic plastic flow, obtained from the hypothesis that the plastic slip rate on a system is proportional to the resolved shear stress acting on it, is the stress deviator. The result is an associated $J_2$ flow theory whose hardening law is not fitted but follows from the kinetics of thermally activated dislocation depinning and the evolution equations for the dislocation density and the effective disorder temperature of thermodynamic dislocation theory, written as rates with respect to time so that arbitrary loading paths can be followed. For torsion the theory yields the torque--twist relation of bars and tubes without the classical reductions to a shear stress--strain curve. The parameters for copper are identified jointly from Hopkinson-bar tension, dynamic compression at room and elevated temperatures, and torque--twist records at several twist rates. One set of material parameters, consistent with the earlier compression-only identification, describes tension, compression and torsion from room temperature to $1173$\,K and from $10$ to $2300$\,s$^{-1}$ to $6$\,\% rms over 142 data points; the tension--torsion discrepancy noted by Johnson and Cook is traced to the initial dislocation state of their torsion specimens.

cond-mat.mtrl-sci↗

Polar-Domain Volume as a Unified Descriptor of Transport in Ionic Liquids

Developing molecular scale structural descriptors that can quantitatively predict the macroscopic transport properties of ionic liquids remains a longstanding question, owing to the complex interplay between molecular interactions, nanoscale organization, and collective ion dynamics. Here, we use all atom molecular dynamics simulations to establish quantitative structure property relationships between the equilibrium microstructure and two key transport properties viscosity and ionic conductivity in a series of imidazolium based ILs with chemically distinct anions and varying alkyl chain lengths. While viscosity and ionic conductivity exhibit distinct dependencies on alkyl chain length and anion chemistry, these seemingly different trends largely collapse onto unified correlations when expressed in terms of the mean polar-domain volume. In particular, both transport properties exhibit systematic power-law scaling with the mean polar domain volume fraction, revealing a common structural origin underlying the variations in ion and momentum transport across chemically distinct ILs. These results establish polar-domain volume as a physically motivated molecular-scale descriptor and provide a general structure property framework for connecting nanoscale organization to macroscopic transport properties in ILs.

cond-mat.mtrl-sci↗