Searcharxiv⌕ Search

arXiv subjects

Search papers

Search indexed arXiv papers on artificial intelligence, large language models, computer vision and robotics. Read source abstracts and follow links to arXiv.

At least 1,369 records · Page 76Linked to original sources

TV-Regulated OPD: Direction Matters in On-Policy Distillation

On-Policy Distillation (OPD) facilitates the transfer of knowledge from domain expert to student in the post-training phase of Large Language Models (LLMs). However, the supervision signals in mainstream OPD methods suffer from high variance and noise which is generally instable during training. In this work, we systematically investigated what really matters to the performance and the fundamental mechanisms behind the instability during training. We found that retaining only the sign of token-level advantages is sufficient to achieve the performance comparable to standard OPD. Meanwhile, smoother and bounded advantages can stabilize the training process without sacrificing its performance. These motivated us to shape the advantages using the Total Variation (TV) and propose a robust TV regulated On-Policy Distillation (TV-OPD) method. Benefiting from the bounded and diminished advantages, TV-OPD exhibits stable training dynamics and steady late-stage performance. We conducted comprehensive experiments and found that, across various settings, TV-OPD consistently achieved better performance and lower variance in the late-stage of training.

cs.LG↗

SRPO: Setwise Relative Policy Optimization for Multi-Agent Systems

Multi-agent systems enable complex reasoning and tool use by coordinating agents that divide roles and refine candidate solutions. Existing methods typically update individual agent responses or treat a complete trajectory as one training example. However, these methods may produce misleading policy updates because they assign the same final outcome to responses or trajectory segments that may play different roles in different team decisions. This is because treating each response as an independent update may separate outputs that jointly determine the next action, while treating an entire trajectory as one update may combine decisions made after different observations. These limitations call for a policy update defined at the level of a team decision, outputs that lead to the same state transition are optimized under a shared objective. In this paper, we propose Setwise Relative Policy Optimization (SRPO) for multi-agent systems. We represent the outputs used together to produce one state transition as an active set. A singleton set covers division of labor, while a larger set covers joint co-evolution, and the set composition can change across decisions. SRPO assigns a shared advantage to each active set, clips the combined policy change, and normalizes its scale according to the set size. This ties each update to the decision that produced the next state. Experiments on mathematical reasoning and multi-turn search demonstrate the effectiveness of SRPO across both tasks. Additional ablation studies analyze the normalization choice and training behavior under changing active-set sizes.

cs.AI↗

TASG-Explore: Traversability-Aware Sector-Guided Exploration for Ground Robot on Uneven Terrain

Autonomous exploration on uneven terrain requires ground robots to balance exploration efficiency, coverage completeness, and terrain safety. Detailed tsrrain reasoning improves local reliability but can slow large-scale exploration, whereas coarse region guidance expands quickly in open areas but can miss narrow passages and irregular traversable boundaries. To address this challenge, this paper presents TASG-Explore, a traversability-aware sector-guided exploration framework for ground robots. The framework first performs hierarchical traversability analysis using variable-voxel ground fitting and adaptive 8-bit obstacle encoding. It then splitting cost map into sectors, incrementally updates sector clusters, extracts terrain-coupled frontier viewpoints, and maintains a dynamic topological roadmap with unknown topological hypotheses. Finally, a sector-guided planner selects region targets and inserts local viewpoints to generate efficient exploration routes. Benchmark experiments in diverse challenging environments, including caves, forests, and rugged hills, show that TASG-Explore achieves the best overall performance among six representative state-of-the-art planners. The proposed traversability analysis improves processing efficiency by 6.3 times while maintaining high accuracy, and the exploration planner improves exploration efficiency by 51% and increases coverage by up to 2.95 times in rugged hill scene. Large-scale real-world experiments further demonstrate the practical value of the proposed method.

cs.RO↗

AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems

Large language model (LLM)-based multi-agent systems (MAS) achieve strong performance by employing specialized multiple agents, yet their performance depends on the prompt design of each agent. For MAS prompt optimization, textual gradient methods that guide prompt updates using natural-language feedback have emerged as a leading paradigm. In this paper, we identify limitations in two stages of existing textual gradient approaches: gradient extraction and gradient aggregation. In gradient extraction, previous works select a target prompt without verifying whether modifying it resolves the failure, and derive gradients without agent-level supervision over the corresponding agent's intermediate output. In gradient aggregation, individual gradients are randomly grouped and concatenated, often mixing unrelated failure modes and producing prompts that fail to generalize. To address these limitations, we propose AgentGrad, a prompt optimization framework for multi-agent systems based on sequential intervention and semantic textual gradient abstraction. For each failure, sequential intervention modifies the behavior of one agent at a time to identify the target agent whose modification resolves the failure. The modified output of the target agent then serves as agent-level supervision for extracting a fine-grained gradient. Semantic textual gradient abstraction clusters semantically similar gradients to prevent mixing unrelated failure modes, and abstracts each cluster into a generalized gradient that captures the shared corrective pattern. Experimental results show that AgentGrad achieves state-of-the-art performance across five MAS benchmarks while reducing wall-clock optimization time by $2.5\times$ and optimization cost by 21.8\% on average compared to the next-best baselines.

cs.AI↗

New Generalizations of Two Ramanujan Series for $1/π$

By utilizing two known hypergeometric summation identities, we extend two \mbox{well-known} Ramanujan series for $1/π$ by making each of them a member of an infinite family of series. We also find the corresponding families involving harmonic numbers and odd harmonic numbers. To further illustrate our method we derive families of \mbox{Ramanujan-like} series for $1/π$ associated with a hypergeometric series derived by Lavoie, two hypergeometric series from Bailey's book and a recent hypergeometric series found by Campbell. Thus, by leveraging known hypergeometric summation identities, our approach bypasses the more tedious, traditional limits or complex Fourier-Legendre expansions; offering a cleaner, unified template for generating Ramanujan-like series for $1/π$.

math.GM↗

A Sum-of-Squares Hierarchy with Quadratic Convergence for Quantum Channel Coding

Computing the optimal success probability for transmitting classical messages through a single use of a quantum channel is NP-hard, even for two messages. An existing semidefinite programming hierarchy based on symmetric extensions provides convergent upper bounds with an a priori error estimate that decays as the inverse square root of the extension level. In this work, we construct a Hermitian sum-of-squares hierarchy for an arbitrary number of messages and prove quadratic convergence in its level. The error bound is proportional to the advantage over random guessing. Our approach combines state-discrimination duality with positive polynomial kernels on products of spheres to construct feasible polynomial dual certificates. For binary messages, the resulting bounds give a multiplicative approximation from above of the trace-norm contraction coefficient.

quant-ph↗

Teacher-Free Self-Distilled Consistency Trajectory Learning for Fast Speech Enhancement

Consistency trajectory models offer a route to fast, high- quality speech enhancement, collapsing the many reverse steps of diffusion-based enhancers into a handful. When instantiated on a Schrödinger bridge (SB), which pins the generative process to fixed clean and noisy endpoints, exist- ing consistency-trajectory enhancers (SBCTMs) still require a pretrained teacher to supply trajectory supervision, which raises training cost and ties the final quality to that of the teacher. We propose a teacher-free, self-distilled consistency- trajectory framework that removes the external teacher result- ing in a 5X reduction in per epoch training time. Our model is trained with a three-stage curriculum of clean speech pre- diction, a self-distilled shortcut objective, and perceptual fine-tuning with a multi-resolution short-time Fourier trans- form (MR-STFT) loss. Using the same NCSN++ backbone as SBCTM, our model attains a wide-band PESQ of 3.01, ES- TOI 0.87, and SI-SDR 19.07 dB on VoiceBank+DEMAND compared to 3.57, 0.87 and 12.8 dB for the teacher based model. Further, we find that a geometric schedule at low reverse step count maximizes perceptual quality, while a higher-step uniform schedule favors signal fidelity.

eess.AS↗

Deformation families of open Calabi-Yau manifolds and Steinness

This paper is to construct deformation families of open Calabi-Yau manifolds together with some discussions on Steinness. Firstly, we construct a nontrivial compactifiable deformation of open Calabi-Yau manifolds from rational elliptic surface with global sections. Secondly, we give a condition for the existence of an elliptic fibration with global sections on the complex analytic family of rational elliptic surface with global sections, construct a holomorphic involution on the compactifiable deformation family and claim the non-Steinness for all the compactifiable deformation families constructed. Finally, we give some discussions on the non-Steinness of the quasi-projective variety from $\text{CP}^2 \text{$\#$9} \overline{\text{CP}}^2$ corresponding to a conjecture by Marco Brunella and construct deformation families whose fibers were conjectured to be Stein by Andreas Höring and Thomas Peternel.

math.AG↗

Testing the Binary Rank with Polynomial Query Complexity

We design an adaptive two-sided error testing algorithm for the binary rank of a $0,1$ matrix $M$ with query complexity $O(d^3\log(d+1)/ε^2)$, where $d$ is the tested binary rank bound and $ε$ is the distance parameter. This answers an open question posed by Parnas, Ron and Shraibman~\cite{parnas2021property}, who asked if the binary rank can be tested with query complexity polynomial in $d$ and $1/ε$. Furthermore, our testing algorithm can be used to find an approximate binary decomposition of $M$ with an additional $d(n+m)$ queries. That is, under the promise that the binary rank of $M$ is at most $d$, we show how to find, with probability at least $5/6$, two $0,1$ matrices $A',B'$ such that $M' = A' \cdot B'$ is a $0,1$ matrix which differs from $M$ on at most an $O(ε)$ fraction of its entries. Our results also imply a testing algorithm with polynomial query complexity for the equivalent problem of testing if the edges of a bipartite graph can be partitioned into at most $d$ bicliques.

cs.DS↗

AcFlow: Controlling Text-to-Image Diffusion Transformers via Learned Conditional Activation Flow

Text-to-image diffusion transformers (DiTs) are powerful generators, yet direct prompting provides limited control interface for style intensity and can fail to suppress unwanted concepts. To enable these controls, we introduce AcFlow, an inference-time controller that transports intermediate layer image-token activations through a learned concept-conditioned velocity field while keeping the base DiT frozen. A textual concept description specifies the desired intervention, while the integration horizon provides a continuous control parameter. The field produces token-varying, activation-dependent updates. With parameters shared across concepts within each task family, one field covers over 15,000 style descriptions or over 1,000 suppression concepts, and generalizes to concepts unseen during training without per-concept fitting. On style control, AcFlow achieves the best style--content trade-off among the evaluated baselines in the high-style-alignment regime. At a fixed operating point, AcFlow attains style--content alignment of 0.5365/0.2860, compared with 0.4397/0.2684 for the baseline with the highest style alignment. On concept suppression, AcFlow reduces the fraction of images showing the concept from 95.3%/82.1% to 41.6%/40.5% on held-in/held-out concepts, including cases where deleting them from the prompt fails to remove them. Our analyses support the learned velocity field as an adaptive control mechanism, with update directions varying across tokens and depending on their activation states. Our code is available at https://github.com/Nove1yst/AcFlow.

cs.CV↗

Detectable Only Where It Is Confounded: What Verified Duplication Counts Say About Membership Evidence in Language Models

When a language model finds a sentence unusually cheap to predict, it is tempting to conclude that the sentence was in its training data. Almost every published test of that inference has had to guess which sentences were in the training data, the members, and which were not. This paper removes the guessing. Two model families, OLMo-2 and Pythia, publish their pretraining corpora, and a public index over those corpora returns the exact number of times any sentence appeared in each. Those counts make three questions answerable directly. The answers form a pincer, closing from two sides. At the duplication levels ordinary text actually has, five models from 1B to 13B parameters carry at most a faint trace of their own exposure. We measure that trace with a design that reads the same sentence through two models, which cancels fluency and quality by construction, and it comes to a rank correlation near -0.08, where -1 would be a perfect relation and 0 none. Where the trace does become strong, above roughly a thousand copies, the two corpora agree on which sentences those are, because they are the famous ones, so exposure can no longer be told apart from fame. Two further measurements show how apparent membership signal gets manufactured. A common way to build a non-member is to change one word of a member. The model does prefer the original, but the gap is the same, within noise, whether the original appeared once or a hundred times, so what the model is rewarding is the author's word choice, not memory. Above a thousand copies the gap grows with model size on the twelve sentences we can test there, at the same boundary where the pincer closes. And swapping the controls for sentences that differ from the members in register moves a detector from 0.83 to 0.94 AUC, on a scale where 0.5 is a coin flip and 1.0 is perfect separation. We release the sentence banks, counts, and code.

cs.CL↗

BEACON: A Versatile Accelerator for Computational Pathology Applications

While accelerators for AI have seen great commercial success, it is challenging to replicate that success for other specialized domains due to a number of factors. We make the case that barriers for new accelerators can be lowered by starting with a baseline AI accelerator, and adding minimal logic to support new operators demanded by new specialized domains. This leads to a versatile chip that can be manufactured at high volume and deployed for a range of popular applications. We refer to this as the AI+X approach. This paper explores its potential for the emerging domain of Computational Pathology, which involves analysis of large whole-slide tissue images with a multi-stage pipeline. The pipeline requires support for a number of different kernels and operators - early stages perform segmentation and feature extraction, followed by graph creation with k nearest neighbor (kNN) algorithms, and finally inference with an iterative graph convolutional network (GCN) that alternates between Aggregation and Combination. We show that these stages execute inefficiently on a range of baseline CPU, GPU, AI, and GCN accelerators. That inefficiency is addressed with a combination of software re-structuring and small modifications to a baseline systolic AI accelerator. Many of the above kernels can be mapped to a systolic accelerator by offering a flexible datapath between processing elements and register access mechanisms. We add support for feature aggregation, load balanced execution, Euclidean distance calculation, binning, and counter aggregation. This additional flexibility and logic grows the area of a baseline AI chiplet by 1.1x, but by avoiding the memory wall and offering high parallelism, the proposed accelerator BEACON yields over an order of magnitude higher throughput for Computational Pathology than baseline CPU and GPU platforms.

cs.AR↗

Fission Modes and Fragment Shell Structures in $^{258}$Md$^*$ from Six-Dimensional Langevin Calculations

The fission of $^{258}\mathrm{Md}^{*}$ is calculated in the excitation energy range of $E^{*} = 6$--$36$~MeV using a six-dimensional Langevin framework with the Cassini shape parametrization. The calculated events are classified into two symmetric and two asymmetric fission modes based on the fragment mass and the quadrupole deformations of the two fragments at scission. The symmetric modes are the short mode with high total kinetic energy (TKE) and the superlong mode with low TKE, whereas the asymmetric modes differ in mass asymmetry. With increasing excitation energy, the yield of the short mode decreases, whereas the combined yield of the two asymmetric modes increases, as observed in the in-beam prompt-fission study of $^{258}\mathrm{Md}^{*}$. From an analysis of the fragment shapes and associated single-particle levels, the short mode and the dominant asymmetric mode with the smaller mass asymmetry are found to involve a compact fragment characterized by deformed shell gaps at $Z = 52$ and $N = 84$, while the complementary fragment is compact in the short mode and strongly elongated in the asymmetric mode.

nucl-th↗

A Dominant Supplier Slows Recursive Drift More Than It Steers It

More and more of the text future language models learn from is written by a few of today's models. If one supplier writes most of a shared corpus, does it pull the models trained on it toward its own writing, or change how fast they drift? We retrain eight open models from their base weights on a shared pool of each other's text for five generations, varying the part written by one model, Phi-2, from an equal share to 90%. The models drift together toward a style with fewer function words, and none starts repeating itself. No share of Phi-2 brings the other models closer to its text than the equal share does. We split each ecosystem's separation from the equal-share one into a delay along its route and a departure from that route, both counted beyond the difference between two equal-share runs. With Phi-2 at 90%, delay outweighs departure 72 to 28 and 64 to 36 in two runs, and the ecosystem falls 2.7 and 2.5 generations behind. With Phi-2 at half the pool the two parts are about equal. When SmolLM2 or Qwen3-1.7B writes half instead, the ecosystem slows less or not at all. The departure leans toward Phi-2 more as its share grows, but more than toward every other model only at 90%. Human text filling a quarter or half of the pool slows the models along the same route.

cs.AI↗

Break Step: Recursive Training Resonates with Replayed Sampling Noise

How fast does a language model degrade when trained on its own outputs? Theory traces it to gradually accumulating errors, while experiments report repeated phrases within ten generations. Under a fixed sampling seed in vLLM, the fast loss of lexical diversity comes from the sampler. When vLLM serves a batch from one seeded sampling configuration, every request receives the same random draws, and a fixed seed replays them every generation. Fine-tuning raises the tokens that won, and the replayed draws let them win by more. Sharing across requests and replay across generations matter only together. Remove either one, by changing the shared seed every generation or by giving each request its own seed that repeats every generation, and the unique-4-gram fraction of two StableLM checkpoints stays near its starting value of about 0.98 through generation 3. Keep both, and the replayed shared seed takes seven checkpoints from five families to between 0.045 and 0.38 by then. Three generations of replay write the favoured phrases into the weights: decoded with one seed per request, the generation-3 weights of the replayed StableLM-2-1.6B chain recover most of their diversity, yet the phrase that filled every sample under the shared seed still opens 46% of them. Without replay, five checkpoints drift slowly, consistent with the gradual accumulation that theory describes, and three turn incoherent though their diversity scores stay high. One peer-reviewed model-collapse pipeline that fine-tunes Gemma-2-27B samples identical prompts under one seeded configuration, and three quarters of the rows it released for one iteration repeat nearly as often as one such batch copies them. A seed per request restores the fresh sample that stability analyses assume.

cs.CL↗

BridgeMatch: Conditional Transport Bridges in Matching Matrix Space for 3D Deformable Registration

Reliable matching between partially observed, deforming point clouds requires global context and fine geometric detail. Coarse candidate selection can exclude correct fine-level correspondences. We present \paper, a unified conditional transport framework with the matching matrix itself as the evolving state. Coarse diffusion establishes global matching hypotheses; hierarchy-preserving lifting expands them into a structured high-resolution source. Geometry-conditioned ODE and SDE bridges continue refinement in the complete fine-level candidate space, allowing coarse errors to be corrected. The deterministic endpoint-parameterized conditional flow matching (CFM) design improves matching through iteration, while the SDE forward drift enables effective few-step refinement. Experiments on 4DMatch and 4DLoMatch demonstrate competitive matching and non-rigid registration, with transfer to CAPE and DeepDeform without target-domain training.

cs.CV↗

Generating the wide sequence of Diffuse Galaxies with de Broglie waves of Dark Matter

Extensive Euclid satellite imaging at low surface brightness has revealed that most nearby galaxies are diffuse-looking spheroids, where the stellar radius increases for over three decades in luminosity. We argue this Diffuse Galaxy sequence results from internal stellar diffusion by Wave Dark Matter ($ψ$DM), as wave energy is transferred to star orbits over time. In particular, the soliton random motion scatters central stars onto radial orbits that become enhanced with each passage through the centre, slowly ``puffing up" the stellar profile. Heating is greater within massive galaxies as $ψ$DM fluctuations are stronger and more frequent, reproducing both the Diffuse Galaxy sequence and the rising velocity dispersion along the sequence, from Ultra-Faint to Dwarf Spheroidal and Ultra Diffuse galaxies, favouring a light boson, $m_ψ=2.88^{+0.14}_{-0.13}\times10^{-22}$eV. Winding back this diffusion, we predict the stellar content of Diffuse Galaxies, including globular clusters, formed near the centre, as anticipated by $ψ$DM simulations, where gas cools efficiently within the dense soliton. This predicted $ψ$DM evolution from compact beginnings towards diffuse spheroidal galaxies today can now be fully charted from JWST to Euclid.

astro-ph.GA↗