Searcharxiv⌕ Search

arXiv · 2609.32708

One-Step Generative Modeling via Unbalanced Optimal Transport

Abstract

Drifting models enable one-step generation by amortizing distribution transport into training, but this efficiency places greater demands on the transport field estimated at each update. In large-scale training, the field is computed from finite mini-batches of generated and real samples, which provide only imperfect approximations to the underlying distributions. Balanced optimal transport enforces exact mass matching within every mini-batch, making the estimated field sensitive to the particular composition of the real-data batch. We find that generated and real samples should be treated asymmetrically: letting the mass assigned to real samples adapt while keeping every generated sample fully transported improves generation across six feature-space metrics in controlled ablations, and is more robust to the relaxation strength than relaxing both marginals simultaneously, which falls below balanced transport under stronger relaxation. Motivated by this observation, we propose Unbalanced Optimal Transport Gradient Flow (UOT-GF), which keeps the generated-sample marginal fixed and relaxes only the real-data marginal. Under identical settings at DiT-B/2 on ImageNet-256, UOT-GF improves Fréchet Inception Distance (FID) from 1.53 to 1.46 over the balanced W-Flow baseline; scaling the same recipe yields 1.34 and 1.22 FID at L/2 and XL/2, the best FID among the one-step models we compare. We further derive the induced UOT transport force, establish a kinetic Vlasov--Fokker--Planck formulation whose overdamped zero-temperature limit recovers the drifting dynamics, and characterize non-target stationary states together with sufficient conditions for convergence.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yirong Shen, Mengfei Xia, Junpeng Jing, Lu Gan, Cong Ling. 2026-09-26. One-Step Generative Modeling via Unbalanced Optimal Transport. https://arxiv.org/abs/2609.32708

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Anonymous Shamir's Secret Sharing via Reed-Solomon Codes Against Permutations, Insertions, and Deletions

In this work, we study the performance of Reed-Solomon codes against an adversary that first permutes the symbols of the codeword and then performs insertions and deletions. This adversarial model is motivated by the recent interest in fully anonymous secret-sharing schemes [EBG+24],[BGI+24]. A fully anonymous secret-sharing scheme has two key properties: (1) the identities of the participants are not revealed before the secret is reconstructed, and (2) the shares of any unauthorized set of participants are uniform and independent. In particular, the shares of any unauthorized subset reveal no information about the identity of the participants who hold them. In this work, we first make the following observation: Reed-Solomon codes that are robust against an adversary that permutes the codeword and then deletes symbols from the permuted codeword can be used to construct ramp threshold secret-sharing schemes that are fully anonymous. Then, we show that over large enough fields of size, there are $[n,k]$ Reed-Solomon codes that are robust against an adversary that arbitrary permutes the codeword and then performs $n-2k+1$ insertions and deletions to the permuted codeword. This implies the existence of a $(k-1, 2k-1, n)$ ramp secret sharing scheme that is fully anonymous. That is, any $k-1$ shares reveal nothing about the secret, and, moreover, this set of shares reveals no information about the identities of the players who hold them. On the other hand, any $2k-1$ shares can reconstruct the secret without revealing their identities. We also provide explicit constructions of such schemes based on previous works on Reed-Solomon codes correcting insertions and deletions. The constructions in this paper give the first gap threshold secret-sharing schemes that satisfy the strongest notion of anonymity together with perfect reconstruction.

cs.IT↗

Unequal Error Protection for Digital Semantic Communication with Channel Coding

This paper investigates unequal error protection (UEP) in digital semantic communication, where semantically important bits require substantially higher reliability than less critical ones. To characterize this heterogeneity, we introduce a novel perspective that treats learned bit-flip probabilities of semantic bits as target error protection levels, thereby directly linking semantic importance to bit-level reliability. This formulation reveals that the required protection levels of the semantic bits may differ by several orders of magnitude, making short-block coding more advantageous than conventional long-block designs. Motivated by this, we develop two UEP frameworks that minimize total blocklength under heterogeneous reliability constraints. First, we propose a bit-level UEP framework based on repetition coding, providing an analytically tractable solution that precisely meets per-bit protection requirements. Second, to improve energy and blocklength efficiency, we design a block-level UEP framework in which the semantic bits are partitioned into short blocks with similar protection levels. Guided by finite blocklength capacity analysis, we derive a closed-form threshold condition for beneficial partitioning and develop a systematic algorithm for integrating modern channel codes. Simulation results on image transmission tasks demonstrate substantial gains in both task performance and transmission efficiency compared with conventional equal-protection schemes.

cs.IT↗

Large Language Models are Shannon Lossy Compressors Not Solomonoff Induction Estimators: Self-improvement and Singularity Are Not Near Without Symbolic Model Synthesis

We connect two questions in Algorithmic Information Theory (AIT), Machine Learning (ML) and Artificial General Intelligence (AGI): whether LLMs estimate Solomonoff induction, and whether they can self-improve towards an AI Singularity. We provide theoretical, methodological and empirical answers in the negative but show how limits can be circumvented. Cross-entropy, negative log-likelihood and related next-token objectives cannot alone implement Solomonoff induction: they fit supplied conditionals rather than a program-weighted universal mixture. More computation can improve fit within a fixed objective but cannot change its inductive principle without external hyperparameter or architectural tuning; they alone do not deliver Solomonoff-Levin optimal prediction. For finite learners and observers, theoretical boundaries become less decisive and approaches diverge. Resource-bounded estimators are finite mechanism-search tools whose divergence does not violate algorithmic information conservation. All 26 served language-model checkpoints across five pre-training families, 0.8-35 billion parameters and 1.9-8.5 bits per weight, evaluated at their commitments over a closed alphabet, violate the dominance guarantee defining a universal mixture. Against a 3.32-bit bound attained by a genuine mixture, the best model trails a Krichevsky-Trofimov code by 4.5 bits, the median by 36 and the worst by 128; excess grows to every stream's end rather than settling to a constant. Served conditionals fail to form a mixture over the declared class in 79 of 91 checkpoint-designs; neither scale nor post-training closes the gap. Frontier developers adopt neurosymbolic approaches, including Fable and Astra, incorporating model synthesis via neurosymbolic computation. They are no longer purely statistical LLMs, making them better, though still limited, candidates for higher forms of induction & model synthesis.

cs.IT↗