SearcharxivSearch

arXiv subjects

Zhenjie Ren

Publications and source records attributed to Zhenjie Ren.

At least 19 recordsLinked to original sources

Exponential convergence of Sinkhorn algorithm for entropy martingale optimal transport

We prove the exponential convergence in relative entropy of the Sinkhorn algorithm for the entropy martingale optimal transport. We assume that the marginals have compact supports and are in strict convex order, the terminal marginal support is convex, and the initial marginal support lies in the relative interior of the terminal marginal support; the reference cost is assumed to be Lipschitz in each variable. Under these assumptions we establish uniform bounds on the dual variables modulo affine gauges. This result allows us to obtain the existence and uniqueness of the optimizer, and its exponential representation in terms of dual variables. We then establish a relative-entropy stability estimate for martingale couplings with different terminal marginals. Proving that the constant in the stability estimate stays uniform throughout the Sinkhorn iteration allows us to establish the exponential convergence of the Sinkhorn algorithm.

math.PR

Self-fictitious-play for Potential Monotone Ergodic Mean-field Games

We investigate long-time learning in ergodic, potential, monotone mean-field games (MFGs) via a self-fictitious-play (SFP) dynamics coupling an optimally controlled diffusion with a slowly evolving belief. At each time, the state follows the optimal feedback associated with the current belief, while the belief is updated using the player's own empirical occupation measure rather than the population distribution. For ergodic monotone potential MFGs on the torus, we prove that the SFP dynamics is contractive and admits a unique invariant law. Moreover, we show that this invariant law is quantitatively close to the MFG Nash equilibrium, with an error of order equal to the square root of the belief-update rate. The proof combines uniform-in-time regularity estimates for the ergodic Hamilton-Jacobi-Bellman equation with an energy argument based on the Lasry-Lions divergence. The linear-quadratic example shows that this rate is sharp, and the numerical experiments illustrate the predicted scaling.

math.OC

Exponential Convergence of the Sinkhorn Algorithm for the Schrödinger Bridge with Regime Switching

This paper studies the convergence of the Sinkhorn algorithm for the Schrödinger bridge problem with regime switching, as introduced in Zlotchevski and Chen (2025). We consider a class of regime-switching stochastic systems on the hybrid state space $E=\mathbb{R}^d\times\{1,\ldots,m\}$, and construct the Sinkhorn iteration through the associated equivalent entropic optimal transport formulation. The main result of this paper is the exponential convergence of the Sinkhorn algorithm in relative entropy under compactness assumptions. Our proofs are inspired by the arguments recently developed for proving exponential convergence of the classical Schrödinger problem in Chiarini, Conforti, Greco and Tamanini (2024), Eckstein (2025). We perform a similar analysis for the partially observed terminal setting, where only the marginal distribution of the continuous component is prescribed at the terminal time, while the discrete regime is unobserved.

math.PR

Generalized specific entropy on Wiener space with application to Martingale Optimal Transport

Classical entropy regularization is poorly suited to continuous-time martingale transport, since relative entropy between diffusion laws typically forces their volatility characteristics to coincide. We introduce a specific-entropy framework based on Poisson jump approximations of continuous martingales. In the Gaussian-mark case, this yields explicit generalized specific entropy functionals on Wiener space, whose limiting costs depend not only on the limiting martingale laws but also on the microscopic approximation mechanism. This Poissonization approach avoids deterministic grid refinement and the associated high-dimensional multimarginal Sinkhorn problems, while allowing jump intensities to reflect local volatility. We prove weak convergence of the Poisson approximations and identify the limiting entropy functionals. For a trace-normalized Poisson scheme, the resulting cost defines a continuous-time specific-entropic martingale optimal transport problem, called SEMOT. This cost yields compactness, existence, and strong duality, and leads formally to a coupled Hamilton-Jacobi-Bellman/Fokker-Planck system. The resulting structure suggests Sinkhorn type numerical schemes, which we implement in one and two dimensions.

math.PR

Generative Transfer for Entropic Optimal Transport with Unknown Costs

This paper addresses the practical challenge in Entropic Optimal Transport (EOT) where the underlying ground cost function is typically latent and unobserved. Rather than assuming a fixed geometric cost, we adopt a data-driven approach where a shared cost is revealed only through samples from a reference optimal coupling. The question is then: given samples from a reference optimal coupling, can we recover the optimal coupling for new marginals under the same latent cost? We introduce a generative transfer framework that recovers the optimal coupling for new marginals by utilizing an iterative path-wise tilting algorithm. Unlike static importance reweighting, this method evolves the coupling jointly with a marginal transport path, allowing mass to move beyond the reference support. We derive sample-level learning rules for these infinitesimal updates, which yield covariance-type evolution equations for the associated transport vector fields. By integrating this dynamics with Conditional Flow Matching (CFM), we produce a practical sampler for paired data. Finally, we provide theoretical guarantees establishing a global convergence rate of \mathcal{O}(δ), ensuring the generated coupling converges to the target EOT plan in W_1 distance.

math.OC

Discrete Flow Matching: Convergence Guarantees Under Minimal Assumptions

Flow Matching has recently emerged as a popular class of generative models for simulating a target distribution $μ_1$ from samples drawn from a source distribution $μ_0$. This framework relies on a fixed coupling between $μ_0$ and $μ_1$, and on a deterministic or stochastic bridge to define an interpolating process between the two distributions. The time marginals of this process can then be approximately sampled by estimating the transition rates, or more generally the generator, of its Markovian projection. This framework has recently been extended to the case of discrete source and target distributions, under the name Discrete Flow Matching (DFM). However, theoretical guarantees for such models remain scarce. In this paper, we study two DFM models on $\mathbb{Z}_m^d = \{0,\ldots,m-1\}^d$, sampled through time discretization, and derive non-asymptotic associated bounds for both of them. In contrast to previous work, we establish non-asymptotic bounds in Kullback--Leibler divergence for the early-stopped version of the target distribution. We also derive explicit convergence guarantees in total variation distance with respect to the true target distribution. Importantly, these bounds rely only on an approximation error assumption, relaxing standard score assumptions used in earlier works, while also yielding improved dependence on the vocabulary size $m$ and the dimension $d$.

cs.LG

Continuous-time q-learning for mean-field control with common noise, part-I: Theoretical foundations

This paper investigates the continuous-time counterpart of the Q-function for entropy-regularized mean-field control (MFC) with controlled common noise, coined as q-function by Jia and Zhou (2023) in the single agent's model. We first show that, under discretely sampled actions, the value function in the exploratory formulation converges to the one in the relaxed control formulation as the time grid refines. Leveraging the relaxed control formulation, we derive the exploratory Hamilton-Jacobi-Bellman (HJB) equation, in which the controlled common noise gives rise to an additional nonlinear functional of policy, rendering the policy iteration intricate. Under certain concavity condition, we establish the existence and uniqueness of the optimal one-step policy iteration via a first-order condition using the partial linear functional derivative with respect to policy. The policy improvement at each iteration is verified by relating to an entropy-regularized optimization problem over the space of policies. In the mean-field setting, we introduce the integrated q-function (Iq-function) defined on the state distribution and the policy, and it is shown that an optimal policy is identified as a two-layer fixed point to the argmax operator of the Iq-function. Finally, we provide the explicit characterization of an optimal policy as a Gaussian distribution in the general linear-quadratic (LQ) setting.

math.OC

Continuous-time q-learning for mean-field control with common noise, part-II: q-learning algorithms

This paper is a continuation work of Ren et al. (2026) aiming to further devise q-learning algorithms for mean-field control (MFC) with controlled common noise. Based on the relaxed control formulation, we first establish the martingale condition of the value function and the Iq-function by evaluating along the conditional state distributions generated by all test policies. As the data in the relaxed control formulation are not observable in practice, we quantify the error incurred when they are replaced by the observable ones in the exploratory formulation under discretely sampled actions. This, together with a two-layer fixed point characterization of an optimal policy in Ren et al. (2026), allows us to propose several algorithms including the Actor-Critic q-learning algorithm, in which the policy is updated in the Actor-step based on the iteration rule induced by the improved Iq-function, and the value function and Iq-function are updated in the Critic-step based on the martingale orthogonality condition using the data from the exploratory formulation. We also establish the convergence of the inner iterations in the Actor-step in an infinite-horizon linear quadratic (LQ) framework. In two examples, within and beyond LQ framework, our q-learning algorithms are implemented with satisfactory performance.

math.OC

Convergence of Sinkhorn's Algorithm for Entropic Martingale Optimal Transport Problem

In this paper, we study the Entropic Martingale Optimal Transport (EMOT) problem on \mathbb{R}. The investigation of the EMOT problem arises in the calibration problem of the Stochastic Volatility Models, where martingale constraints reflect no-arbitrage pricing conditions under the risk-neutral measure, as originally proposed by Henry-Labordere. We first establish the dual formulation of the EMOT problem and prove that Sinkhorn's algorithm achieves an exponential convergence rate under mild conditions. Notably, our analysis does not presuppose the existence of optimal potentials and rigorously confirms the absence of a primal-dual gap. These results provide a theoretical foundation for solving EMOT via Sinkhorn's method and constructing the optimal distribution from dual coefficients.

math.PR

Quantitative weak propagation of chaos for McKean--Vlasov branching diffusion processes

We study in this paper the weak propagation of chaos for McKean--Vlasov diffusions with branching, whose induced marginal measures are nonnegative finite measures but not necessary probability measures. The flow of marginal measures satisfies a non-linear Fokker--Planck equation, along which we provide a functional Itô's formula. We then consider a functional of the terminal marginal measure of the branching process, whose conditional value is solution to a Kolmogorov backward master equation. By using Itô's formula and based on the estimates of second-order linear and intrinsic functional derivatives of the value function, we finally derive a quantitative weak convergence rate for the empirical measures of the branching diffusion processes with finite population.

math.PR

Entropic Optimal Transport Problem with Convex Functional Cost

We study an entropic optimal transport problem in which the transport plan is penalized by a nonlinear convex functional acting on the coupling. We establish existence, uniqueness, and uniform a priori bounds for minimizers, and we show that each minimizer satisfies a fixed-point first-order optimality system associated with an exponentially tilted reference measure. Building on this variational structure, we introduce the Sinkhorn-Frank-Wolfe (SFW) flow, prove its global well-posedness, and derive an energy-dissipation inequality yielding exponential convergence toward the unique optimal transport plan. As an application, we implement the SFW algorithm to solve an optimal routing problem for unmanned aerial vehicles with congestion aversion.

math.OC

Dimension-free error estimate for diffusion model and optimal scheduling

Diffusion generative models have emerged as powerful tools for producing synthetic data from an empirically observed distribution. A common approach involves simulating the time-reversal of an Ornstein-Uhlenbeck (OU) process initialized at the true data distribution. Since the score function associated with the OU process is typically unknown, it is approximated using a trained neural network. This approximation, along with finite time simulation, time discretization and statistical approximation, introduce several sources of error whose impact on the generated samples must be carefully understood. Previous analyses have quantified the error between the generated and the true data distributions in terms of Wasserstein distance or Kullback-Leibler (KL) divergence. However, both metrics present limitations: KL divergence requires absolute continuity between distributions, while Wasserstein distance, though more general, leads to error bounds that scale poorly with dimension, rendering them impractical in high-dimensional settings. In this work, we derive an explicit, dimension-free bound on the discrepancy between the generated and the true data distributions. The bound is expressed in terms of a smooth test functional with bounded first and second derivatives. The key novelty lies in the use of this weaker, functional metric to obtain dimension-independent guarantees, at the cost of higher regularity on the test functions. As an application, we formulate and solve a variational problem to minimize the time-discretization error, leading to the derivation of an optimal time-scheduling strategy for the reverse-time diffusion. Interestingly, this scheduler has appeared previously in the literature in a different context; our analysis provides a new justification for its optimality, now grounded in minimizing the discretization bias in generative sampling.

stat.ML

Uniform-in-time propagation of chaos for mean field Langevin dynamics

We study the mean field Langevin dynamics and the associated particle system. By assuming the functional convexity of the energy, we obtain the $L^p$-convergence of the marginal distributions towards the unique invariant measure for the mean field dynamics. Furthermore, we prove the uniform-in-time propagation of chaos in both the $L^2$-Wasserstein metric and relative entropy.

math.PR

Size of chaos for Gibbs measures of mean field interacting diffusions

We investigate Gibbs measures for diffusive particles interacting through a two-body mean field energy. By identifying a gradient structure for the conditional law, we derive sharp bounds on the size of chaos, providing a quantitative characterization of particle independence. To handle interaction forces that are unbounded at infinity, we study the concentration of measure phenomenon for Gibbs measures via a defective Talagrand inequality, which may hold independent interest. Our approach provides a unified framework for both the flat semi-convex and displacement convex cases. Additionally, we establish sharp chaos bounds for the quartic Curie-Weiss model in the sub-critical regime, demonstrating the generality of this method.

math.PR

Self-interacting approximation to McKean-Vlasov long-time limit: a Markov chain Monte Carlo method

For a certain class of McKean-Vlasov processes, we introduce proxy processes that substitute the mean-field interaction with self-interaction, employing a weighted occupation measure. Our study encompasses two key achievements. First, we demonstrate the ergodicity of the self-interacting dynamics, under broad conditions, by applying the reflection coupling method. Second, in scenarios where the drifts are negative intrinsic gradients of convex mean-field potential functionals, we use entropy and functional inequalities to demonstrate that the stationary measures of the self-interacting processes approximate the invariant measures of the corresponding McKean-Vlasov processes. As an application, we show how to learn the optimal weights of a two-layer neural network by training a single neuron.

math.PR

Deciding Bank Interest Rates -- A Major-Minor Impulse Control Mean-Field Game Perspective

Deciding bank interest rates has been a long-standing challenge in finance. It is crucial to ensure that the selected rates balance market share and profitability. However, traditional approaches typically focus on the interest rate changes of individual banks, often neglecting the interactions with other banks in the market. This work proposes a novel framework that models the interest rate problem as a major-minor mean field game within the context of an interbank game. To incorporate the complex interactions between banks, we utilize mean-field theory and employ impulsive control to model the overhead in rate adjustments. Ultimately, we solve this optimal control problem using a new deep Q-network method, which iterates the parameterized action value functions for major and minor players and updates the networks in a Fictitious Play way. Our proposed algorithm converges, offering a solution that enables the analysis of strategies for major and minor players in the market under the Nash Equilibrium.

math.OC

Time-uniform log-Sobolev inequalities and applications to propagation of chaos

Time-uniform log-Sobolev inequalities (LSI) satisfied by solutions of semi-linear mean-field equations have recently appeared to be a key tool to obtain time-uniform propagation of chaos estimates. This work addresses the more general settings of time-inhomogeneous Fokker-Planck equations. Time-uniform LSI are obtained in two cases, either with the bounded-Lipschitz perturbation argument with respect to a reference measure, or with a coupling approach at high temperature. These arguments are then applied to mean-field equations, where, on the one hand, sharp marginal propagation of chaos estimates are obtained in smooth cases and, on the other hand, time-uniform global propagation of chaos is shown in the case of vortex interactions with quadratic confinement potential on the whole space. In this second case, an important point is to establish global gradient and Hessian estimates, which is of independent interest. We prove these bounds in the more general situation of non-attractive logarithmic and Riesz singular interactions.

math.PR

Uniform-in-time propagation of chaos for kinetic mean field Langevin dynamics

We study the kinetic mean field Langevin dynamics under the functional convexity assumption of the mean field energy functional. Using hypocoercivity, we first establish the exponential convergence of the mean field dynamics and then show the corresponding $N$-particle system converges exponentially in a rate uniform in $N$ modulo a small error. Finally we study the short-time regularization effects of the dynamics and prove its uniform-in-time propagation of chaos property in both the Wasserstein and entropic sense. Our results can be applied to the training of two-layer neural networks with momentum and we include the numerical experiments.

math.PR