SearcharxivSearch

arXiv subjects

Hugo Lavenant

Publications and source records attributed to Hugo Lavenant.

At least 19 recordsLinked to original sources

Continuous transformations of probability measures and their transport representations

Given a function $F$ transforming a probability measure $\mu$ into another one $F(\mu)$, we study the existence and regularity of a transport representation of it. That is, we ask whether we can represent the image $F(\mu)$ of the input probability measure $\mu$ as the push-forward of $\mu$ by a map $f(\cdot,\mu)$ which may depend on $\mu$; and furthermore, how regular $f$ can be chosen depending on $F$. Even if $F$ is continuous and a transport representative exists, it cannot necessarily be chosen in a continuous way; however, if $F$ is Lipschitz continuous with respect to the Wasserstein distance, then $f$ can be chosen continuous. We provide several examples to illustrate the sharpness of our assumptions. This question is motivated by approximation results for transformations of probability distributions with transformers.

math.FA

Gradient Flows of Potential Energies in the Geometry of Sinkhorn Divergences

We analyze the gradient flow of a potential energy in the space of probability measures when we substitute the optimal transport geometry with a geometry based on Sinkhorn divergences, a debiased version of entropic optimal transport. This gradient flow appears formally as the limit of the minimizing movement scheme, a.k.a. JKO scheme, when the squared Wasserstein distance is substituted by the Sinkhorn divergence. We prove well-posedness and stability of the flow, and that, in the long term, the energy always converges to its minimal value. The analysis is based on a change of variable to study the flow in a Reproducing Kernel Hilbert Space, in which the evolution is no longer a gradient flow but described by a monotone operator. Under a restrictive assumption we prove the convergence of our modified JKO scheme towards this flow as the time step vanishes. We also provide numerical illustrations of the intriguing properties of this newly defined gradient flow.

math.AP

Error Bounds and Optimal Schedules for Masked Diffusions with Factorized Approximations

Recently proposed generative models for discrete data, such as Masked Diffusion Models (MDMs), exploit conditional independence approximations to reduce the computational cost of popular Auto-Regressive Models (ARMs), at the price of some bias in the sampling distribution. We study the resulting computation-vs-accuracy trade-off, providing general error bounds (in relative entropy) that depend only on the average number of tokens generated per iteration and are independent of the data dimensionality (i.e. sequence length), thus supporting the empirical success of MDMs. We then investigate the gain obtained by using non-constant schedule sizes (i.e. varying the number of unmasked tokens during the generation process) and identify the optimal schedule as a function of a so-called information profile of the data distribution, thus allowing for a principled optimization of schedule sizes. We define methods directly as sampling algorithms and do not use classical derivations as time-reversed diffusion processes, leading us to simple and transparent proofs.

stat.ML

Measures of Dependence based on Wasserstein distances

Measuring dependence between random variables is a fundamental problem in Statistics, with applications across diverse fields. While classical measures such as Pearson's correlation have been widely used for over a century, they have notable limitations, particularly in capturing nonlinear relationships and extending to general metric spaces. In recent years, the theory of Optimal Transport and Wasserstein distances has provided new tools to define measures of dependence that generalize beyond Euclidean settings. This survey explores recent proposals, outlining two main approaches: one based on the distance between the joint distribution and the product of marginals, and another leveraging conditional distributions. We discuss key properties, including characterization of independence, normalization, invariances, robustness, sample, and computational complexity. Additionally, we propose an alternative perspective that measures deviation from maximal dependence rather than independence, leading to new insights and potential extensions. Our work highlights recent advances in the field and suggests directions for further research in the measurement of dependence using Optimal Transport.

math.ST

Measuring Partial Exchangeability with Reproducing Kernel Hilbert Spaces

In Bayesian multilevel models, the data are structured in interconnected groups, and their posteriors borrow information from one another due to prior dependence between latent parameters. However, little is known about the behaviour of the dependence a posteriori. In this work, we develop a general framework for measuring partial exchangeability for parametric and nonparametric models, both a priori and a posteriori. We define an index that detects exchangeability for common models, is invariant by reparametrization, can be estimated through samples, and, crucially, is well-suited for posteriors. We achieve these properties through the use of Reproducing Kernel Hilbert Spaces, which map any random probability to a random object on a Hilbert space. This leads to many convenient properties and tractable expressions, especially a priori and under mixing. We apply our general framework to i) investigate the dependence a posteriori for the hierarchical Dirichlet process, retrieving a parametric convergence rate under very mild assumptions on the data; ii) eliciting the dependence structure of a parametric model for a principled comparison with a nonparametric alternative.

math.ST

Entropy contraction of the Gibbs sampler under log-concavity

The Gibbs sampler (a.k.a. Glauber dynamics and heat-bath algorithm) is a popular Markov Chain Monte Carlo algorithm which iteratively samples from the conditional distributions of a probability measure $\pi$ of interest. Under the assumption that $\pi$ is strongly log-concave, we show that the random scan Gibbs sampler contracts in relative entropy and provide a sharp characterization of the associated contraction rate. Assuming that evaluating conditionals is cheap compared to evaluating the joint density, our results imply that the number of full evaluations of $\pi$ needed for the Gibbs sampler to mix grows linearly with the condition number and is independent of the dimension. If $\pi$ is non-strongly log-concave, the convergence rate in entropy degrades from exponential to polynomial. Our techniques are versatile and extend to Metropolis-within-Gibbs schemes and the Hit-and-Run algorithm. A comparison with gradient-based schemes and the connection with the optimization literature are also discussed.

math.PR

Convergence rate of random scan Coordinate Ascent Variational Inference under log-concavity

The Coordinate Ascent Variational Inference scheme is a popular algorithm used to compute the mean-field approximation of a probability distribution of interest. We analyze its random scan version, under log-concavity assumptions on the target density. Our approach builds on the recent work of M. Arnese and D. Lacker, \emph{Convergence of coordinate ascent variational inference for log-concave measures via optimal transport} [arXiv:2404.08792] which studies the deterministic scan version of the algorithm, phrasing it as a block-coordinate descent algorithm in the space of probability distributions endowed with the geometry of optimal transport. We obtain tight rates for the random scan version, which imply that the total number of factor updates required to converge scales linearly with the condition number and the number of blocks of the target distribution. By contrast, available bounds for the deterministic scan case scale quadratically in the same quantities, which is analogue to what happens for optimization of convex functions in Euclidean spaces.

stat.ML

The Riemannian geometry of Sinkhorn divergences

We propose a new metric between probability measures on a compact metric space that mirrors the Riemannian manifold-like structure of quadratic optimal transport but includes entropic regularization. Its metric tensor is given by the Hessian of the Sinkhorn divergence, a debiased variant of entropic optimal transport. We precisely identify the tangent space it induces, which turns out to be related to a Reproducing Kernel Hilbert Space (RKHS). As usual in Riemannian geometry, the distance is built by looking for shortest paths. We prove that our distance is geodesic, metrizes the weak-star topology, and is equivalent to a RKHS norm. Still it retains the geometric flavor of optimal transport: as a paradigmatic example, translations are geodesics for the quadratic cost on $\mathbb{R}^d$. We also show two negative results on the Sinkhorn divergence that may be of independent interest: that it is not jointly convex, and that its square root is not a distance because it fails to satisfy the triangle inequality.

math.OC

Hierarchical Integral Probability Metrics: A distance on random probability measures with low sample complexity

Random probabilities are a key component to many nonparametric methods in Statistics and Machine Learning. To quantify comparisons between different laws of random probabilities several works are starting to use the elegant Wasserstein over Wasserstein distance. In this paper we prove that the infinite dimensionality of the space of probabilities drastically deteriorates its sample complexity, which is slower than any polynomial rate in the sample size. We propose a new distance that preserves many desirable properties of the former while achieving a parametric rate of convergence. In particular, our distance 1) metrizes weak convergence; 2) can be estimated numerically through samples with low complexity; 3) can be bounded analytically from above and below. The main ingredient are integral probability metrics, which lead to the name hierarchical IPM.

math.ST

Quantitative convergence of a discretization of dynamic optimal transport using the dual formulation

We present a discretization of the dynamic optimal transport problem for which we can obtain the convergence rate for the value of the transport cost to its continuous value when the temporal and spatial stepsize vanish. This convergence result does not require any regularity assumption on the measures, though experiments suggest that the rate is not sharp. Via an analysis of the duality gap we also obtain the convergence rates for the gradient of the optimal potentials and the velocity field under mild regularity assumptions. To obtain such rates, we discretize the dual formulation of the dynamic optimal transport problem and use the mature literature related to the error due to discretizing the Hamilton-Jacobi equation.

math.NA

Lifting functionals defined on maps to measure-valued maps via optimal transport

How can one lift a functional defined on maps from a space X to a space Y into a functional defined on maps from X into P(Y) the space of probability distributions over Y? Looking at measure-valued maps can be interpreted as knowing a classical map with uncertainty, and from an optimization point of view the main gain is the convexification of Y into P(Y). We will explain why trying to single out the largest convex lifting amounts to solve an optimal transport problem with an infinity of marginals which can be interesting by itself. Moreover we will show that, to recover previously proposed liftings for functionals depending on the Jacobian of the map, one needs to add a restriction of additivity to the lifted functional.

math.OC

Merging Rate of Opinions via Optimal Transport on Random Measures

Random measures provide flexible parameters for Bayesian nonparametric models. Given two different priors for a random measure, we develop a natural framework to investigate the rate at which the corresponding posteriors merge, as the sample size increases. We define a new distance between the laws of random measures that is built as a Wasserstein distance on the ground space of unbalanced measures, endowed with the bounded Lipschitz metric. We develop tight analytical bounds for its specification to completely random measures, including the special case of Poisson and gamma random measures. The bounds are interpreted in terms of an adapted extended Wasserstein distance between the L\'evy measures and are used to investigate the merging between the posteriors of normalized gamma and generalized gamma priors. After a careful study on the identifiability of the law of the random measure, interesting asymptotic and finite-sample insights are derived without putting any assumption on the true data generating process.

math.ST

Regularized unbalanced optimal transport as entropy minimization with respect to branching Brownian motion

We consider the problem of minimizing the entropy of a law with respect to the law of a reference branching Brownian motion under density constraints at an initial and final time. We call this problem the branching Schr\"odinger problem by analogy with the Schr\"odinger problem, where the reference process is a Brownian motion. Whereas the Schr\"odinger problem is related to regularized (a.k.a. entropic) optimal transport, we investigate here the link of the branching Schr\"odinger problem with regularized unbalanced optimal transport. This link is shown at two levels. First, relying on duality arguments, the values of these two problems of calculus of variations are linked, in the sense that the value of the regularized unbalanced optimal transport (seen as a function of the initial and final measure) is the lower semi-continuous relaxation of the value of the branching Schr\"odinger problem. Second, we also explicit a correspondence between the competitors of these two problems, and to that end we provide a fine description of laws having a finite entropy with respect to a reference branching Brownian motion. We investigate the small noise limit, when the noise intensity of the branching Brownian motion goes to $0$: in this case we show, at the level of the optimal transport model, that there is convergence to partial optimal transport. We also provide formal arguments about why looking at the branching Brownian motion, and not at other measure-valued branching Markov processes, like superprocesses, yields the problem closest to optimal transport. Finally, we explain how this problem can be solved numerically: the dynamical formulation of regularized unbalanced optimal transport can be discretized and solved via convex optimization.

math.PR

A Wasserstein index of dependence for random measures

Optimal transport and Wasserstein distances are flourishing in many scientific fields as a means for comparing and connecting random structures. Here we pioneer the use of an optimal transport distance between L\'{e}vy measures to solve a statistical problem. Dependent Bayesian nonparametric models provide flexible inference on distinct, yet related, groups of observations. Each component of a vector of random measures models a group of exchangeable observations, while their dependence regulates the borrowing of information across groups. We derive the first statistical index of dependence in $[0,1]$ for (completely) random measures that accounts for their whole infinite-dimensional distribution, which is assumed to be equal across different groups. This is accomplished by using the geometric properties of the Wasserstein distance to solve a max-min problem at the level of the underlying L\'{e}vy measures. The Wasserstein index of dependence sheds light on the models' deep structure and has desirable properties: (i) it is $0$ if and only if the random measures are independent; (ii) it is $1$ if and only if the random measures are completely dependent; (iii) it simultaneously quantifies the dependence of $d \ge 2$ random measures, avoiding the need for pairwise comparisons; (iv) it can be evaluated numerically. Moreover, the index allows for informed prior specifications and fair model comparisons for Bayesian nonparametric models.

math.ST

Convex functions defined on metric spaces are pulled back to subharmonic ones by harmonic maps

If $u : \Omega\subset \mathbb{R}^d \to {\rm X}$ is a harmonic map valued in a metric space ${\rm X}$ and ${\sf E} : {\rm X} \to \mathbb{R}$ is a convex function, in the sense that it generates an ${\rm EVI}_0$-gradient flow, we prove that the pullback ${\sf E} \circ u : \Omega \to \mathbb{R}$ is subharmonic. This property was known in the smooth Riemannian manifold setting or with curvature restrictions on ${\rm X}$, while we prove it here in full generality. In addition, we establish generalized maximum principles, in the sense that the $L^q$ norm of ${\sf E} \circ u$ on $\partial \Omega$ controls the $L^p$ norm of ${\sf E} \circ u$ in $\Omega$ for some well-chosen exponents $p \geq q$, including the case $p=q=+\infty$. In particular, our results apply when ${\sf E}$ is a geodesically convex entropy over the Wasserstein space, and thus settle some conjectures of Y. Brenier, "Extended Monge-Kantorovich theory" in Optimal transportation and applications (Martina Franca, 2001), volume 1813 of Lecture Notes in Math., pages 91-121. Springer, Berlin, 2003.

math.MG

Towards a mathematical theory of trajectory inference

We devise a theoretical framework and a numerical method to infer trajectories of a stochastic process from samples of its temporal marginals. This problem arises in the analysis of single cell RNA-sequencing data, which provide high dimensional measurements of cell states but cannot track the trajectories of the cells over time. We prove that for a class of stochastic processes it is possible to recover the ground truth trajectories from limited samples of the temporal marginals at each time-point, and provide an efficient algorithm to do so in practice. The method we develop, Global Waddington-OT (gWOT), boils down to a smooth convex optimization problem posed globally over all time-points involving entropy-regularized optimal transport. We demonstrate that this problem can be solved efficiently in practice and yields good reconstructions, as we show on several synthetic and real datasets.

stat.ML

Hidden convexity in a problem of nonlinear elasticity

We study compressible and incompressible nonlinear elasticity variational problems in a general context. Our main result gives a sufficient condition for an equilibrium to be a global energy minimizer, in terms of convexity properties of the pressure in the deformed configuration. We also provide a convex relaxation of the problem together with its dual formulation, based on measure-valued mappings, which coincides with the original problem under our condition.

math.AP

Unconditional convergence for discretizations of dynamical optimal transport

The dynamical formulation of optimal transport, also known as Benamou-Brenier formulation or Computational Fluid Dynamics formulation, amounts to write the optimal transport problem as the optimization of a convex functional under a PDE constraint, and can handle \emph{a priori} a vast class of cost functions and geometries. Several discretizations of this problem have been proposed, leading to computations on flat spaces as well as Riemannian manifolds, with extensions to mean field games and gradient flows in the Wasserstein space. In this article, we provide a framework which guarantees convergence under mesh refinement of the solutions of the space-time discretized problems to the one of the infinite-dimensional one for quadratic optimal transport. The convergence holds without condition on the ratio between spatial and temporal step sizes, and can handle arbitrary positive measures as input, while the underlying space can be a Riemannian manifold. Both the finite volume discretization proposed by Gladbach, Kopfer and Maas, as well as the discretization over triangulations of surfaces studied by the present author in collaboration with Claici, Chien and Solomon fit in this framework.

math.NA