SearcharxivSearch

arXiv subjects

Gilles Mordant

Publications and source records attributed to Gilles Mordant.

At least 19 recordsLinked to original sources

Fluctuations of the quadratic matching cost on the flat torus: dimensions two, three and four

We establish sharp variance asymptotics and limiting distributions for the quadratic optimal matching cost between the uniform measure and an empirical measure counterpart on the flat torus $\mathbb T^d$ in dimensions $d=2,3,4$. More precisely, let $μ$ be Haar probability measure and $μ_n$ the empirical measure of $n$ independent samples drawn from $μ$. Then, for $d=2,3$, $$ \begin{aligned} n\bigl(W_2^2(μ_n,μ)-\mathbb E[W_2^2(μ_n,μ)]\bigr) &\xrightarrow{\mathrm d} L_d,\\ n^2\operatorname{Var}(W_2^2(μ_n,μ)) &\longrightarrow \frac{1}{8π^4} \sum_{k\in\mathbb Z^d\setminus\{0\}}\frac{1}{|k|^4}, \end{aligned} $$ where $L_d$ is an explicit non-Gaussian weighted sum of independent centered exponential random variables. In dimension four, $$ \frac{n}{\sqrt{\log n}} \bigl(W_2^2(μ_n,μ)-\mathbb E[W_2^2(μ_n,μ)]\bigr) \xrightarrow{\mathrm d} N\left(0,\frac{1}{16π^2}\right). $$ The main idea of the proof is to rely on a Hoeffding decomposition of the optimal transport cost, which turns out to be asymptotically equivalent to U-statistics of order 2.

math.PR

Convergence of the Sinkhorn Riemannian metric for finitely supported measures

The Sinkhorn Riemannian metric characterizes the local behavior of the Sinkhorn divergence, a debiased version of entropy-regularized optimal transport. Although the Sinkhorn divergence is known to approximate the squared Wasserstein-2 distance as the regularization parameter $\varepsilon$ goes to $0$, the behavior of the Sinkhorn Riemannian metric as $\varepsilon\to 0$ remains largely open. To the best of our knowledge, the only existing argument is for measures with a density and is still a formal computation. In this work, we prove the convergence of the Sinkhorn Riemannian metric as $\varepsilon\to 0$ in the case of finitely supported measures undergoing horizontal perturbations. Under our assumptions, our techniques additionally allow us to prove the convergence of higher order derivatives of the Sinkhorn Riemannian metric, thereby ensuring the convergence of the associated Riemann curvature tensor and Christoffel symbols as $\varepsilon\to 0$. We also compare with the absolutely continuous setting and show where our approach fails for measures with a density.

math.OC

On the discretization of the object space in inverse problems with application to cryo-electron microscopy

In many inverse problems, the aim is to recover a probability distribution on a latent (or object state) space from indirect, noisy observations. When the observations can be modelled as noisy samples from the mixing latent distribution, the recovery problem is a deconvolution problem on the space of probability measures. A common strategy is to fix a finite set of candidate points and estimate a weight for each, turning the problem into a finite-dimensional concave maximum likelihood problem on the simplex. We study the combined effect of this discretization and of the noise in the context of cryo-electron microscopy (cryo-EM), where the candidate points are biomolecular conformations and the weights describe the relative frequency of each conformation. Our results pertain to both statistical and algorithmic aspects of the estimator. We analyze the weight-recovery problem in which the candidate states and their likelihoods are known. A nearby pair of candidates forces a near-null direction. More generally, the grid and the noise level impose a uniform lower bound on the achievable Kullback--Leibler divergence between observation densities, even with infinite data. The finite-grid estimator is asymptotically normal when its population target is in the interior of the simplex; at a boundary target, its limit is a cone-projected Gaussian. Finally, the exact proximal form of Expectation--Maximization leads to a global high-noise comparison between an early iterate and a KL-penalized likelihood, without a basin assumption or a linearization of the recursion. Tests on synthetic images of the Hsp90 molecule illustrate the theoretical findings and translate them into practical guidelines for interpreting reweighted ensembles.

stat.ME

A central limit theorem for the random assignment problem

Let \(C_n\) be the minimum cost of a perfect matching in an \(n\times n\) matrix of independent uniform random variables. We prove that \[ \sqrt n\{C_n-ζ(2)\} \ \Longrightarrow\ \mathcal N\bigl(0,4ζ(2)-4ζ(3)\bigr). \] The proof begins with an exact change of variables based on a uniformly rooted shortest-path selection of an optimal dual potential. After the unused reduced costs are integrated out, a reference law separates the rows conditionally on the potential field, while the ordered potential gaps become independent exponentials. The only residual dependence is a directed-tree factor. Ordering the potentials turns its zero--one support into a Ferrers matrix, whose matrix-tree determinant is triangular. A singular inverse-degree estimate and exact normalization then yield total-variation convergence to the reference law. Finally, a conditional triangular-array central limit theorem accounts for row noise, and a second triangular array accounts for the linear response of the potential field. The strategy used here is likely to be applicable to other problems.

math.PR

Empirical optimal transport potentials: fast rates and a functional central limit theorem

Optimal transport potentials are fundamental objects in statistics, economics, and machine learning: their gradients generate optimal transport maps, while the potentials themselves act as location-dependent dual prices and sensitivity variables. We study the estimation of the quadratic optimal transport potential when a fixed absolutely continuous reference distribution $μ$ is transported to an unknown distribution $ν$, accessed to via its empirical measure. Our main ingredient is a stability inequality that controls the $L^1(μ)$ distance, modulo additive constants, between a strongly convex potential $φ$ and a convex potential $\widetildeφ$ by a weak dual norm of $(\nabla\widetildeφ)_\#μ-(\nablaφ)_\#μ,$ together with a second-order Wasserstein remainder of logarithmic type. This separation between the leading empirical-process term and the Wasserstein remainder yields faster convergence for potentials than for the corresponding transport maps. Under smoothness and uniform convexity assumptions, the exact semidiscrete Brenier potential converges in $L^1(μ)$ at rate $n^{-1/2}$ for $d\leq3$, at rate $n^{-1/2}(\log n)^{5/2}$ for $d=4$, and at rate $n^{-2/d}(\log n)^{(d+2)/d}$ for $d\geq5$. The polynomial exponents are sharp. In dimensions $d\leq3$, we further establish a nondegenerate function-space central limit theorem and prove consistency of the nonparametric bootstrap. These results yield joint root-$n$ inference for every fixed finite collection of normalization-invariant weighted contrasts of the potential, including regional shadow premia in reference-based risk problems. Finally, we prove matching upper and lower bounds of order $\varepsilon\log(1/\varepsilon)$ for the normalization-invariant sum of the entropic dual potentials.

math.ST

Quantifying the noise sensitivity of the Wasserstein metric for images

Wasserstein metrics are increasingly adopted as similarity scores for images. We consider the sensitivity of Wasserstein metrics with respect to pixel-wise additive noise when the images are treated as discrete measures on the pixel grid. We derive finite-sample expectation bounds for a Gaussian noise model. Among other results, we prove that the error in the signed 2-Wasserstein discrepancy scales with the square root of the noise standard deviation. This is favorable compared to the Euclidean metric that scales linearly, and thus provides a theoretical basis for the benefits of optimal transport distances in noisy settings. We present experiments that support our theoretical findings and point to a peculiar phenomenon where increasing the level of noise can decrease the Wasserstein distance. A case study on cryo-electron microscopy images demonstrates that the Wasserstein metric can capture the geometry of the data manifold in high noise settings even when the Euclidean metric fails.

math.ST

The entropic optimal (self-)transport problem: Limit distributions for decreasing regularization with application to score function estimation

We study the statistical properties of the entropic optimal (self) transport problem for smooth probability measures. We provide an accurate description of the limit distribution for entropic (self-)potentials and plans as the regularization parameter shrinks with the sample size; this regime is largely unexplored in the prior statistical literature, where $ε$ is typically held fixed. Additionally, we show that a rescaling of the barycentric projection of the empirical entropic optimal self-transport plans converges to the score function, a central object for diffusion models, and characterize the asymptotic fluctuations both pointwise and in $L^2$. Finally, we describe under what conditions the methods used enable to derive (pointwise) limiting distribution results for the empirical entropic optimal transport potentials in the case of two different measures and appropriately chosen shrinking regularization parameter. This endeavour requires a better understanding of the composition of Sinkhorn operators in the small $\eps$-limit, a result of independent interest.

math.ST

The Catastrophic Failure of The k-Means Algorithm in High Dimensions, and How Hartigan's Algorithm Avoids It

Lloyd's k-means algorithm is one of the most widely used clustering methods. We prove that in high-dimensional, high-noise settings, the algorithm exhibits catastrophic failure: with high probability, essentially every partition of the data is a fixed point. Consequently, Lloyd's algorithm simply returns its initial partition - even when the underlying clusters are trivially recoverable by other methods. In contrast, we prove that Hartigan's k-means algorithm does not exhibit this pathology. Our results show the stark difference between these algorithms and offer a theoretical explanation for the empirical difficulties often observed with k-means in high dimensions.

stat.ML

p-Wasserstein distances on networks and 3D to 1D convergence

We study transport distances on metric graphs representing gas networks. Starting from the dynamic formulation of the Wasserstein distance, we review extensions to networks, with and without the possibility of storing mass on the vertices. Next, we examine the asymptotic behavior of the static Wasserstein distance on a three-dimensional network domain that converges to a metric graph. We show convergence of the distance with a proof that is based on the characterization of optimal transport plans as $c$-cyclically monotone sets. We conclude by illustrating our finding with several numerical examples.

math.AP

The Riemannian geometry of Sinkhorn divergences

We propose a new metric between probability measures on a compact metric space that mirrors the Riemannian manifold-like structure of quadratic optimal transport but includes entropic regularization. Its metric tensor is given by the Hessian of the Sinkhorn divergence, a debiased variant of entropic optimal transport. We precisely identify the tangent space it induces, which turns out to be related to a Reproducing Kernel Hilbert Space (RKHS). As usual in Riemannian geometry, the distance is built by looking for shortest paths. We prove that our distance is geodesic, metrizes the weak-star topology, and is equivalent to a RKHS norm. Still it retains the geometric flavor of optimal transport: as a paradigmatic example, translations are geodesics for the quadratic cost on $\mathbb{R}^d$. We also show two negative results on the Sinkhorn divergence that may be of independent interest: that it is not jointly convex, and that its square root is not a distance because it fails to satisfy the triangle inequality.

math.OC

Estimation of Algebraic Sets: Extending PCA Beyond Linearity

An algebraic set is defined as the zero locus of a system of real polynomial equations. In this paper we address the problem of recovering an unknown algebraic set $\mathcal{A}$ from noisy observations of latent points lying on $\mathcal{A}$ -- a task that extends principal component analysis, which corresponds to the purely linear case. Our procedure consists of three steps: (i) constructing the {\it moment matrix} from the Vandermonde matrix associated with the data set and the degree of the fitted polynomials, (ii) debiasing this moment matrix to remove the noise-induced bias, (iii) extracting its kernel via an eigenvalue decomposition of the debiased moment matrix. These steps yield $n^{-1/2}$-consistent estimators of the coefficients of a set of generators for the ideal of polynomials vanishing on $\mathcal{A}$. To reconstruct $\mathcal{A}$ itself, we propose three complementary strategies: (a) compute the zero set of the fitted polynomials; (b) build a semi-algebraic approximation that encloses $\mathcal{A}$; (c) when structural prior information is available, project the estimated coefficients onto the corresponding constrained space. We prove (nearly) parametric asymptotic error bounds and show that each approach recovers $\mathcal{A}$ under mild regularity conditions.

math.ST

Manifold Learning with Sparse Regularised Optimal Transport

Manifold learning is a central task in modern statistics and data science. Many datasets (cells, documents, images, molecules) can be represented as point clouds embedded in a high dimensional ambient space, however the degrees of freedom intrinsic to the data are usually far fewer than the number of ambient dimensions. The task of detecting a latent manifold along which the data are embedded is a prerequisite for a wide family of downstream analyses. Real-world datasets are subject to noisy observations and sampling, so that distilling information about the underlying manifold is a major challenge. We propose a method for manifold learning that utilises a symmetric version of optimal transport with a quadratic regularisation that constructs a sparse and adaptive affinity matrix, that can be interpreted as a generalisation of the bistochastic kernel normalisation. We prove that the resulting kernel is consistent with a Laplace-type operator in the continuous limit, establish robustness to heteroskedastic noise and exhibit these results in numerical experiments. We identify a highly efficient computational scheme for computing this optimal transport for discrete data and demonstrate that it outperforms competing methods in a set of examples.

stat.ML

Infinitesimal behavior of Quadratically Regularized Optimal Transport and its relation with the Porous Medium Equation

The quadratically regularized optimal transport problem has recently been considered in various applications where the coupling needs to be \emph{sparse}, i.e., the density of the coupling needs to be zero for a large subset of the product of the supports of the marginals. However, unlike the acclaimed entropy-regularized optimal transport, the effect of quadratic regularization on the transport problem is not well understood from a mathematical standpoint. In this work, we take a first step towards its understanding. We prove that the difference between the cost of optimal transport and its regularized version multiplied by the ratio $\varepsilon^{-\frac{2}{d+2}}$ converges to a nontrivial limit as the regularization parameter $\varepsilon$ tends to 0. The proof confirms a conjecture from Zhang et al. (2023) where it is claimed that a modification of the self-similar solution of the porous medium equation, the Barenblatt--Pattle solution, can be used as an approximate solution of the regularized transport cost for small values of $\varepsilon$.

math.AP

Regularised optimal self-transport is approximate Gaussian mixture maximum likelihood

We investigate the link between regularised self-transport problems and maximum likelihood estimation in Gaussian mixture models (GMM). This link suggests that self-transport followed by a clustering technique leads to principled estimators at a reasonable computational cost. Also, robustness, sparsity and stability properties of the optimal transport plan arguably make the regularised self-transport a statistical tool of choice for the GMM.

math.ST

Empirical Optimal Transport under Estimated Costs: Distributional Limits and Statistical Applications

Optimal transport (OT) based data analysis is often faced with the issue that the underlying cost function is (partially) unknown. This paper is concerned with the derivation of distributional limits for the empirical OT value when the cost function and the measures are estimated from data. For statistical inference purposes, but also from the viewpoint of a stability analysis, understanding the fluctuation of such quantities is paramount. Our results find direct application in the problem of goodness-of-fit testing for group families, in machine learning applications where invariant transport costs arise, in the problem of estimating the distance between mixtures of distributions, and for the analysis of empirical sliced OT quantities. The established distributional limits assume either weak convergence of the cost process in uniform norm or that the cost is determined by an optimization problem of the OT value over a fixed parameter space. For the first setting we rely on careful lower and upper bounds for the OT value in terms of the measures and the cost in conjunction with a Skorokhod representation. The second setting is based on a functional delta method for the OT value process over the parameter space. The proof techniques might be of independent interest.

math.ST

Quadratically Regularized Optimal Transport: nearly optimal potentials and convergence of discrete Laplace operators

We consider the conjecture proposed in Matsumoto, Zhang and Schiebinger (2022) suggesting that optimal transport with quadratic regularisation can be used to construct a graph whose discrete Laplace operator converges to the Laplace--Beltrami operator. We derive first order optimal potentials for the problem under consideration and find that the resulting solutions exhibit a surprising resemblance to the well-known Barenblatt--Prattle solution of the porous medium equation. Then, relying on these first order optimal potentials, we derive the pointwise $L^2$-limit of such discrete operators built from an i.i.d. random sample on a smooth compact manifold. Simulation results complementing the limiting distribution results are also presented.

math.AP

Center-Outward Multiple-Output Lorenz Curves and Gini Indices a measure transportation approach

Based on measure transportation ideas and the related concepts of center-outward quantile functions, we propose multiple-output center-outward generalizations of the traditional univariate concepts of Lorenz and concentration functions, and the related Gini and Kakwani coefficients. These new concepts have a natural interpretation, either in terms of contributions of central ("middle-class") regions to the expectation of some variable of interest, or in terms of the physical notions of work and energy, which sheds new light on the nature of economic and social inequalities. Importantly, the proposed concepts pave the way to statistically sound definitions, based on multiple variables, of quantiles and quantile regions, and the concept of "middle class," of high relevance in various socio-economic contexts.

math.ST

Extrema of multinomial assignment process

We study the asymptotic behavior of the expectation of the maxima and minima of random assignment process generated by a large matrix with multinomial entries. A variety of results is obtained for different sparsity regimes.

math.PR