SearcharxivSearch

arXiv subjects

Vincent Divol

Publications and source records attributed to Vincent Divol.

16 recordsLinked to original sources

Spectral stability of empirical metric-measure Laplacians

The variance of nonparametric estimators is typically insensitive to the regularity of the object being estimated. We establish such a property for the spectra of graph Laplacian matrices at a fixed bandwidth $h>0$. Specifically, given $n$ i.i.d. samples from a probability measure $\mu$ on a Polish metric space, we compare the eigenvalues of the empirical weighted Laplacian operator $\Delta_{\mu_n}^h$ to those of the population counterpart $\Delta_\mu^h$ under a spectral gap condition, bounding the relative error by $1/\sqrt{nv_\mu(h)}$ for eigenvalues of order smaller than $h^{-2}$, where $v_\mu(h)$ is the smallest mass of a ball of radius $h$. This bound requires very weak regularity conditions on $\mu$: it is satisfied if $\mu$ belongs to the class of coarse PI measures that we introduce. This class contains measures on metric graphs, spaces with sufficiently regular boundaries, corners, or branch points, together with discretizations or thickenings of these at scale $O(h)$. Even for measures having densities of regularity $s>2$ on manifolds (the only known case so far), our bound improves on the state-of-the-art by shaving off logarithmic factors.

math.ST

Minimax spectral estimation of weighted Laplace operators

Given $n$ i.i.d. observations, we study the problem of estimating the spectrum of weighted Laplace operators of the form $\Delta_f=\Delta + \alpha \nabla \log f\cdot \nabla$, where $f$ is a positive probability density on a known compact $d$-dimensional manifold without boundary and $\alpha\in \mathbb{R}$ is a hyperparameter. These operators arise as continuum limits of graph Laplacian matrices and provide valuable geometric information on the underlying data distribution. We establish the exact minimax rates of estimation for this problem, by exhibiting two different rates of convergence for eigenfunctions and eigenvalues. When $f$ belongs to a H\"older-Zygmund class $\mathscr{C}^s$ of regularity $s\geqslant 2$, the eigenfunctions can be estimated with respect to the $\mathrm{L}^q$-norm ($q\geqslant 1$) via plug-in methods at the minimax rate $n^{-\frac{s+1}{2s+d}}$ for $d\geqslant 3$ (with different rates for $d\leqslant 2$). Moreover, eigenvalues can be estimated at the minimax rate $n^{-\frac{4s}{4s+d}}+n^{-\frac 12}$. In the regime $s>\frac d4$, we further show that asymptotically efficient estimators exist. We also present a general framework for estimating nonlinear functionals over H\"older-Zygmund spaces, with potential applications to a broad class of statistical problems.

math.ST

Measure estimation on a manifold explored by a diffusion process

From the observation of a diffusion path $(X_t)_{t\in [0,T]}$ on a compact connected $d$-dimensional manifold $\mathcal{M}$ without boundary, we consider the problem of estimating the stationary measure $\mu$ of the process. Wang and Zhu (2023) showed that for the Wasserstein metric $\mathcal{W}_2$ and for $d\geq 5$, the convergence rate of $T^{-1/(d-2)}$ is attained by the occupation measure of the path $(X_t)_{t\in [0,T]}$ when $(X_t)_{t\in [0,T]}$ is a Langevin diffusion. We extend their result in several directions. First, we show that the rate of convergence holds for a large class of diffusion paths, whose generators are uniformly elliptic. Second, the regularity of the density $p$ of the stationary measure $\mu$ with respect to the volume measure of $\mathcal{M}$ can be leveraged to obtain faster estimators: when $p$ belongs to a Sobolev space of order $\ell\geq 2$, smoothing the occupation measure by convolution with a kernel yields an estimator whose rate of convergence is of order $T^{-(\ell+1)/(2\ell+d-2)}$. We further show that this rate is the minimax rate of estimation for this problem.

math.ST

Demographic parity in regression and classification within the unawareness framework

This paper explores the theoretical foundations of fair regression under the constraint of demographic parity within the unawareness framework, where disparate treatment is prohibited, extending existing results where such treatment is permitted. Specifically, we aim to characterize the optimal fair regression function when minimizing the quadratic loss. Our results reveal that this function is given by the solution to a barycenter problem with optimal transport costs. Additionally, we study the connection between optimal fair cost-sensitive classification, and optimal fair regression. We demonstrate that nestedness of the decision sets of the classifiers is both necessary and sufficient to establish a form of equivalence between classification and regression. Under this nestedness assumption, the optimal classifiers can be derived by applying thresholds to the optimal fair regression function; conversely, the optimal fair regression function is characterized by the family of cost-sensitive classifiers.

stat.ML

Wasserstein convergence of \v{C}ech persistence diagrams for samplings of submanifolds

\v{C}ech Persistence diagrams (PDs) are topological descriptors routinely used to capture the geometry of complex datasets. They are commonly compared using the Wasserstein distances $OT_{p}$; however, the extent to which PDs are stable with respect to these metrics remains poorly understood. We partially close this gap by focusing on the case where datasets are sampled on an $m$-dimensional submanifold of $\mathbb{R}^{d}$. Under this manifold hypothesis, we show that convergence with respect to the $OT_{p}$ metric happens exactly when $p\gt m$. We also provide improvements upon the bottleneck stability theorem in this case and prove new laws of large numbers for the total $\alpha$-persistence of PDs. Finally, we show how these theoretical findings shed new light on the behavior of the feature maps on the space of PDs that are used in ML-oriented applications of Topological Data Analysis.

cs.CG

Tight stability bounds for entropic Brenier maps

Entropic Brenier maps are regularized analogues of Brenier maps (optimal transport maps) which converge to Brenier maps as the regularization parameter shrinks. In this work, we prove quantitative stability bounds between entropic Brenier maps under variations of the target measure. In particular, when all measures have bounded support, we establish the optimal Lipschitz constant for the mapping from probability measures to entropic Brenier maps. This provides an exponential improvement to a result of Carlier, Chizat, and Laborde (2024). As an application, we prove near-optimal bounds for the stability of semi-discrete \emph{unregularized} Brenier maps for a family of discrete target measures.

math.PR

Critical points of the distance function to a generic submanifold

In general, the critical points of the distance function $d_{\mathsf{M}}$ to a compact submanifold $\mathsf{M} \subset \mathbb{R}^D$ can be poorly behaved. In this article, we show that this is generically not the case by listing regularity conditions on the critical and $\mu$-critical points of a submanifold and by proving that they are generically satisfied and stable with respect to small $C^2$ perturbations. More specifically, for any compact abstract manifold $M$, the set of embeddings $i:M\rightarrow \mathbb{R}^D$ such that the submanifold $i(M)$ satisfies those conditions is open and dense in the Whitney $C^2$-topology. When those regularity conditions are fulfilled, we prove that the distance function to $i(M)$ satisfies Morse-like conditions and that the critical points of the distance function to an $\varepsilon$-dense subset of the submanifold (e.g., obtained via some sampling process) are well-behaved. We also provide many examples that showcase how the absence of these conditions allows for pathological situations.

math.DG

Minimax estimation of discontinuous optimal transport maps: The semi-discrete case

We consider the problem of estimating the optimal transport map between two probability distributions, $P$ and $Q$ in $\mathbb R^d$, on the basis of i.i.d. samples. All existing statistical analyses of this problem require the assumption that the transport map is Lipschitz, a strong requirement that, in particular, excludes any examples where the transport map is discontinuous. As a first step towards developing estimation procedures for discontinuous maps, we consider the important special case where the data distribution $Q$ is a discrete measure supported on a finite number of points in $\mathbb R^d$. We study a computationally efficient estimator initially proposed by Pooladian and Niles-Weed (2021), based on entropic optimal transport, and show in the semi-discrete setting that it converges at the minimax-optimal rate $n^{-1/2}$, independent of dimension. Other standard map estimation techniques both lack finite-sample guarantees in this setting and provably suffer from the curse of dimensionality. We confirm these results in numerical experiments, and provide experiments for other settings, not covered by our theory, which indicate that the entropic estimator is a promising methodology for other discontinuous transport map estimation problems.

math.ST

Optimal transport map estimation in general function spaces

We study the problem of estimating a function $T$ given independent samples from a distribution $P$ and from the pushforward distribution $T_\sharp P$. This setting is motivated by applications in the sciences, where $T$ represents the evolution of a physical system over time, and in machine learning, where, for example, $T$ may represent a transformation learned by a deep neural network trained for a generative modeling task. To ensure identifiability, we assume that $T = \nabla \varphi_0$ is the gradient of a convex function, in which case $T$ is known as an \emph{optimal transport map}. Prior work has studied the estimation of $T$ under the assumption that it lies in a H\"older class, but general theory is lacking. We present a unified methodology for obtaining rates of estimation of optimal transport maps in general function spaces. Our assumptions are significantly weaker than those appearing in the literature: we require only that the source measure $P$ satisfy a Poincar\'e inequality and that the optimal map be the gradient of a smooth convex function that lies in a space whose metric entropy can be controlled. As a special case, we recover known estimation rates for H\"older transport maps, but also obtain nearly sharp results in many settings not covered by prior work. For example, we provide the first statistical rates of estimation when $P$ is the normal distribution and the transport map is given by an infinite-width shallow neural network.

math.ST

Estimation and Quantization of Expected Persistence Diagrams

Persistence diagrams (PDs) are the most common descriptors used to encode the topology of structured data appearing in challenging learning tasks; think e.g. of graphs, time series or point clouds sampled close to a manifold. Given random objects and the corresponding distribution of PDs, one may want to build a statistical summary-such as a mean-of these random PDs, which is however not a trivial task as the natural geometry of the space of PDs is not linear. In this article, we study two such summaries, the Expected Persistence Diagram (EPD), and its quantization. The EPD is a measure supported on R 2 , which may be approximated by its empirical counterpart. We prove that this estimator is optimal from a minimax standpoint on a large class of models with a parametric rate of convergence. The empirical EPD is simple and efficient to compute, but possibly has a very large support, hindering its use in practice. To overcome this issue, we propose an algorithm to compute a quantization of the empirical EPD, a measure with small support which is shown to approximate with near-optimal rates a quantization of the theoretical EPD.

math.ST

Measure estimation on manifolds: an optimal transport approach

Assume that we observe i.i.d.~points lying close to some unknown $d$-dimensional $\mathcal{C}^k$ submanifold $M$ in a possibly high-dimensional space. We study the problem of reconstructing the probability distribution generating the sample. After remarking that this problem is degenerate for a large class of standard losses ($L_p$, Hellinger, total variation, etc.), we focus on the Wasserstein loss, for which we build an estimator, based on kernel density estimation, whose rate of convergence depends on $d$ and the regularity $s\leq k-1$ of the underlying density, but not on the ambient dimension. In particular, we show that the estimator is minimax and matches previous rates in the literature in the case where the manifold $M$ is a $d$-dimensional cube. The related problem of the estimation of the volume measure of $M$ for the Wasserstein loss is also considered, for which a minimax estimator is exhibited.

math.ST

Minimax adaptive estimation in manifold inference

We focus on the problem of manifold estimation: given a set of observations sampled close to some unknown submanifold $M$, one wants to recover information about the geometry of $M$. Minimax estimators which have been proposed so far all depend crucially on the a priori knowledge of some parameters quantifying the underlying distribution generating the sample (such as bounds on its density), whereas those quantities will be unknown in practice. Our contribution to the matter is twofold: first, we introduce a one-parameter family of manifold estimators $(\hat{M}_t)_{t\geq 0}$ based on a localized version of convex hulls, and show that for some choice of $t$, the corresponding estimator is minimax on the class of models of $C^2$ manifolds introduced in [Genovese et al., Manifold estimation and singular deconvolution under Hausdorff loss]. Second, we propose a completely data-driven selection procedure for the parameter $t$, leading to a minimax adaptive manifold estimator on this class of models. This selection procedure actually allows us to recover the Hausdorff distance between the set of observations and $M$, and can therefore be used as a scale parameter in other settings, such as tangent space estimation.

math.ST

Understanding the Topology and the Geometry of the Space of Persistence Diagrams via Optimal Partial Transport

Despite the obvious similarities between the metrics used in topological data analysis and those of optimal transport, an optimal-transport based formalism to study persistence diagrams and similar topological descriptors has yet to come. In this article, by considering the space of persistence diagrams as a space of discrete measures, and by observing that its metrics can be expressed as optimal partial transport problems, we introduce a generalization of persistence diagrams, namely Radon measures supported on the upper half plane. Such measures naturally appear in topological data analysis when considering continuous representations of persistence diagrams (e.g.\ persistence surfaces) but also as limits for laws of large numbers on persistence diagrams or as expectations of probability distributions on the persistence diagrams space. We explore topological properties of this new space, which will also hold for the closed subspace of persistence diagrams. New results include a characterization of convergence with respect to Wasserstein metrics, a geometric description of barycenters (Fr\'echet means) for any distribution of diagrams, and an exhaustive description of continuous linear representations of persistence diagrams. We also showcase the strength of this framework to study random persistence diagrams by providing several statistical results made meaningful thanks to this new formalism.

cs.CG

On the choice of weight functions for linear representations of persistence diagrams

Persistence diagrams are efficient descriptors of the topology of a point cloud. As they do not naturally belong to a Hilbert space, standard statistical methods cannot be directly applied to them. Instead, feature maps (or representations) are commonly used for the analysis. A large class of feature maps, which we call linear, depends on some weight functions, the choice of which is a critical issue. An important criterion to choose a weight function is to ensure stability of the feature maps with respect to Wasserstein distances on diagrams. We improve known results on the stability of such maps, and extend it to general weight functions. We also address the choice of the weight function by considering an asymptotic setting; assume that $\mathbb{X}_n$ is an i.i.d. sample from a density on $[0,1]^d$. For the \v{C}ech and Rips filtrations, we characterize the weight functions for which the corresponding feature maps converge as $n$ approaches infinity, and by doing so, we prove laws of large numbers for the total persistences of such diagrams. Those two approaches (stability and convergence) lead to the same simple heuristic for tuning weight functions: if the data lies near a $d$-dimensional manifold, then a sensible choice of weight function is the persistence to the power $\alpha$ with $\alpha \geq d$.

math.PR

The density of expected persistence diagrams and its kernel based estimation

Persistence diagrams play a fundamental role in Topological Data Analysis where they are used as topological descriptors of filtrations built on top of data. They consist in discrete multisets of points in the plane $\mathbb{R}^2$ that can equivalently be seen as discrete measures in $\mathbb{R}^2$. When the data come as a random point cloud, these discrete measures become random measures whose expectation is studied in this paper. First, we show that for a wide class of filtrations, including the \v{C}ech and Rips-Vietoris filtrations, the expected persistence diagram, that is a deterministic measure on $\mathbb{R}^2$ , has a density with respect to the Lebesgue measure. Second, building on the previous result we show that the persistence surface recently introduced in [Adams & al., Persistence images: a stable vector representation of persistent homology] can be seen as a kernel estimator of this density. We propose a cross-validation scheme for selecting an optimal bandwidth, which is proven to be a consistent procedure to estimate the density.

cs.CG