SearcharxivSearch

arXiv subjects

Eustasio del Barrio

Publications and source records attributed to Eustasio del Barrio.

At least 19 recordsLinked to original sources

Distributional Limit Theory for Optimal Transport

Optimal Transport (OT) is a resource allocation problem with applications in biology, data science, economics and statistics, among others. In some of the applications, practitioners have access to samples which approximate the continuous measure. Hence the quantities of interest derived from OT -- plans, maps and costs -- are only available in their empirical versions. Statistical inference on OT aims at finding confidence intervals of the population plans, maps and costs. In recent years this topic gained an increasing interest in the statistical community. In this paper we provide a comprehensive review of the most influential results on this research field, underlying the some of the applications. Finally, we provide a list of open problems.

math.ST

An Optimal Transportation Approach for Improved Confidence Intervals

Optimal transport methods have recently attracted a lot of attention in statistics. Their appeal lies in providing a geometric framework for comparing probability measures, leading to new perspectives on classical problems. A central problem in statistics is the construction of valid confidence sets as fundamental inferential tools in practice. A well-known problem is that for complex problems or relatively small samples, their asymptotic approximations often show poor performance. This suggests to apply optimal transport methods when constructing confidence sets for hard problems to improve their coverage properties. We introduce such a procedure, derive the theoretical framework studying consistency and error bounds for the coverage probability of the resulting intervals. To guarantee feasibility in practice, we propose data-driven choices for our hyper parameters. This approach extends classical quantile-based confidence intervals by leveraging optimal couplings to minimize coverage deviations. Simulations demonstrate striking performance in different estimation problems, outperforming standard methods in accuracy and robustness.

stat.ME

An improved central limit theorem for the empirical sliced Wasserstein distance

Wasserstein distances are widely used in modern data analysis but pose significant computational and statistical challenges in high dimensions. The sliced Wasserstein distance alleviates these challenges by leveraging one-dimensional projections. Building on the Efron-Stein inequality-a technique proven effective in related problems-and a non-trivial control of the optimal transport potentials across directions, we establish a central limit theorem for the p-sliced Wasserstein distance, for p>1, centered at the expected empirical cost. Unlike for the general Wasserstein distance, the centering can be replaced by the population cost, enabling valid statistical inference. This generalizes and refines existing one-dimensional results, providing the first asymptotically valid inference framework for the sliced Wasserstein distance between possibly non-compact measures. Finally, we address other practical aspects crucial for inference, including Monte Carlo approximation of the slicing integral and consistent variance estimation.

math.ST

Sample Complexity of Quadratically Regularized Optimal Transport

It is well known that optimal transport suffers from the curse of dimensionality: when the prescribed marginals are approximated by i.i.d. samples, the convergence of the empirical optimal transport problem to the population counterpart slows exponentially with increasing dimension. Entropically regularized optimal transport (EOT) has become the standard bearer in many statistical applications as it avoids this curse. Indeed, EOT has parametric sample complexity, as has been shown in a series of works based on the smoothness of the EOT potentials or the strong concavity of the dual EOT problem. However, EOT produces full-support approximations to the (sparse) OT problem, leading to overspreading in applications, and is computationally unstable for small regularization parameters. The most popular alternative is quadratically regularized optimal transport (QOT), which penalizes couplings by $L^2$ norm instead of relative entropy. QOT produces sparse approximations of OT and is computationally stable. However, its potentials are not smooth (do not belong to a Donsker class) and its dual problem is not strongly concave, hence QOT is often assumed to suffer from the curse of dimensionality. In this paper, we show that QOT nevertheless has parametric sample complexity. More precisely, we establish central limit theorems for its dual potentials, optimal couplings, and optimal costs. Our analysis is based on novel arguments that focus on the regularity of the support of the optimal QOT coupling. Specifically, we establish a Lipschitz property of its sections and leverage VC theory to bound its statistical complexity. Our analysis also leads to gradient estimates of independent interest, including $C^{1,1}$ regularity of the population potentials.

math.ST

Regularity of center-outward distribution functions in non-convex domains

For a probability P in $R^d$ its center outward distribution function $F_{\pm}$, introduced in Chernozhukov et al. (2017) and Hallin et al. (2021), is a new and successful concept of multivariate distribution function based on mass transportation theory. This work proves, for a probability P with density locally bounded away from zero and infinity in its support, the continuity of the center-outward map on the interior of the support of P and the continuity of its inverse, the quantile, $Q_{\pm}$. This relaxes the convexity assumption in del Barrio et al. (2020). Some important consequences of this continuity are Glivenko-Cantelli type theorems and characterisation of weak convergence by the stability of the center-outward map.

math.PR

An improved central limit theorem and fast convergence rates for entropic transportation costs

We prove a central limit theorem for the entropic transportation cost between subgaussian probability measures, centered at the population cost. This is the first result which allows for asymptotically valid inference for entropic optimal transport between measures which are not necessarily discrete. In the compactly supported case, we complement these results with new, faster, convergence rates for the expected entropic transportation cost between empirical measures. Our proof is based on strengthening convergence results for dual solutions to the entropic optimal transport problem.

math.ST

Nonparametric Multiple-Output Center-Outward Quantile Regression

Based on the novel concept of multivariate center-outward quantiles introduced recently in Chernozhukov et al. (2017) and Hallin et al. (2021), we are considering the problem of nonparametric multiple-output quantile regression. Our approach defines nested conditional center-outward quantile regression contours and regions with given conditional probability content irrespective of the underlying distribution; their graphs constitute nested center-outward quantile regression tubes. Empirical counterparts of these concepts are constructed, yielding interpretable empirical regions and contours which are shown to consistently reconstruct their population versions in the Pompeiu-Hausdorff topology. Our method is entirely non-parametric and performs well in simulations including heteroskedasticity and nonlinear trends; its power as a data-analytic tool is illustrated on some real datasets.

stat.ME

Central Limit Theorems for Semidiscrete Wasserstein Distances

We prove a Central Limit Theorem for the empirical optimal transport cost, $\sqrt{\frac{nm}{n+m}}\{\mathcal{T}_c(P_n,Q_m)-\mathcal{T}_c(P,Q)\}$, in the semi discrete case, i.e when the distribution $P$ is supported in $N$ points, but without assumptions on $Q$. We show that the asymptotic distribution is the supremun of a centered Gaussian process, which is Gaussian under some additional conditions on the probability $Q$ and on the cost. Such results imply the central limit theorem for the $p$-Wassertein distance, for $p\geq 1$. This means that, for fixed $N$, the curse of dimensionality is avoided. To better understand the influence of such $N$, we provide bounds of $E|\mathcal{W}_1(P,Q_m)-\mathcal{W}_1(P,Q)|$ depending on $m$ and $N$. Finally, the semidiscrete framework provides a control on the second derivative of the dual formulation, which yields the first central limit theorem for the optimal transport potentials. The results are supported by simulations that help to visualize the given limits and bounds. We analyse also the cases where classical bootstrap works.

math.PR

Attraction-Repulsion clustering with applications to fairness

We consider the problem of diversity enhancing clustering, i.e, developing clustering methods which produce clusters that favour diversity with respect to a set of protected attributes such as race, sex, age, etc. In the context of fair clustering, diversity plays a major role when fairness is understood as demographic parity. To promote diversity, we introduce perturbations to the distance in the unprotected attributes that account for protected attributes in a way that resembles attraction-repulsion of charged particles in Physics. These perturbations are defined through dissimilarities with a tractable interpretation. Cluster analysis based on attraction-repulsion dissimilarities penalizes homogeneity of the clusters with respect to the protected attributes and leads to an improvement in diversity. An advantage of our approach, which falls into a pre-processing set-up, is its compatibility with a wide variety of clustering methods and whit non-Euclidean data. We illustrate the use of our procedures with both synthetic and real data and provide discussion about the relation between diversity, fairness, and cluster structure. Our procedures are implemented in an R package freely available at https://github.com/HristoInouzhe/AttractionRepulsionClustering.

stat.ML

A Central Limit Theorem for Semidiscrete Wasserstein Distances

We address the problem of proving a Central Limit Theorem for the empirical optimal transport cost, $\sqrt{n}\{\mathcal{T}_c(P_n,Q)-\mathcal{W}_c(P,Q)\}$, in the semi discrete case, i.e when the distribution $P$ is finitely supported. We show that the asymptotic distribution is the supremun of a centered Gaussian process which is Gaussian under some additional conditions on the probability $Q$ and on the cost. Such results imply the central limit theorem for the $p$-Wassertein distance, for $p\geq 1$. Finally, the semidiscrete framework provides a control on the second derivative of the dual formulation, which yields the first central limit theorem for the optimal transport potentials.

math.PR

Achieving robustness in classification using optimal transport with hinge regularization

Adversarial examples have pointed out Deep Neural Networks vulnerability to small local noise. It has been shown that constraining their Lipschitz constant should enhance robustness, but make them harder to learn with classical loss functions. We propose a new framework for binary classification, based on optimal transport, which integrates this Lipschitz constraint as a theoretical requirement. We propose to learn 1-Lipschitz networks using a new loss that is an hinge regularized version of the Kantorovich-Rubinstein dual formulation for the Wasserstein distance estimation. This loss function has a direct interpretation in terms of adversarial robustness together with certifiable robustness bound. We also prove that this hinge regularized version is still the dual formulation of an optimal transportation problem, and has a solution. We also establish several geometrical properties of this optimal solution, and extend the approach to multi-class problems. Experiments show that the proposed approach provides the expected guarantees in terms of robustness without any significant accuracy drop. The adversarial examples, on the proposed models, visibly and meaningfully change the input providing an explanation for the classification.

cs.LG

Central Limit Theorems for General Transportation Costs

We consider the problem of optimal transportation with general cost between a empirical measure and a general target probability on R d , with d $\ge$ 1. We extend results in [19] and prove asymptotic stability of both optimal transport maps and potentials for a large class of costs in R d. We derive a central limit theorem (CLT) towards a Gaussian distribution for the empirical transportation cost under minimal assumptions, with a new proof based on the Efron-Stein inequality and on the sequential compactness of the closed unit ball in L 2 (P) for the weak topology. We provide also CLTs for empirical Wassertsein distances in the special case of potential costs | $\bullet$ | p , p > 1.

math.ST

The statistical effect of entropic regularization in optimal transportation

We propose to tackle the problem of understanding the effect of regularization in Sinkhorn algotihms. In the case of Gaussian distributions we provide a closed form for the regularized optimal transport which enables to provide a better understanding of the effect of the regularization from a statistical framework.

math.ST

Review of Mathematical frameworks for Fairness in Machine Learning

A review of the main fairness definitions and fair learning methodologies proposed in the literature over the last years is presented from a mathematical point of view. Following our independence-based approach, we consider how to build fair algorithms and the consequences on the degradation of their performance compared to the possibly unfair case. This corresponds to the price for fairness given by the criteria $\textit{statistical parity}$ or $\textit{equality of odds}$. Novel results giving the expressions of the optimal fair classifier and the optimal fair predictor (under a linear regression gaussian model) in the sense of $\textit{equality of odds}$ are presented.

stat.ML

optimalFlow: Optimal-transport approach to flow cytometry gating and population matching

Data obtained from Flow Cytometry present pronounced variability due to biological and technical reasons. Biological variability is a well-known phenomenon produced by measurements on different individuals, with different characteristics such as illness, age, sex, etc. The use of different settings for measurement, the variation of the conditions during experiments and the different types of flow cytometers are some of the technical causes of variability. This mixture of sources of variability makes the use of supervised machine learning for identification of cell populations difficult. The present work is conceived as a combination of strategies to facilitate the task of supervised gating. We propose $optimalFlowTemplates$, based on a similarity distance and $\text{Wasserstein barycenters}$, which clusters cytometries and produces prototype cytometries for the different groups. We show that supervised learning, restricted to the new groups, performs better than the same techniques applied to the whole collection. We also present $optimalFlowClassification$, which uses a database of gated cytometries and optimalFlowTemplates to assign cell types to a new cytometry. We show that this procedure can outperform state of the art techniques in the proposed datasets. Our code is freely available as $optimalFlow$ a Bioconductor R package at https://bioconductor.org/packages/optimalFlow. optimalFlowTemplates+optimalFlowClassification addresses the problem of using supervised learning while accounting for biological and technical variability. Our methodology provides a robust automated gating workflow that handles the intrinsic variability of flow cytometry data well. Our main innovation is the methodology itself and the optimal-transport techniques that we apply to flow cytometry analysis.

stat.ML

A survey of bias in Machine Learning through the prism of Statistical Parity for the Adult Data Set

Applications based on Machine Learning models have now become an indispensable part of the everyday life and the professional world. A critical question then recently arised among the population: Do algorithmic decisions convey any type of discrimination against specific groups of population or minorities? In this paper, we show the importance of understanding how a bias can be introduced into automatic decisions. We first present a mathematical framework for the fair learning problem, specifically in the binary classification setting. We then propose to quantify the presence of bias by using the standard Disparate Impact index on the real and well-known Adult income data set. Finally, we check the performance of different approaches aiming to reduce the bias in binary classification outcomes. Importantly, we show that some intuitive methods are ineffective. This sheds light on the fact trying to make fair machine learning models may be a particularly challenging task, in particular when the training observations contain a bias.

stat.ML

Center-Outward Distribution Functions, Quantiles, Ranks, and Signs in $\mathbb{R}^d$

Univariate concepts as quantile and distribution functions involving ranks and signs, do not canonically extend to $\mathbb{R}^d, d\geq 2$. Palliating that has generated an abundant literature. Chapter 1 shows that, unlike the many definitions that have been proposed so far, the measure transportation-based ones introduced in Chernozhukov et al. (2017) enjoy all the properties that make univariate quantiles and ranks successful tools for semiparametric statistical inference. We therefore propose a new center-outward definition of multivariate distribution and quantile functions, along with their empirical counterparts, for which we obtain a Glivenko-Cantelli result. Our approach is geometric and, contrary to the Monge-Kantorovich one in Chernozhukov et al. (2017), does not require any moment assumptions. The resulting ranks and signs are strictly distribution-free, and maximal invariant under the action of a data-driven class of (order-preserving) transformations generating the family of absolutely continuous distributions; that property is the theoretical foundation of the semiparametric efficiency preservation property of ranks. The corresponding quantiles are equivariant under the same transformations. The empirical proposed distribution functions are defined at observed values only. A continuous extension to the entire $\mathbb{R}^d$, yielding continuous empirical quantile contours while preserving the monotonicity and Glivenko-Cantelli features is desirable. Such extension requires solving a nontrivial problem of smooth interpolation under cyclical monotonicity constraints. A complete solution of that problem is given in Chapter 2; we show that the resulting distribution and quantile functions are Lipschitz, and provide a sharp lower bound for the Lipschitz constants. A numerical study of empirical center-outward quantile contours and their consistency is conducted.

stat.ME

A note on the Regularity of Center-Outward Distribution and Quantile Functions

We provide sufficient conditions under which the center-outward distribution and quantile functions introduced in Chernozhukov et al.~(2017) and Hallin~(2017) are homeomorphisms, thereby extending a recent result by Figalli \cite{Fi2}. Our approach relies on Cafarelli's classical regularity theory for the solutions of the Monge-Ampère equation, but has to deal with difficulties related with the unboundedness at the origin of the density of the spherical uniform reference measure. Our conditions are satisfied by probabillities on Euclidean space with a general (bounded or unbounded) convex support which are not covered in~\cite{Fi2}. We provide some additional results about center-outward distribution and quantile functions, including the fact that quantile sets exhibit some weak form of convexity.

math.ST