SearcharxivSearch

arXiv subjects

E. del Barrio

Publications and source records attributed to E. del Barrio.

8 recordsLinked to original sources

Invariant measures of disagreement with stochastic dominance

Stochastic dominance has not been too employed in practice due to its important limitations. To increase its versatility, the concept has recently been adapted by introducing various indices that measure the degree to which one probability distribution stochastically dominates another. In this paper, starting from the fundamentals and using very simple examples, we present and discuss some of these indices when one intends to maintain invariance through increasing functions. This naturally leads to consideration of the appealing common representation, $θ(F,G)=P(X>Y)$, where $(X, Y)$ is a random vector with marginal distributions $F$ and $G$. The indices considered here arise from different dependencies between X and Y. This includes the case of independent marginals, as well as other indices related to a contamination model or to a joint quantile representation. We emphasize the complementary role of some of these indices, which, in addition to measuring disagreement with respect to stochastic dominance, enable us to describe the maximum possible difference in the status of a value $x\in \Rea$ under $F$ or $G$. We apply these indices to simulated and real-world datasets, exploring their practical advantages and limitations. The tour includes lesser-known facets of well-known statistics such as Mann-Whitney, one-tailed Kolmogorov-Smirnov and Galton's rank statistics, even providing additional theory for the latter.

stat.ME

Using the Sinkhorn divergence in permutation tests for the multivariate two-sample problem

In order to adapt the Wasserstein distance to the large sample multivariate non-parametric two-sample problem, making its application computationally feasible, permutation tests based on the Sinkhorn divergence between probability vectors associated to data dependent partitions are considered. Different ways of implementing these tests are evaluated and the asymptotic distribution of the underlying statistic is established in some cases. The statistics proposed are compared, in simulated examples, with the test of Schilling's, one of the best non-parametric tests available in the literature.

math.ST

The complex behaviour of Galton rank order statistic

Galton's rank order statistic is one of the oldest statistical tools for two-sample comparisons. It is also a very natural index to measure departures from stochastic dominance. Yet, its asymptotic behaviour has been investigated only partially, under restrictive assumptions. This work provides a comprehensive {study} of this behaviour, based on the analysis of the so-called contact set (a modification of the set in which the quantile functions coincide). We show that a.s. convergence to the population counterpart holds if and only if {the} contact set has zero Lebesgue measure. When this set is finite we show that the asymptotic behaviour is determined by the local behaviour of a suitable reparameterization of the quantile functions in a neighbourhood of the contact points. Regular crossings result in standard rates and Gaussian limiting distributions, but higher order contacts (in the sense introduced in this work) or contacts at the extremes of the supports may result in different rates and non-Gaussian limits.

math.ST

Wide Consensus for Parallelized Inference

We develop a general theory to address a consensus-based combination of estimations in a parallelized or distributed estimation setting. Taking into account the possibility of very discrepant estimations, instead of a full consensus we consider a "wide consensus" procedure. The approach is based on the consideration of trimmed barycenters in the Wasserstein space of probability distributions on R^d with finite second order moments. We include general existence and consistency results as well as characterizations of barycenters of probabilities that belong to (non necessarily elliptical) location and scatter familes. On these families, the effective computation of barycenters and distances can be addressed through a consistent iterative algorithm. Since, once a shape has been chosen, these computations just depend on the locations and scatters, the theory can be applied to cover with great generality a wide consensus approach for location and scatter estimation or for obtaining confidence regions.

stat.ME

An optimal transportation approach for assessing almost stochastic order

When stochastic dominance $F\leq_{st}G$ does not hold, we can improve agreement to stochastic order by suitably trimming both distributions. In this work we consider the $L_2-$Wasserstein distance, $\mathcal W_2$, to stochastic order of these trimmed versions. Our characterization for that distance naturally leads to consider a $\mathcal W_2$-based index of disagreement with stochastic order, $\varepsilon_{\mathcal W_2}(F,G)$. We provide asymptotic results allowing to test $H_0: \varepsilon_{\mathcal W_2}(F,G)\geq \varepsilon_0$ vs $H_a: \varepsilon_{\mathcal W_2}(F,G)<\varepsilon_0$, that, under rejection, would give statistical guarantee of almost stochastic dominance. We include a simulation study showing a good performance of the index under the normal model.

stat.ME

Models for the assessment of treatment improvement: the ideal and the feasible

Comparisons of different treatments or production processes are the goals of a significant fraction of applied research. Unsurprisingly, two-sample problems play a main role in Statistics through natural questions such as `Is the the new treatment significantly better than the old?'. However, this is only partially answered by some of the usual statistical tools for this task. More importantly, often practitioners are not aware of the real meaning behind these statistical procedures. We analyze these troubles from the point of view of the order between distributions, the stochastic order, showing evidence of the limitations of the usual approaches, paying special attention to the classical comparison of means under the normal model. We discuss the unfeasibility of statistically proving stochastic dominance, but show that it is possible, instead, to gather statistical evidence to conclude that slightly relaxed versions of stochastic dominance hold.

stat.ME

Robust clustering tools based on optimal transportation

A robust clustering method for probabilities in Wasserstein space is introduced. This new "trimmed $k$-barycenters" approach relies on recent results on barycenters in Wasserstein space that allow intensive computation, as required by clustering algorithms. The possibility of trimming the most discrepant distributions results in a gain in stability and robustness, highly convenient in this setting. As a remarkable application we consider a parallelized estimation setup in which each of $m$ units processes a portion of the data, producing an estimate of $k$-features, encoded as $k$ probabilities. We prove that the trimmed $k$-barycenter of the $m\times k$ estimates produces a consistent aggregation. We illustrate the methodology with simulated and real data examples. These include clustering populations by age distributions and analysis of cytometric data.

stat.ME

A fixed-point approach to barycenters in Wasserstein space

Let $\mathcal{P}_{2,ac}$ be the set of Borel probabilities on $\mathbb{R}^d$ with finite second moment and absolutely continuous with respect to Lebesgue measure. We consider the problem of finding the barycenter (or Fréchet mean) of a finite set of probabilities $ν_1,\ldots,ν_k \in \mathcal{P}_{2,ac}$ with respect to the $L_2-$Wasserstein metric. For this task we introduce an operator on $\mathcal{P}_{2,ac}$ related to the optimal transport maps pushing forward any $μ\in \mathcal{P}_{2,ac}$ to $ν_1,\ldots,ν_k$. Under very general conditions we prove that the barycenter must be a fixed point for this operator and introduce an iterative procedure which consistently approximates the barycenter. The procedure allows effective computation of barycenters in any location-scatter family, including the Gaussian case. In such cases the barycenter must belong to the family, thus it is characterized by its mean and covariance matrix. While its mean is just the weighted mean of the means of the probabilities, the covariance matrix is characterized in terms of their covariance matrices $Σ_1,\dots,Σ_k$ through a nonlinear matrix equation. The performance of the iterative procedure in this case is illustrated through numerical simulations, which show fast convergence towards the barycenter.

stat.CO