SearcharxivSearch

arXiv subjects

Quentin Paris

Publications and source records attributed to Quentin Paris.

11 recordsLinked to original sources

Online learning with exponential weights in metric spaces

This paper addresses the problem of online learning in metric spaces using exponential weights. We extend the analysis of the exponentially weighted average forecaster, traditionally studied in a Euclidean settings, to a more abstract framework. Our results rely on the notion of barycenters, a suitable version of Jensen's inequality and a synthetic notion of lower curvature bound in metric spaces known as the measure contraction property. We also adapt the online-to-batch conversion principle to apply our results to a statistical learning framework.

stat.ML

Jensen's inequality in geodesic spaces with lower bounded curvature

Let $(M,d)$ be a separable and complete geodesic space with curvature lower bounded, by $\kappa\in \mathbb R$, in the sense of Alexandrov. Let $\mu$ be a Borel probability measure on $M$, such that $\mu\in\mathcal P_2(M)$, and that has at least one barycenter $x^{*}\in M$. We show that for any geodesically $\alpha$-convex function $f:M\to \mathbb R$, for $\alpha\in \mathbb R$, the inequality \[f(x^*)\le \int_M (f -\frac{\alpha}{2}d^2(x^*,.))\,{\rm d}\mu,\] holds provided $f$ is locally Lipschitz at $x^*$ and either positive or in $L^1(\mu)$. Our proof relies on the properties of tangent cones at barycenters and on the existence of gradients for semi-concave functions in spaces with lower bounded curvature.

math.MG

The exponentially weighted average forecaster in geodesic spaces of non-positive curvature

This paper addresses the problem of prediction with expert advice for outcomes in a geodesic space with non-positive curvature in the sense of Alexandrov. Via geometric considerations, and in particular the notion of barycenters, we extend to this setting the definition and analysis of the classical exponentially weighted average forecaster. We also adapt the principle of online to batch conversion to this setting. We shortly discuss the application of these results in the context of aggregation and for the problem of barycenter estimation.

math.ST

Fast convergence of empirical barycenters in Alexandrov spaces and the Wasserstein space

This work establishes fast rates of convergence for empirical barycenters over a large class of geodesic spaces with curvature bounds in the sense of Alexandrov. More specifically, we show that parametric rates of convergence are achievable under natural conditions that characterize the bi-extendibility of geodesics emanating from a barycenter. These results largely advance the state-of-the-art on the subject both in terms of rates of convergence and the variety of spaces covered. In particular, our results apply to infinite-dimensional spaces such as the 2-Wasserstein space, where bi-extendibility of geodesics translates into regularity of Kantorovich potentials.

math.ST

On the occupancy problem for a regime switching model

This article studies the expected occupancy probabilities on an alphabet. Unlike the standard situation, where observations are assumed to be independent and identically distributed (iid), we assume that they follow a regime switching Markov chain. For this model, we 1) give finite sample bounds on the occupancy probabilities, and 2) provide detailed asymptotics in the case where the underlying distribution is regularly varying. We find that, in the regularly varying case, the finite sample bounds are rate optimal and have, up to a constant, the same rate of decay as the asymptotic result.

math.PR

Convergence rates for empirical barycenters in metric spaces: curvature, convexity and extendible geodesics

This paper provides rates of convergence for empirical (generalised) barycenters on compact geodesic metric spaces under general conditions using empirical processes techniques. Our main assumption is termed a variance inequality and provides a strong connection between usual assumptions in the field of empirical processes and central concepts of metric geometry. We study the validity of variance inequalities in spaces of non-positive and non-negative Aleksandrov curvature. In this last scenario, we show that variance inequalities hold provided geodesics, emanating from a barycenter, can be extended by a constant factor. We also relate variance inequalities to strong geodesic convexity. While not restricted to this setting, our results are largely discussed in the context of the $2$-Wasserstein space.

math.ST

A notion of stability for k-means clustering

In this paper, we define and study a new notion of stability for the $k$-means clustering scheme building upon the notion of quantization of a probability measure. We connect this notion of stability to a geometric feature of the underlying distribution of the data, named absolute margin condition, inspired by recent works on the subject.

math.ST

On the Exponentially Weighted Aggregate with the Laplace Prior

In this paper, we study the statistical behaviour of the Exponentially Weighted Aggregate (EWA) in the problem of high-dimensional regression with fixed design. Under the assumption that the underlying regression vector is sparse, it is reasonable to use the Laplace distribution as a prior. The resulting estimator and, specifically, a particular instance of it referred to as the Bayesian lasso, was already used in the statistical literature because of its computational convenience, even though no thorough mathematical analysis of its statistical properties was carried out. The present work fills this gap by establishing sharp oracle inequalities for the EWA with the Laplace prior. These inequalities show that if the temperature parameter is small, the EWA with the Laplace prior satisfies the same type of oracle inequality as the lasso estimator does, as long as the quality of estimation is measured by the prediction loss. Extensions of the proposed methodology to the problem of prediction with low-rank matrices are considered.

math.ST

On the prediction loss of the lasso in the partially labeled setting

In this paper we revisit the risk bounds of the lasso estimator in the context of transductive and semi-supervised learning. In other terms, the setting under consideration is that of regression with random design under partial labeling. The main goal is to obtain user-friendly bounds on the off-sample prediction risk. To this end, the simple setting of bounded response variable and bounded (high-dimensional) covariates is considered. We propose some new adaptations of the lasso to these settings and establish oracle inequalities both in expectation and in deviation. These results provide non-asymptotic upper bounds on the risk that highlight the interplay between the bias due to the mis-specification of the linear model, the bias due to the approximate sparsity and the variance. They also demonstrate that the presence of a large number of unlabeled features may have significant positive impact in the situations where the restricted eigenvalue of the design matrix vanishes or is very small.

math.ST

Finite sample properties of the mean occupancy counts and probabilities

For a probability distribution $P$ on an at most countable alphabet $\mathcal A$, this article gives finite sample bounds for the expected occupancy counts $\mathbb E K_{n,r}$ and probabilities $\mathbb E M_{n,r}$. Both upper and lower bounds are given in terms of the counting function $\nu$ of $P$. Special attention is given to the case where $\nu$ is bounded by a regularly varying function. In this case, it is shown that our general results lead to an optimal-rate control of the expected occupancy counts and probabilities with explicit constants. Our results are also put in perspective with Turing's formula and recent concentration bounds to deduce bounds in probability. At the end of the paper, we discuss an extension of the occupancy problem to arbitrary distributions in a metric space.

math.ST

Cox process functional learning

This article addresses the problem of functional supervised classification of Cox process trajectories, whose random intensity is driven by some exogenous random covariable. The classification task is achieved through a regularized convex empirical risk minimization procedure, and a nonasymptotic oracle inequality is derived. We show that the algorithm provides a Bayes-risk consistent classifier. Furthermore, it is proved that the classifier converges at a rate which adapts to the unknown regularity of the intensity process. Our results are obtained by taking advantage of martingale and stochastic calculus arguments, which are natural in this context and fully exploit the functional nature of the problem.

math.ST