SearcharxivSearch

arXiv subjects

Gabriel Romon

Publications and source records attributed to Gabriel Romon.

5 recordsLinked to original sources

Statistical Inference via T-Posterior Randomised Estimators

Given a statistical model, we propose a novel estimation method that yields randomised estimators for the unknown distribution of an observed random variable. We establish non-asymptotic bounds for the performance of these estimators and demonstrate their robustness to potential model misspecification. Notably, these properties are established by circumventing the use of concentration inequalities and empirical process theory. We provide an illustration of this approach to the problem of estimating the intensity of a Poisson process.

math.ST

Convex generalized Fr\'echet means in a metric tree

We are interested in measures of central tendency for a population $\mu$ on a network, which is modeled by a metric tree. The location parameters that we study are generalized Fr\'echet means defined as minimizers of the objective function $\alpha \mapsto \mathbb E[\ell(d(\alpha,X)) - \ell(d(o,X))]$, where $\ell$ is a convex, strictly increasing loss, $X\sim \mu$ and $o$ is an arbitrary origin. We leverage the geometry of the tree and the geodesic convexity of the objective to develop a notion of directional derivative in the tree, which helps us locate and characterize the minimizers. We then extend to a metric tree the concept of stickiness defined by Hotz et al. (2013). When $\mu$ is nondegenerate, any Fr\'echet $\ell$-mean of $\mu$ is either sticky, one-sided partly sticky or two-sided partly sticky, depending on the number of vanishing directional derivatives. Estimation is performed using a sample analog. We establish laws of large numbers and central limit theorems that reveal distinct asymptotic behaviors of empirical means across these three regimes: collapse onto the population mean in the sticky case, partial collapse accompanied by one-sided fluctuations in the one-sided case, and two-sided fluctuations without collapse in the two-sided case. We further develop two consistent estimators of the stickiness regime, based respectively on pairwise collisions among empirical means and on empirical difference quotients. For the particular case of the Fr\'echet median, we develop distribution-free non-asymptotic confidence regions for the entire set of minimizers.

math.ST

Statistical properties of approximate geometric quantiles in infinite-dimensional Banach spaces

Geometric quantiles are location parameters which extend classical univariate quantiles to normed spaces (possibly infinite-dimensional) and which include the geometric median as a special case. The infinite-dimensional setting is highly relevant in the modeling and analysis of functional data, as well as for kernel methods. We begin by providing new results on the existence and uniqueness of geometric quantiles. Estimation is then performed with an approximate M-estimator and we investigate its large-sample properties in infinite dimension. When the population quantile is not uniquely defined, we leverage the theory of variational convergence to obtain asymptotic statements on subsequences in the weak topology. When there is a unique population quantile, we show, under minimal assumptions, that the estimator is consistent in the norm topology for a wide range of Banach spaces including every separable uniformly convex space. In separable Hilbert spaces, we establish weak Bahadur-Kiefer representations of the estimator, from which $\sqrt n$-asymptotic normality follows. As a consequence, we obtain the first central limit theorem valid in a generic Hilbert space and under minimal assumptions that exactly match those of the finite-dimensional case. Our consistency and asymptotic normality results significantly improve the state of the art, even for exact geometric medians in Hilbert spaces.

math.ST

Noise Covariance Estimation in Multi-Task High-dimensional Linear Models

This paper studies the multi-task high-dimensional linear regression models where the noise among different tasks is correlated, in the moderately high dimensional regime where sample size $n$ and dimension $p$ are of the same order. Our goal is to estimate the covariance matrix of the noise random vectors, or equivalently the correlation of the noise variables on any pair of two tasks. Treating the regression coefficients as a nuisance parameter, we leverage the multi-task elastic-net and multi-task lasso estimators to estimate the nuisance. By precisely understanding the bias of the squared residual matrix and by correcting this bias, we develop a novel estimator of the noise covariance that converges in Frobenius norm at the rate $n^{-1/2}$ when the covariates are Gaussian. This novel estimator is efficiently computable. Under suitable conditions, the proposed estimator of the noise covariance attains the same rate of convergence as the "oracle" estimator that knows in advance the regression coefficients of the multi-task model. The Frobenius error bounds obtained in this paper also illustrate the advantage of this new estimator compared to a method-of-moments estimator that does not attempt to estimate the nuisance. As a byproduct of our techniques, we obtain an estimate of the generalization error of the multi-task elastic-net and multi-task lasso estimators. Extensive simulation studies are carried out to illustrate the numerical performance of the proposed method.

math.ST

Chi-square and normal inference in high-dimensional multi-task regression

The paper proposes chi-square and normal inference methodologies for the unknown coefficient matrix $B^*$ of size $p\times T$ in a Multi-Task (MT) linear model with $p$ covariates, $T$ tasks and $n$ observations under a row-sparse assumption on $B^*$. The row-sparsity $s$, dimension $p$ and number of tasks $T$ are allowed to grow with $n$. In the high-dimensional regime $p\ggg n$, in order to leverage row-sparsity, the MT Lasso is considered. We build upon the MT Lasso with a de-biasing scheme to correct for the bias induced by the penalty. This scheme requires the introduction of a new data-driven object, coined the interaction matrix, that captures effective correlations between noise vector and residuals on different tasks. This matrix is psd, of size $T\times T$ and can be computed efficiently. The interaction matrix lets us derive asymptotic normal and $χ^2_T$ results under Gaussian design and $\frac{sT+s\log(p/s)}{n}\to0$ which corresponds to consistency in Frobenius norm. These asymptotic distribution results yield valid confidence intervals for single entries of $B^*$ and valid confidence ellipsoids for single rows of $B^*$, for both known and unknown design covariance $Σ$. While previous proposals in grouped-variables regression require row-sparsity $s\lesssim\sqrt n$ up to constants depending on $T$ and logarithmic factors in $n,p$, the de-biasing scheme using the interaction matrix provides confidence intervals and $χ^2_T$ confidence ellipsoids under the conditions ${\min(T^2,\log^8p)}/{n}\to 0$ and $$ \frac{sT+s\log(p/s)+\|Σ^{-1}e_j\|_0\log p}{n}\to0, \quad \frac{\min(s,\|Σ^{-1}e_j\|_0)}{\sqrt n} \sqrt{[T+\log(p/s)]\log p}\to 0, $$ allowing row-sparsity $s\ggg\sqrt n$ when $\|Σ^{-1}e_j\|_0 \sqrt T\lll \sqrt{n}$ up to logarithmic factors.

math.ST