SearcharxivSearch

arXiv subjects

Perrine Lacroix

Publications and source records attributed to Perrine Lacroix.

4 recordsLinked to original sources

Non-asymptotic two-sample kernel testing with the spectrally truncated normalized MMD

Kernel methods provide a flexible and powerful framework for nonparametric statistical testing by embedding probability distributions into a reproducing kernel Hilbert space (RKHS). In this work, we study the kernel two-sample testing problem and focus on a normalized version of the Maximum Mean Discrepancy (MMD) as a test statistic, which scales the discrepancy by the within-group covariance operator to account for data variability. This normalization has been shown to improve test power in both theoretical and empirical settings. Because this normalization requires regularization, we study the non-asymptotic properties of the spectrally truncated normalized MMD (st-nMMD) and derive an exponential upper bound under the null hypothesis. Thanks to this result we propose a sharp and explicit upper bound for the corresponding non-asymptotic quantile, along with a data-adaptive estimator. We further propose an algorithm to tune the hyperparameters involved in the quantile estimation, including the truncation level, without requiring data splitting. We demonstrate the performance of the st-nMMD through numerical experiments under both the null and alternative hypotheses.

math.ST

Eve, Adam and the Preferential Attachment Tree

We consider the problem of finding the initial vertex (Adam) in a Barabási--Albert tree process $(\mathcal{T}(n) : n \geq 1)$ at large times. More precisely, given $ \varepsilon>0$, one wants to output a subset $ \mathcal{P}_{ \varepsilon}(n)$ of vertices of $ \mathcal{T}(n)$ so that the initial vertex belongs to $ \mathcal{P}_ \varepsilon(n)$ with probability at least $1- \varepsilon$ when $n$ is large. It has been shown by Bubeck, Devroye & Lugosi, refined later by Banerjee & Huang, that one needs to output at least $ \varepsilon^{-1 + o(1)}$ and at most $\varepsilon^{-2 + o(1)}$ vertices to succeed. We prove that the exponent in the lower bound is sharp and the key idea is that Adam is either a ``large degree" vertex or is a neighbor of a ``large degree" vertex (Eve).

math.PR

Trade-off between predictive performance and FDR control for high-dimensional Gaussian model selection

In the context of high-dimensional Gaussian linear regression for ordered variables, we study the variable selection procedure via the minimization of the penalized least-squares criterion. We focus on model selection where the penalty function depends on an unknown multiplicative constant commonly calibrated for prediction. We propose a new proper calibration of this hyperparameter to simultaneously control predictive risk and false discovery rate. We obtain non-asymptotic bounds on the False Discovery Rate with respect to the hyperparameter and we provide an algorithm to calibrate it. This algorithm is based on quantities that can typically be observed in real data applications. The algorithm is validated in an extensive simulation study and is compared with several existing variable selection procedures. Finally, we study an extension of our approach to the case in which an ordering of the variables is not available.

math.ST

A comprehensive guideline for regularization-path variable selection in high-dimensional Gaussian linear regression

This paper provides a comprehensive comparison of complete regularization-path-based variable selection procedures in high-dimensional Gaussian linear regression. Our simulation study stands out from previous ones for several reasons. First, it offers comprehensive guidelines for complete variable selection, from regularization path construction to final variable subset selection. In particular, we compare jointly two convex regularization functions (Lasso and Elastic-Net) with two optimization algorithms (LARS and cyclic coordinate descent) and a range of model selection and variable identification techniques. Second, we also incorporate methods with non-asymptotic theoretical guarantees, which are typically not included in existing reviews. It is unrealistic to expect a single method that works best. Still, we provide detailed guidance on which methods perform best under different evaluation criteria, including pROC-AUC, MSE, recall, specificity, and FDP, and that are robust to data scenarios, including varying variable dependency structures. Overall, Elastic-Net combined with the LARS algorithm provides the most reliable regularization path, while the preferred final selection procedure depends on whether prediction performances, support recovery or false discovery control is prioritized. We show that some methods are empirically better for a given criterion than others, even though the latter were designed to control for it theoretically. Regarding new developments in a non-asymptotic setting, we highlight the quality of LinSelect and the need, as future work, to fine-tune the unknown constants of data-dependent penalties in high-dimensional Gaussian linear regression models.

stat.AP