Searcharxiv⌕ Search

arXiv subjects

Clément Marteau

Publications and source records attributed to Clément Marteau.

At least 19 recordsLinked to original sources

Riemannian Gradient Descent for Gaussian Mixture Models with unknown diagonal covariances

This paper investigates the numerical resolution of the Beurling-LASSO (BLASSO), a convex optimization framework that promotes sparsity in the space of measures. We consider its application to the estimation of Gaussian mixture models (GMMs) with an unknown number of components and unknown diagonal covariance matrices. Our approach combines the Conic Particle Gradient Descent (CPGD) principle with Riemannian gradient descent, to account for the underlying Fisher-Rao geometry of Gaussian distributions. Our contributions are twofold. First, we provide theoretical guarantees for the convergence of our algorithm. In particular, we establish exponential local convergence under a non-degeneracy condition on the solution and relate this assumption to a separation condition on the underlying statistical target. Second, we address practical implementation aspects of CPGD and present numerical experiments illustrating its performance. On the test cases considered, these experiments suggest that CPGD is more robust to overspecification of the number of components than the EM algorithm. We also investigate the impact of component separation on recovery accuracy.

math.OC↗

Gaussian Mixture Model with unknown diagonal covariances via continuous sparse regularization

This paper addresses the statistical estimation of Gaussian Mixture Models (GMMs) with unknown diagonal covariances from independent and identically distributed samples. We employ the Beurling-LASSO (BLASSO), a convex optimization framework that promotes sparsity in the space of measures, to simultaneously estimate the number of components and their parameters. Our main contribution extends the BLASSO methodology to multivariate GMMs with component-specific unknown diagonal covariance matrices. This setting is significantly more flexible than previous approaches, which required known and identical covariances. We establish non-asymptotic recovery guarantees with nearly parametric convergence rates for component means, diagonal covariances, and weights, as well as for density prediction. A key theoretical contribution is the identification of an explicit separation condition on mixture components that enables the construction of non-degenerate dual certificates-essential tools for establishing statistical guarantees for the BLASSO. Our analysis leverages the Fisher-Rao geometry of the statistical model and introduces a novel semi-distance adapted to our framework, providing new insights into the interplay between component separation, parameter space geometry, and achievable statistical recovery.

math.ST↗

Fast Spawn\&Prune (FS\&P): Global convergence of stochastic conic particle gradient descent via birth/death process

We investigate the global optimization of the objective function arising in continuous sparse regression, specifically the Beurling LASSO (BLASSO), over the space of measures. While Conic Particle Gradient Descent (CPGD) methods are computationally efficient, they may become trapped in local minima due to the non-convexity of the parameterization. To overcome this limitation, we introduce Fast Spawn\&Prune (FS\&P), a stochastic algorithm that extends FastPart introduced in De Castro et al. (2025) and combines CPGD with a birth-death process. The birth mechanism ensures asymptotic global exploration by introducing particles in regions where first-order optimality conditions are violated, while the death process preserves computational efficiency by pruning non-informative particles. We provide the first theoretical guarantee of global convergence for this class of discrete-time stochastic algorithms, without requiring exponentially large initializations. Furthermore, we derive explicit convergence rates for the excess risk, which scale as $\mathcal{O}\big(\left(\log K / K\right)^{\frac{1}{2(2+d)}}\big)$, where $K$ denotes the number of iterations and d the dimension of the domain, thereby quantifying the trade-off between global exploration and local refinement. Moreover, the sample complexity is $\mathcal{O}\big(N^{-\frac{1}{4(2+d)}}\big)$ (up to logarithmic factors). We also propose a horizon-free variant that does not require prior knowledge of the iteration budget.

math.OC↗

FastPart: Over-Parameterized Stochastic Gradient Descent for Sparse optimisation on Measures

This paper presents a novel algorithm that leverages Stochastic Gradient Descent strategies in conjunction with Random Features to augment the scalability of Conic Particle Gradient Descent (CPGD) specifically tailored for solving sparse optimization problems on measures. By formulating the CPGD steps within a variational framework, we provide rigorous mathematical proofs demonstrating the following key findings: $\mathrm{(i)}$ The total variation norms of the solution measures along the descent trajectory remain bounded, ensuring stability and preventing undesirable divergence; $\mathrm{(ii)}$ We establish a global convergence guarantee with a convergence rate of ${O}(\log(K)/\sqrt{K})$ over $K$ iterations, showcasing the efficiency and effectiveness of our algorithm, $\mathrm{(iii)}$ Additionally, we analyse and establish local control over the first-order condition discrepancy, contributing to a deeper understanding of the algorithm's behaviour and reliability in practical applications.

math.OC↗

Minimax testing in a statistical inverse problem with unknown operator

We study minimax testing in a statistical inverse problem when the associated operator is unknown. In particular, we consider observations from an inverse Gaussian regression model where the associated operator is unknown but contained in a given dictionary B of finite cardinality. Using the non-asymptotic framework for minimax testing (that is, for any fixed value of the noise level), we provide optimal separation conditions for the goodness-of-fit testing problem. We restrict our attention to the specific case where the dictionary contains only two members. As we will demonstrate, even this simple case is quite intrigued and reveals an interesting phase transition phenomenon. The general case is even more involved, requires different strategies, and it is only briefly discussed.

math.ST↗

A non-asymptotic upper bound in prediction for the PLS estimator

We investigate the theoretical performances of the Partial Least Square (PLS) algorithm in a high dimensional context. We provide upper bounds on the risk in prediction for the statistical linear model when considering the PLS estimator. Our bounds are non-asymptotic and are expressed in terms of the number of observations, the noise level, the properties of the design matrix, and the number of considered PLS components. In particular, we exhibit some scenarios where the variability of the PLS may explode and prove that we can get round of these situations by introducing a Ridge regularization step. These theoretical findings are illustrated by some numerical simulations.

math.ST↗

A non-asymptotic analysis of the single component PLS regression

This paper investigates some theoretical properties of the Partial Least Square (PLS) method. We focus our attention on the single component case, that provides a useful framework to understand the underlying mechanism. We provide a non-asymptotic upper bound on the quadratic loss in prediction with high probability in a high dimensional regression context. The bound is attained thanks to a preliminary test on the first PLS component. In a second time, we extend these results to the sparse partial least squares (sPLS) approach. In particular, we exhibit upper bounds similar to those obtained with the lasso algorithm, up to an additional restricted eigenvalue constraint on the design matrix.

math.ST↗

Dual-sPLS: a family of Dual Sparse Partial Least Squares regressions for feature selection and prediction with tunable sparsity; evaluation on simulated and near-infrared (NIR) data

Relating a set of variables X to a response y is crucial in chemometrics. A quantitative prediction objective can be enriched by qualitative data interpretation, for instance by locating the most influential features. When high-dimensional problems arise, dimension reduction techniques can be used. Most notable are projections (e.g. Partial Least Squares or PLS ) or variable selections (e.g. lasso). Sparse partial least squares combine both strategies, by blending variable selection into PLS. The variant presented in this paper, Dual-sPLS, generalizes the classical PLS1 algorithm. It provides balance between accurate prediction and efficient interpretation. It is based on penalizations inspired by classical regression methods (lasso, group lasso, least squares, ridge) and uses the dual norm notion. The resulting sparsity is enforced by an intuitive shrinking ratio parameter. Dual-sPLS favorably compares to similar regression methods, on simulated and real chemical data. Code is provided as an open-source package in R: \url{https://CRAN.R-project.org/package=dual.spls}.

stat.ML↗

SuperMix: Sparse Regularization for Mixtures

This paper investigates the statistical estimation of a discrete mixing measure $μ$0 involved in a kernel mixture model. Using some recent advances in l1-regularization over the space of measures, we introduce a "data fitting and regularization" convex program for estimating $μ$0 in a grid-less manner from a sample of mixture law, this method is referred to as Beurling-LASSO. Our contribution is twofold: we derive a lower bound on the bandwidth of our data fitting term depending only on the support of $μ$0 and its so-called "minimum separation" to ensure quantitative support localization error bounds; and under a so-called "non-degenerate source condition" we derive a non-asymptotic support stability property. This latter shows that for a sufficiently large sample size n, our estimator has exactly as many weighted Dirac masses as the target $μ$0 , converging in amplitude and localization towards the true ones. Finally, we also introduce some tractable algorithms for solving this convex program based on "Sliding Frank-Wolfe" or "Conic Particle Gradient Descent". Statistical performances of this estimator are investigated designing a so-called "dual certificate", which is appropriate to our setting. Some classical situations, as e.g. mixtures of super-smooth distributions (e.g. Gaussian distributions) or ordinary-smooth distributions (e.g. Laplace distributions), are discussed at the end of the paper.

math.ST↗

Parameter recovery in two-component contamination mixtures: the $\mathbb{L}^2$ strategy

In this paper, we consider a parametric density contamination model. We work with a sample of i.i.d. data with a common density, $f^\star =(1-λ^\star) ϕ+ λ^\star ϕ(.-μ^\star)$, where the shape $ϕ$ is assumed to be known. We establish the optimal rates of convergence for the estimation of the mixture parameters $(λ^\star,μ^\star)$. In particular, we prove that the classical parametric rate $1/\sqrt{n}$ cannot be reached when at least one of these parameters is allowed to tend to $0$ with $n$.

math.ST↗

Maxiset point of view for signal detection in inverse problems

This paper extends the successful maxiset paradigm from function estimation to signal detection in inverse problems. In this context, the maxisets do not have the same shape compared to the classical estimation framework. Nevertheless, we introduce a robustified version of these maxisets, allowing to exhibit tail conditions on the signals of interest. Under this novel paradigm we are able to compare direct and indirect testing procedures.

math.ST↗

Optimal functional supervised classification with separation condition

We consider the binary supervised classification problem with the Gaussian functional model introduced in [7]. Taking advantage of the Gaussian structure, we design a natural plug-in classifier and derive a family of upper bounds on its worst-case excess risk over Sobolev spaces. These bounds are parametrized by a separation distance quantifying the difficulty of the problem, and are proved to be optimal (up to logarithmic factors) through matching minimax lower bounds. Using the recent works of [9] and [14] we also derive a logarithmic lower bound showing that the popular k-nearest neighbors classifier is far from optimality in this specific functional setting.

math.ST↗

Non-asymptotic detection of two-component mixtures with unknown means

This work is concerned with the detection of a mixture distribution from a $\mathbb{R}$-valued sample. Given a sample $X_1,\dots,X_n$ and an even density $ϕ$, our aim is to detect whether the sample distribution is $ϕ(\cdot-μ)$ for some unknown mean $μ$, or is defined as a two-component mixture based on translations of $ϕ$. We propose a procedure which is based on several spacings of the order statistics, which provides a level-$α$ test for all $n$. Our test is therefore a multiple testing procedure and we prove from a theoretical and practical point of view that it automatically adapts to the proportion of the mixture and to the difference of the means of the two components of the mixture under the alternative. From a theoretical point of view, we prove the optimality of the power of our procedure in various situations. A simulation study shows the good performances of our test compared with several classical procedures.

math.ST↗

Multidimensional two-component Gaussian mixtures detection

Let $(X\_1,\ldots,X\_n)$ be a $d$-dimensional i.i.d sample from a distribution with density $f$. The problem of detection of a two-component mixture is considered. Our aim is to decide whether $f$ is the density of a standard Gaussian random $d$-vector ($f=ϕ\_d$) against $f$ is a two-component mixture: $f=(1-\varepsilon)ϕ\_d +\varepsilon ϕ\_d (.-μ)$ where $(\varepsilon,μ)$ are unknown parameters. Optimal separation conditions on $\varepsilon, μ, n$ and the dimension $d$ are established, allowing to separate both hypotheses with prescribed errors. Several testing procedures are proposed and two alternative subsets are considered.

math.ST↗

Minimax Goodness-of-Fit Testing in Ill-Posed Inverse Problems with Partially Unknown Operators

We consider a Gaussian sequence model that contains ill-posed inverse problems as special cases. We assume that the associated operator is partially unknown in the sense that its singular functions are known and the corresponding singular values are unknown but observed with Gaussian noise. For the considered model, we study the minimax goodness-of-fit testing problem. Working with certain ellipsoids in the space of squared-summable sequences of real numbers, with a ball of positive radius removed, we obtain lower and upper bounds for the minimax separation radius in the non-asymptotic framework, i.e., for fixed values of the involved noise levels. Examples of mildly and severely ill-posed inverse problems with ellipsoids of ordinary-smooth and super-smooth sequences are examined in detail and minimax rates of goodness-of-fit testing are obtained for illustrative purposes.

math.ST↗

Minimax fast rates for discriminant analysis with errors in variables

The effect of measurement errors in discriminant analysis is investigated. Given observations $Z=X+ε$, where $ε$ denotes a random noise, the goal is to predict the density of $X$ among two possible candidates $f$ and $g$. We suppose that we have at our disposal two learning samples. The aim is to approach the best possible decision rule $G^\star$ defined as a minimizer of the Bayes risk. In the free-noise case $(ε=0)$, minimax fast rates of convergence are well-known under the margin assumption in discriminant analysis (see \cite{mammen}) or in the more general classification framework (see \cite{tsybakov2004,AT}). In this paper we intend to establish similar results in the noisy case, i.e. when dealing with errors in variables. We prove minimax lower bounds for this problem and explain how can these rates be attained, using in particular an Empirical Risk Minimizer (ERM) method based on deconvolution kernel estimators.

math.ST↗

Classification with the nearest neighbor rule in general finite dimensional spaces: necessary and sufficient conditions

Given an $n$-sample of random vectors $(X_i,Y_i)_{1 \leq i \leq n}$ whose joint law is unknown, the long-standing problem of supervised classification aims to \textit{optimally} predict the label $Y$ of a given a new observation $X$. In this context, the nearest neighbor rule is a popular flexible and intuitive method in non-parametric situations. Even if this algorithm is commonly used in the machine learning and statistics communities, less is known about its prediction ability in general finite dimensional spaces, especially when the support of the density of the observations is $\mathbb{R}^d$. This paper is devoted to the study of the statistical properties of the nearest neighbor rule in various situations. In particular, attention is paid to the marginal law of $X$, as well as the smoothness and margin properties of the \textit{regression function} $η(X) = \mathbb{E}[Y | X]$. We identify two necessary and sufficient conditions to obtain uniform consistency rates of classification and to derive sharp estimates in the case of the nearest neighbor rule. Some numerical experiments are proposed at the end of the paper to help illustrate the discussion.

math.ST↗