SearcharxivSearch

arXiv subjects

Matthias Löffler

Publications and source records attributed to Matthias Löffler.

10 recordsLinked to original sources

AdaBoost and robust one-bit compressed sensing

This paper studies binary classification in robust one-bit compressed sensing with adversarial errors. It is assumed that the model is overparameterized and that the parameter of interest is effectively sparse. AdaBoost is considered, and, through its relation to the max-$\ell_1$-margin-classifier, prediction error bounds are derived. The developed theory is general and allows for heavy-tailed feature distributions, requiring only a weak moment assumption and an anti-concentration condition. Improved convergence rates are shown when the features satisfy a small deviation lower bound. In particular, the results provide an explanation why interpolating adversarial noise can be harmless for classification problems. Simulations illustrate the presented theory.

math.ST

Reconstruction of the neutrino mass as a function of redshift

We reconstruct the neutrino mass as a function of redshift, z, from current cosmological data using both standard binned priors and linear spline priors with variable knots. Using cosmic microwave background temperature, polarization and lensing data, in combination with distance measurements from baryonic acoustic oscillations and supernovae, we find that the neutrino mass is consistent with $\sum m_ν(z)$ = const. We obtain a larger bound on the neutrino mass at low redshifts coinciding with the onset of dark energy domination, $\sum m_ν(z = 0)$ < 1.46 eV (95% CL). This result can be explained either by the well-known degeneracy between $\sum m_ν$ and $Ω_Λ$ at low redshifts, or by models in which neutrino masses are generated very late in the Universe. We finally convert our results into cosmological limits for models with non-relativistic neutrino decay and find $\sum m_ν$ < 0.21 eV (95% CL), which would be out of reach for the KATRIN experiment.

astro-ph.CO

Spectral thresholding for the estimation of Markov chain transition operators

We consider nonparametric estimation of the transition operator $P$ of a Markov chain and its transition density $p$ where the singular values of $P$ are assumed to decay exponentially fast. This is for instance the case for periodised, reversible multi-dimensional diffusion processes observed in low frequency. We investigate the performance of a spectral hard thresholded Galerkin-type estimator for $P$ and ${p}$, discarding most of the estimated singular triplets. The construction is based on smooth basis functions such as wavelets or B-splines. We show its statistical optimality by establishing matching minimax upper and lower bounds in $L^2$-loss. Particularly, the effect of the dimensionality $d$ of the state space on the nonparametric rate improves from $2d$ to $d$ compared to the case without singular value decay.

math.ST

On the robustness of minimum norm interpolators and regularized empirical risk minimizers

This article develops a general theory for minimum norm interpolating estimators and regularized empirical risk minimizers (RERM) in linear models in the presence of additive, potentially adversarial, errors. In particular, no conditions on the errors are imposed. A quantitative bound for the prediction error is given, relating it to the Rademacher complexity of the covariates, the norm of the minimum norm interpolator of the errors and the size of the subdifferential around the true parameter. The general theory is illustrated for Gaussian features and several norms: The $\ell_1$, $\ell_2$, group Lasso and nuclear norms. In case of sparsity or low-rank inducing norms, minimum norm interpolators and RERM yield a prediction error of the order of the average noise level, provided that the overparameterization is at least a logarithmic factor larger than the number of samples and that, in case of RERM, the regularization parameter is small enough. Lower bounds that show near optimality of the results complement the analysis.

math.ST

Computationally efficient sparse clustering

We study statistical and computational limits of clustering when the means of the centres are sparse and their dimension is possibly much larger than the sample size. Our theoretical analysis focuses on the model $X_i = z_i θ+ \varepsilon_i, ~z_i \in \{-1,1\}, ~\varepsilon_i \thicksim \mathcal{N}(0,I)$, which has two clusters with centres $θ$ and $-θ$. We provide a finite sample analysis of a new sparse clustering algorithm based on sparse PCA and show that it achieves the minimax optimal misclustering rate in the regime $\|θ\| \rightarrow \infty$. Our results require the sparsity to grow slower than the square root of the sample size. Using a recent framework for computational lower bounds -- the low-degree likelihood ratio -- we give evidence that this condition is necessary for any polynomial-time clustering algorithm to succeed below the BBP threshold. This complements existing evidence based on reductions and statistical query lower bounds. Compared to these existing results, we cover a wider set of parameter regimes and give a more precise understanding of the runtime required and the misclustering error achievable. Our results imply that a large class of tests based on low-degree polynomials fail to solve even the weak testing task.

math.ST

Optimality of Spectral Clustering in the Gaussian Mixture Model

Spectral clustering is one of the most popular algorithms to group high dimensional data. It is easy to implement and computationally efficient. Despite its popularity and successful applications, its theoretical properties have not been fully understood. In this paper, we show that spectral clustering is minimax optimal in the Gaussian Mixture Model with isotropic covariance matrix, when the number of clusters is fixed and the signal-to-noise ratio is large enough. Spectral gap conditions are widely assumed in the literature to analyze spectral clustering. On the contrary, these conditions are not needed to establish optimality of spectral clustering in this paper.

math.ST

Efficient Estimation of Linear Functionals of Principal Components

We study principal component analysis (PCA) for mean zero i.i.d. Gaussian observations $X_1,\dots, X_n$ in a separable Hilbert space $\mathbb{H}$ with unknown covariance operator $Σ.$ The complexity of the problem is characterized by its effective rank ${\bf r}(Σ):= \frac{{\rm tr}(Σ)}{\|Σ\|},$ where ${\rm tr}(Σ)$ denotes the trace of $Σ$ and $\|Σ\|$ denotes its operator norm. We develop a method of bias reduction in the problem of estimation of linear functionals of eigenvectors of $Σ.$ Under the assumption that ${\bf r}(Σ)=o(n),$ we establish the asymptotic normality and asymptotic properties of the risk of the resulting estimators and prove matching minimax lower bounds, showing their semi-parametric optimality.

math.ST

Wald Statistics in high-dimensional PCA

In this note we consider PCA for Gaussian observations $X_1,\dots, X_n$ with covariance $Σ=\sum_i λ_i P_i$ in the 'effective rank' setting with model complexity governed by $\mathbf{r}(Σ):=\text{tr}(Σ)/\| Σ\|$. We prove a Berry-Essen type bound for a Wald Statistic of the spectral projector $\hat P_r$. This can be used to construct non-asymptotic confidence ellipsoids and tests for spectral projectors $P_r$. Using higher order pertubation theory we are able to show that our Theorem remains valid even when $\mathbf{r}(Σ) \gg \sqrt{n}$.

math.ST

Constructing confidence sets for the matrix completion problem

In the present note we consider the problem of constructing honest and adaptive confidence sets for the matrix completion problem. For the Bernoulli model with known variance of the noise we provide a realizable method for constructing confidence sets that adapt to the unknown rank of the true matrix.

math.ST

Adaptive confidence sets for matrix completion

In the present paper we study the problem of existence of honest and adaptive confidence sets for matrix completion. We consider two statistical models: the trace regression model and the Bernoulli model. In the trace regression model, we show that honest confidence sets that adapt to the unknown rank of the matrix exist even when the error variance is unknown. Contrary to this, we prove that in the Bernoulli model, honest and adaptive confidence sets exist only when the error variance is known a priori. In the course of our proofs we obtain bounds for the minimax rates of certain composite hypothesis testing problems arising in low rank inference.

math.ST