SearcharxivSearch

arXiv subjects

Martin Wahl

Publications and source records attributed to Martin Wahl.

At least 19 recordsLinked to original sources

Higher-order spectral perturbation expansions II: Kernel matrices and manifold learning

We study spectral concentration bounds for kernel matrices as approximation of the corresponding kernel integral operator. Results are established under weak assumptions on the data setting and the reproducing kernel relying only on a Mercer condition and a local Weyl law. This allows us to deal with key features of kernel matrices, such as large multiplicities, large effective dimension, and heavy-tailed distributions. Our results apply to infinite dimensional principal component analysis, manifold learning, and Bayesian nonparametric statistics. We illustrate this via two prototypical examples: The heat kernel on the sphere and a wavelet prior from Bayesian nonparametrics.

math.ST

On empirical Hodge Laplacians under the manifold hypothesis

Given i.i.d. observations uniformly distributed on a closed submanifold of the Euclidean space, we study higher-order generalizations of graph Laplacians, so-called Hodge Laplacians on graphs, as approximations of the Laplace-Beltrami operator on differential forms. Our main result is a high-probability error bound for the associated Dirichlet forms. This bound improves existing Dirichlet form error bounds for graph Laplacians in the context of Laplacian Eigenmaps, and it provides insights into the Betti numbers studied in topological data analysis and the complementing positive part of the spectrum.

math.ST

Concentration and moment inequalities for sums of independent heavy-tailed random matrices

We prove Fuk-Nagaev and Rosenthal-type inequalities for sums of independent random matrices, focusing on the situation when the norms of the matrices possess finite moments of only low orders. Our bounds depend on the ``intrinsic'' dimensional characteristics such as the effective rank, as opposed to the dimension of the ambient space. We illustrate the advantages of such results through several applications, including new moment inequalities for sample covariance matrices and their eigenvectors when the underlying distribution is heavy-tailed. Moreover, we demonstrate that our techniques yield sharpened versions of moment inequalities for empirical processes.

math.PR

A kernel-based analysis of Laplacian Eigenmaps

Given i.i.d. observations uniformly distributed on a closed manifold $\mathcal{M}\subseteq \mathbb{R}^p$, we study the spectral properties of the associated empirical graph Laplacian based on a Gaussian kernel. Our main results are non-asymptotic error bounds, showing that the eigenvalues and eigenspaces of the empirical graph Laplacian are close to the eigenvalues and eigenspaces of the Laplace-Beltrami operator of $\mathcal{M}$. In our analysis, we connect the empirical graph Laplacian to kernel principal component analysis, and consider the heat kernel of $\mathcal{M}$ as reproducing kernel feature map. This leads to novel points of view and allows to leverage results for empirical covariance operators in infinite dimensions.

math.ST

A note on the prediction error of principal component regression in high dimensions

We analyze the prediction error of principal component regression (PCR) and prove high probability bounds for the corresponding squared risk conditional on the design. Our first main result shows that PCR performs comparably to the oracle method obtained by replacing empirical principal components by their population counterparts, provided that an effective rank condition holds. On the other hand, if the latter condition is violated, then empirical eigenvalues start to have a significant upward bias, resulting in a self-induced regularization of PCR. Our approach relies on the behavior of empirical eigenvalues, empirical eigenvectors and the excess risk of principal component analysis in high-dimensional regimes.

math.ST

Optimal parameter estimation for linear SPDEs from multiple measurements

The coefficients in a second order parabolic linear stochastic partial differential equation (SPDE) are estimated from multiple spatially localised measurements. Assuming that the spatial resolution tends to zero and the number of measurements is non-decreasing, the rate of convergence for each coefficient depends on its differential order and is faster for higher order coefficients. Based on an explicit analysis of the reproducing kernel Hilbert space of a general stochastic evolution equation, a Gaussian lower bound scheme is introduced. As a result, minimax optimality of the rates as well as sufficient and necessary conditions for consistent estimation are established.

math.ST

Quantitative limit theorems and bootstrap approximations for empirical spectral projectors

Given finite i.i.d.~samples in a Hilbert space with zero mean and trace-class covariance operator $\Sigma$, the problem of recovering the spectral projectors of $\Sigma$ naturally arises in many applications. In this paper, we consider the problem of finding distributional approximations of the spectral projectors of the empirical covariance operator $\hat \Sigma$, and offer a dimension-free framework where the complexity is characterized by the so-called relative rank of $\Sigma$. In this setting, novel quantitative limit theorems and bootstrap approximations are presented subject only to mild conditions in terms of moments and spectral decay. In many cases, these even improve upon existing results in a Gaussian setting.

math.PR

Relative perturbation bounds with applications to empirical covariance operators

The goal of this paper is to establish relative perturbation bounds, tailored for empirical covariance operators. Our main results are expansions for empirical eigenvalues and spectral projectors, leading to concentration inequalities and limit theorems. One of the key ingredients is a specific separation measure for population eigenvalues, which we call the relative rank, giving rise to a sharp invariance principle in terms of limit theorems, concentration inequalities and inconsistency results. Our framework is very general, requiring only $p > 4$ moments and allows for a huge variety of dependence structures.

math.PR

Functional estimation in log-concave location families

Let $\{P_θ:θ\in {\mathbb R}^d\}$ be a log-concave location family with $P_θ(dx)=e^{-V(x-θ)}dx,$ where $V:{\mathbb R}^d\mapsto {\mathbb R}$ is a known convex function and let $X_1,\dots, X_n$ be i.i.d. r.v. sampled from distribution $P_θ$ with an unknown location parameter $θ.$ The goal is to estimate the value $f(θ)$ of a smooth functional $f:{\mathbb R}^d\mapsto {\mathbb R}$ based on observations $X_1,\dots, X_n.$ In the case when $V$ is sufficiently smooth and $f$ is a functional from a ball in a Hölder space $C^s,$ we develop estimators of $f(θ)$ with minimax optimal error rates measured by the $L_2({\mathbb P}_θ)$-distance as well as by more general Orlicz norm distances. Moreover, we show that if $d\leq n^α$ and $s>\frac{1}{1-α},$ then the resulting estimators are asymptotically efficient in Hájek-LeCam sense with the convergence rate $\sqrt{n}.$ This generalizes earlier results on estimation of smooth functionals in Gaussian shift models. The estimators have the form $f_k(\hat θ),$ where $\hat θ$ is the maximum likelihood estimator and $f_k: {\mathbb R}^d\mapsto {\mathbb R}$ (with $k$ depending on $s$) are functionals defined in terms of $f$ and designed to provide a higher order bias reduction in functional estimation problem. The method of bias reduction is based on iterative parametric bootstrap and it has been successfully used before in the case of Gaussian models.

math.ST

Lower bounds for invariant statistical models with applications to principal component analysis

This paper develops nonasymptotic information inequalities for the estimation of the eigenspaces of a covariance operator. These results generalize previous lower bounds for the spiked covariance model, and they show that recent upper bounds for models with decaying eigenvalues are sharp. The proof relies on lower bound techniques based on group invariance arguments which can also deal with a variety of other statistical models.

math.ST

Nonparametric goodness-of-fit testing for parametric covariate models in pharmacometric analyses

The characterization of covariate effects on model parameters is a crucial step during pharmacokinetic/pharmacodynamic analyses. While covariate selection criteria have been studied extensively, the choice of the functional relationship between covariates and parameters, however, has received much less attention. Often, a simple particular class of covariate-to-parameter relationships (linear, exponential, etc.) is chosen ad hoc or based on domain knowledge, and a statistical evaluation is limited to the comparison of a small number of such classes. Goodness-of-fit testing against a nonparametric alternative provides a more rigorous approach to covariate model evaluation, but no such test has been proposed so far. In this manuscript, we derive and evaluate nonparametric goodness-of-fit tests for parametric covariate models, the null hypothesis, against a kernelized Tikhonov regularized alternative, transferring concepts from statistical learning to the pharmacological setting. The approach is evaluated in a simulation study on the estimation of the age-dependent maturation effect on the clearance of a monoclonal antibody. Scenarios of varying data sparsity and residual error are considered. The goodness-of-fit test correctly identified misspecified parametric models with high power for relevant scenarios. The case study provides proof-of-concept of the feasibility of the proposed approach, which is envisioned to be beneficial for applications that lack well-founded covariate models.

stat.ME

Analyzing the discrepancy principle for kernelized spectral filter learning algorithms

We investigate the construction of early stopping rules in the nonparametric regression problem where iterative learning algorithms are used and the optimal iteration number is unknown. More precisely, we study the discrepancy principle, as well as modifications based on smoothed residuals, for kernelized spectral filter learning algorithms including gradient descent. Our main theoretical bounds are oracle inequalities established for the empirical estimation error (fixed design), and for the prediction error (random design). From these finite-sample bounds it follows that the classical discrepancy principle is statistically adaptive for slow rates occurring in the hard learning scenario, while the smoothed discrepancy principles are adaptive over ranges of faster rates (resp. higher smoothness parameters). Our approach relies on deviation inequalities for the stopping rules in the fixed design setting, combined with change-of-norm arguments to deal with the random design setting.

math.ST

Statistical inference in sparse high-dimensional additive models

In this paper we discuss the estimation of a nonparametric component $f_1$ of a nonparametric additive model $Y=f_1(X_1) + ...+ f_q(X_q) + ε$. We allow the number $q$ of additive components to grow to infinity and we make sparsity assumptions about the number of nonzero additive components. We compare this estimation problem with that of estimating $f_1$ in the oracle model $Z= f_1(X_1) + ε$, for which the additive components $f_2,\dots,f_q$ are known. We construct a two-step presmoothing-and-resmoothing estimator of $f_1$ and state finite-sample bounds for the difference between our estimator and some smoothing estimators $\hat f_1^{\text{(oracle)}}$ in the oracle model. In an asymptotic setting these bounds can be used to show asymptotic equivalence of our estimator and the oracle estimators; the paper thus shows that, asymptotically, under strong enough sparsity conditions, knowledge of $f_2,\dots,f_q$ has no effect on estimation accuracy. Our first step is to estimate $f_1$ with an undersmoothed estimator based on near-orthogonal projections with a group Lasso bias correction. We then construct pseudo responses $\hat Y$ by evaluating a debiased modification of our undersmoothed estimator of $f_1$ at the design points. In the second step the smoothing method of the oracle estimator $\hat f_1^{\text{(oracle)}}$ is applied to a nonparametric regression problem with responses $\hat Y$ and covariates $X_1$. Our mathematical exposition centers primarily on establishing properties of the presmoothing estimator. We present simulation results demonstrating close-to-oracle performance of our estimator in practical applications.

math.ST

On the perturbation series for eigenvalues and eigenprojections

A standard perturbation result states that perturbed eigenvalues and eigenprojections admit a perturbation series provided that the operator norm of the perturbation is smaller than a constant times the corresponding eigenvalue isolation distance. In this paper, we show that the same holds true under a weighted condition, where the perturbation is symmetrically normalized by the square-root of the reduced resolvent. This weighted condition originates in random perturbations where it leads to significant improvements.

math.FA

A note on the prediction error of principal component regression

We analyse the prediction error of principal component regression (PCR) and prove non-asymptotic upper bounds for the corresponding squared risk. Under mild assumptions, we show that PCR performs as well as the oracle method obtained by replacing empirical principal components by their population counterparts. Our approach relies on upper bounds for the excess risk of principal component analysis.

math.ST

Non-asymptotic upper bounds for the reconstruction error of PCA

We analyse the reconstruction error of principal component analysis (PCA) and prove non-asymptotic upper bounds for the corresponding excess risk. These bounds unify and improve existing upper bounds from the literature. In particular, they give oracle inequalities under mild eigenvalue conditions. The bounds reveal that the excess risk differs significantly from usually considered subspace distances based on canonical angles. Our approach relies on the analysis of empirical spectral projectors combined with concentration inequalities for weighted empirical covariance operators and empirical eigenvalues.

math.ST