SearcharxivSearch

arXiv subjects

Markus Reiß

Publications and source records attributed to Markus Reiß.

At least 19 recordsLinked to original sources

Nonparametric Diffusivity Estimation for the Stochastic Heat Equation from Noisy Observations

We estimate nonparametrically the spatially varying diffusivity of a stochastic heat equation from observations perturbed by additional noise. To that end, we employ a two-step localization procedure, more precisely, we combine local state estimates into a locally linear regression approach. Our analysis relies on quantitative Trotter--Kato type approximation results for the heat semigroup that are of independent interest. The presence of observational noise leads to non-standard scaling behaviour of the model. Numerical simulations illustrate the results.

math.ST

Comparing regularisation paths of (conjugate) gradient estimators in ridge regression

We consider standard gradient descent, gradient flow and conjugate gradients as iterative algorithms for minimising a penalised ridge criterion in linear regression. While it is well known that conjugate gradients exhibit fast numerical convergence, the statistical properties of their iterates are more difficult to assess due to inherent non-linearities and dependencies. On the other hand, standard gradient flow is a linear method with well-known regularising properties when stopped early. By an explicit non-standard error decomposition we are able to bound the prediction error for conjugate gradient iterates by a corresponding prediction error of gradient flow at transformed iteration indices. This way, the risk along the entire regularisation path of conjugate gradient iterations can be compared to that for regularisation paths of standard linear methods like gradient flow and ridge regression. In particular, the oracle conjugate gradient iterate shares the optimality properties of the gradient flow and ridge regression oracles up to a constant factor. Numerical examples show the similarity of the regularisation paths in practice.

stat.ML

Rank tests for time-varying covariance matrices observed under noise

We consider a $d$-dimensional continuous martingale $X(t)$ with quadratic variation matrix $\langle X\rangle_t=\int_0^t Σ(s)\,ds$ and develop tests for the rank of its spot covariance matrix $Σ(t)$, $t\in[0,1]$. The process $X$ is observed under observational noise, as is standard for microstructure noise models in high-frequency finance. We test the null hypothesis ${\mathcal H}_0:rank(Σ(t))\le r$ against local alternatives ${\mathcal H}_{1,n}:λ_{r+1}(Σ(t))\ge v_n$, where $λ_{r+1}$ denotes the $(r+1)$st eigenvalue and $v_n\downarrow 0$ as the sample size $n\to\infty$. We construct test statistics based on eigenvalues of carefully calibrated localized spectral covariance matrix estimates. Critical values are provided non-asymptotically as well as asymptotically via maximal eigenvalues of Gaussian orthogonal ensembles. The power analysis establishes asymptotic consistency for a separation rate $v_n\thicksim (\underlineλ_r^{-1/(β+1)}n^{-β/(β+1)})\wedge n^{-β/(β+2)}$, depending on the Hölder-regularity $β$ of $Σ$ and a possible spectral gap $\underlineλ_r\ge 0$ under ${\mathcal H}_0$. A lower bound shows the optimality of this rate. We discuss why the rate is much faster than conventional estimation rates. The theory is illustrated by simulations and a real data example with German government bonds of varying maturity.

math.ST

Early Stopping for Regression Trees

We develop early stopping rules for growing regression tree estimators. The fully data-driven stopping rule is based on monitoring the global residual norm. The best-first search and the breadth-first search algorithms together with linear interpolation give rise to generalized projection or regularization flows. A general theory of early stopping is established. Oracle inequalities for the early-stopped regression tree are derived without any smoothness assumption on the regression function, assuming the original CART splitting rule, yet with a much broader scope. The remainder terms are of smaller order than the best achievable rates for Lipschitz functions in dimension $d\ge 2$. In real and synthetic data the early stopping regression tree estimators attain the statistical performance of cost-complexity pruning while significantly reducing computational costs.

math.ST

Information bounds for inference in stochastic evolution equations observed under noise

We consider statistics for stochastic evolution equations in Hilbert space with emphasis on stochastic partial differential equations (SPDEs). We observe a solution process under additional measurement errors and want to estimate a real or functional parameter in the drift. Main targets of estimation are the diffusivity, transport or source coefficient in a parabolic SPDE. By bounding the Hellinger distance between observation laws under different parameters we derive lower bounds on the estimation error, which reveal the underlying information structure. The estimation rates depend on the measurement noise level, the observation time, the covariance of the dynamic noise, the dimension and the order, at which the parametrised coefficient appears in the differential operator. A general estimation procedure attains these rates in many parametric cases and proves their minimax optimality. For nonparametric estimation problems, where the parameter is an unknown function, the lower bounds exhibit an even more complex information structure. The proofs are to a large extent based on functional calculus, perturbation theory and monotonicity of the semigroup generators.

math.ST

Early stopping for conjugate gradients in statistical inverse problems

We consider estimators obtained by iterates of the conjugate gradient (CG) algorithm applied to the normal equation of prototypical statistical inverse problems. Stopping the CG algorithm early induces regularisation, and optimal convergence rates of prediction and reconstruction error are established in wide generality for an ideal oracle stopping time. Based on this insight, a fully data-driven early stopping rule $τ$ is constructed, which also attains optimal rates, provided the error in estimating the noise level is not dominant. The error analysis of CG under statistical noise is subtle due to its nonlinear dependence on the observations. We provide an explicit error decomposition and identify two terms in the prediction error, which share important properties of classical bias and variance terms. Together with a continuous interpolation between CG iterates, this paves the way for a comprehensive error analysis of early stopping. In particular, a general oracle-type inequality is proved for the prediction error at $τ$. For bounding the reconstruction error, a more refined probabilistic analysis, based on concentration of self-normalised Gaussian processes, is developed. The methodology also provides some new insights into early stopping for CG in deterministic inverse problems. A numerical study for standard examples shows good results in practice for early stopping at $τ$.

math.ST

Change point estimation for a stochastic heat equation

We study a change point model based on a stochastic partial differential equation (SPDE) corresponding to the heat equation governed by the weighted Laplacian $Δ_\vartheta = \nabla\vartheta\nabla$, where $\vartheta=\vartheta(x)$ is a space-dependent diffusivity. As a basic problem the domain $(0,1)$ is considered with a piecewise constant diffusivity with a jump at an unknown point $τ$. Based on local measurements of the solution in space with resolution $δ$ over a finite time horizon, we construct a simultaneous M-estimator for the diffusivity values and the change point. The change point estimator converges at rate $δ$, while the diffusivity constants can be recovered with convergence rate $δ^{3/2}$. Moreover, when the diffusivity parameters are known and the jump height vanishes with the spatial resolution tending to zero, we derive a limit theorem for the change point estimator and identify the limiting distribution. For the mathematical analysis, a precise understanding of the SPDE with discontinuous $\vartheta$, tight concentration bounds for quadratic functionals in the solution, and a generalisation of classical M-estimators are developed.

math.ST

Parameter estimation for the stochastic heat equation with multiplicative noise from local measurements

For the stochastic heat equation with multiplicative noise we consider the problem of estimating the diffusivity parameter in front of the Laplace operator. Based on local observations in space, we first study an estimator that was derived for additive noise. A stable central limit theorem shows that this estimator is consistent and asymptotically mixed normal. By taking into account the quadratic variation, we propose two new estimators. Their limiting distributions exhibit a smaller (conditional) variance and the last estimator also works for vanishing noise levels. The proofs are based on local approximation results to overcome the intricate nonlinearities and on a stable central limit theorem for stochastic integrals with respect to cylindrical Brownian motion. Simulation results illustrate the theoretical findings.

math.ST

Estimation for the reaction term in semi-linear SPDEs under small diffusivity

We consider the estimation of a non-linear reaction term in the stochastic heat or more generally in a semi-linear stochastic partial differential equation (SPDE). Consistent inference is achieved by studying a small diffusivity level, which is realistic in applications. Our main result is a central limit theorem for the estimation error of a parametric estimator, from which confidence intervals can be constructed. Statistical efficiency is demonstrated by establishing local asymptotic normality. The estimation method is extended to local observations in time and space, which allows for non-parametric estimation of a reaction intensity varying in time and space. Furthermore, discrete observations in time and space can be handled. The statistical analysis requires advanced tools from stochastic analysis like Malliavin calculus for SPDEs, the infinite-dimensional Gaussian Poincaré inequality and regularity results for SPDEs in $L^p$-interpolation spaces.

math.ST

Inference on the maximal rank of time-varying covariance matrices using high-frequency data

We study the rank of the instantaneous or spot covariance matrix $Σ_X(t)$ of a multidimensional continuous semi-martingale $X(t)$. Given high-frequency observations $X(i/n)$, $i=0,\ldots,n$, we test the null hypothesis $rank(Σ_X(t))\le r$ for all $t$ against local alternatives where the average $(r+1)$st eigenvalue is larger than some signal detection rate $v_n$. A major problem is that the inherent averaging in local covariance statistics produces a bias that distorts the rank statistics. We show that the bias depends on the regularity and a spectral gap of $Σ_X(t)$. We establish explicit matrix perturbation and concentration results that provide non-asymptotic uniform critical values and optimal signal detection rates $v_n$. This leads to a rank estimation method via sequential testing. For a class of stochastic volatility models, we determine data-driven critical values via normed p-variations of estimated local covariance matrices. The methods are illustrated by simulations and an application to high-frequency data of U.S. government bonds.

math.ST

Parameter Estimation in an SPDE Model for Cell Repolarisation

As a concrete setting where stochastic partial differential equations (SPDEs) are able to model real phenomena, we propose a stochastic Meinhardt model for cell repolarisation and study how parameter estimation techniques developed for simple linear SPDE models apply in this situation. We establish the existence of mild SPDE solutions and we investigate the impact of the driving noise process on pattern formation in the solution. We then pursue estimation of the diffusion term and show asymptotic normality for our estimator as the space resolution becomes finer. The finite sample performance is investigated for synthetic and real data.

math.ST

Nonparametric estimation for linear SPDEs from local measurements

The coefficient function of the leading differential operator is estimated from observations of a linear stochastic partial differential equation (SPDE). The estimation is based on continuous time observations which are localised in space. For the asymptotic regime with fixed time horizon and with the spatial resolution of the observations tending to zero, we provide rate-optimal estimators and establish scaling limits of the deterministic PDE and of the SPDE on growing domains. The estimators are robust to lower order perturbations of the underlying differential operator and achieve the parametric rate even in the nonparametric setup with a spatially varying coefficient. A numerical example illustrates the main results.

math.ST

Non-asymptotic upper bounds for the reconstruction error of PCA

We analyse the reconstruction error of principal component analysis (PCA) and prove non-asymptotic upper bounds for the corresponding excess risk. These bounds unify and improve existing upper bounds from the literature. In particular, they give oracle inequalities under mild eigenvalue conditions. The bounds reveal that the excess risk differs significantly from usually considered subspace distances based on canonical angles. Our approach relies on the analysis of empirical spectral projectors combined with concentration inequalities for weighted empirical covariance operators and empirical eigenvalues.

math.ST

Functional estimation and hypothesis testing in nonparametric boundary models

Consider a Poisson point process with unknown support boundary curve $g$, which forms a prototype of an irregular statistical model. We address the problem of estimating non-linear functionals of the form $\int Φ(g(x))\,dx$. Following a nonparametric maximum-likelihood approach, we construct an estimator which is UMVU over Hölder balls and achieves the (local) minimax rate of convergence. These results hold under weak assumptions on $Φ$ which are satisfied for $Φ(u)=|u|^p$, $p\ge 1$. As an application, we consider the problem of estimating the $L^p$-norm and derive the minimax separation rates in the corresponding nonparametric hypothesis testing problem. Structural differences to results for regular nonparametric models are discussed.

math.ST

Early stopping for statistical inverse problems via truncated SVD estimation

We consider truncated SVD (or spectral cut-off, projection) estimators for a prototypical statistical inverse problem in dimension $D$. Since calculating the singular value decomposition (SVD) only for the largest singular values is much less costly than the full SVD, our aim is to select a data-driven truncation level $\widehat m\in\{1,\ldots,D\}$ only based on the knowledge of the first $\widehat m$ singular values and vectors. We analyse in detail whether sequential {\it early stopping} rules of this type can preserve statistical optimality. Information-constrained lower bounds and matching upper bounds for a residual based stopping rule are provided, which give a clear picture in which situation optimal sequential adaptation is feasible. Finally, a hybrid two-step approach is proposed which allows for classical oracle inequalities while considerably reducing numerical complexity.

math.ST

Wasserstein and total variation distance between marginals of Lévy processes

We present upper bounds for the Wasserstein distance of order $p$ between the marginals of Lévy processes, including Gaussian approximations for jumps of infinite activity. Using the convolution structure, we further derive upper bounds for the total variation distance between the marginals of Lévy processes. Connections to other metrics like Zolotarev and Toscani-Fourier distances are established. The theory is illustrated by concrete examples and an application to statistical lower bounds.

math.PR

Optimal adaptation for early stopping in statistical inverse problems

For linear inverse problems $Y=\mathsf{A}μ+ξ$, it is classical to recover the unknown signal $μ$ by iterative regularisation methods $(\widehat μ^{(m)}, m=0,1,\ldots)$ and halt at a data-dependent iteration $τ$ using some stopping rule, typically based on a discrepancy principle, so that the weak (or prediction) squared-error $\|\mathsf{A}(\widehat μ^{(τ)}-μ)\|^2$ is controlled. In the context of statistical estimation with stochastic noise $ξ$, we study oracle adaptation (that is, compared to the best possible stopping iteration) in strong squared-error $E[\|\hat μ^{(τ)}-μ\|^2]$. For a residual-based stopping rule oracle adaptation bounds are established for general spectral regularisation methods. The proofs use bias and variance transfer techniques from weak prediction error to strong $L^2$-error, as well as convexity arguments and concentration bounds for the stochastic part. Adaptive early stopping for the Landweber method is studied in further detail and illustrated numerically.

math.ST

Estimating the Spot Covariation of Asset Prices - Statistical Theory and Empirical Evidence

We propose a new estimator for the spot covariance matrix of a multi-dimensional continuous semi-martingale log asset price process which is subject to noise and non-synchronous observations. The estimator is constructed based on a local average of block-wise parametric spectral covariance estimates. The latter originate from a local method of moments (LMM) which recently has been introduced. We prove consistency and a point-wise stable central limit theorem for the proposed spot covariance estimator in a very general setup with stochastic volatility, leverage effects and general noise distributions. Moreover, we extend the LMM estimator to be robust against autocorrelated noise and propose a method to adaptively infer the autocorrelations from the data. Based on simulations we provide empirical guidance on the effective implementation of the estimator and apply it to high-frequency data of a cross-section of Nasdaq blue chip stocks. Employing the estimator to estimate spot covariances, correlations and volatilities in normal but also unusual periods yields novel insights into intraday covariance and correlation dynamics. We show that intraday (co-)variations (i) follow underlying periodicity patterns, (ii) reveal substantial intraday variability associated with (co-)variation risk, and (iii) can increase strongly and nearly instantaneously if new information arrives.

math.ST