SearcharxivSearch

arXiv subjects

Siegfried Hörmann

Publications and source records attributed to Siegfried Hörmann.

At least 19 recordsLinked to original sources

Inference for sparsely sampled Gauss-Markov processes under outcome-dependent dropout

We consider sparsely observed functional data generated by a latent Gauss-Markov process modeled through a linear stochastic differential equation (SDE). Motivated by an application to tumor growth data, we allow for outcome-dependent dropout, where sampling terminates once the most recent observation exceeds a prescribed threshold. Such observation schemes arise naturally in longitudinal studies and violate the fundamental missing completely at random assumption commonly imposed in the analysis of partially observed functional data. We develop a likelihood-based estimation framework that, in contrast to existing moment-based methods, avoids the bias induced by outcome-dependent dropout. We establish convergence rates for the proposed estimators and show that they are mini-max optimal up to logarithmic factors, with the diffusion coefficient admitting a faster rate than the drift coefficients under a random initial condition for the SDE. An application to tumor growth data further demonstrates the advantages of a fully probabilistic functional data model for tasks beyond second-order inference, including prediction bands, first-passage times, and tumor-age estimation.

math.ST

Testing the Missing Completely at Random Assumption for Functional Data

We consider functional data which have only been observed on a subset of their domain. This paper aims to develop statistical tests to determine whether the function and the domain over which it is observed are independent. The assumption that data are missing completely at random (MCAR) is essential for many functional data methods handling incomplete observations. However, no general testing procedures have been established to validate this assumption. We address this critical gap by introducing a testing framework which is generally based on a partition of the observation patterns. Besides deterministic partitions, we also consider a systematic approach based on clustering. We establish asymptotic results for our tests and illustrate the methodology in several real data applications.

stat.ME

Quantifying and testing dependence to categorical variables

We suggest a dependence coefficient between a categorical variable and some general variable taking values in a metric space. We derive important theoretical properties and study the large sample behaviour of our suggested estimator. Moreover, we develop an independence test which has an asymptotic $χ^2$-distribution if the variables are independent and prove that this test is consistent against any violation of independence. The test is also applicable to the classical~$K$-sample problem with possibly high- or infinite-dimensional distributions. We discuss some extensions, including a variant of the coefficient for measuring conditional dependence.

math.ST

Azadkia-Chatterjee's dependence coefficient for infinite dimensional data

We extend the scope of Azadkia-Chatterjee's dependence coefficient between a scalar response $Y$ and a multivariate covariate $X$ to the case where $X$ takes values in a general metric space. Particular attention is paid to the case where $X$ is a curve. Although extending this framework at the population level is relatively straightforward, analyzing the asymptotic behavior of the estimator proves to be complex. This complexity is largely related to the nearest neighbor structure of the infinite-dimensional covariate sample, leading us to explore a topic that has not been previously addressed in the literature. The primary contribution of this paper is to provide insights into this issue and propose strategies to address it. Our findings also have significant implications for other graph-based methods facing similar challenges.

math.ST

Covariate-informed reconstruction of partially observed functional data via factor models

This paper studies linear reconstruction of partially observed functional data which are recorded on a discrete grid. We propose a novel estimation approach based on approximate factor models with increasing rank taking into account potential covariate information. Whereas alternative reconstruction procedures commonly involve some preliminary smoothing, our method separates the signal from noise and reconstructs missing fragments at once. We establish uniform convergence rates of our estimator and introduce a new method for constructing simultaneous prediction bands for the missing trajectories. A simulation study examines the performance of the proposed methods in finite samples. Finally, a real data application of temperature curves demonstrates that our theory provides a simple and effective method to recover missing fragments.

math.ST

Prediction in functional regression with discretely observed and noisy covariates

In practice functional data are sampled on a discrete set of observation points and often susceptible to noise. We consider in this paper the setting where such data are used as explanatory variables in a regression problem. If the primary goal is prediction, we show that the gain by embedding the problem into a scalar-on-function regression is limited. Instead we impose a factor model on the predictors and suggest regressing the response on an appropriate number of factor scores. This approach is shown to be consistent under mild technical assumptions, numerically efficient and gives good practical performance in both simulations as well as real data settings.

stat.ME

Preprocessing noisy functional data: a multivariate perspective

We consider functional data which are measured on a discrete set of observation points. Often such data are measured with additional noise. We explore in this paper the factor structure underlying this type of data. We show that the latent signal can be attributed to the common components of a corresponding factor model and can be estimated accordingly, by borrowing methods from factor model literature. We also show that principal components, which play a key role in functional data analysis, can be accurately estimated after taking such a multivariate instead of a `functional' perspective. In addition to the estimation problem, we also address testing of the null-hypothesis of iid noise. While this assumption is largely prevailing in the literature, we believe that it is often unrealistic and not supported by a residual analysis.

stat.ME

Consistently recovering the signal from noisy functional data

In practice most functional data cannot be recorded on a continuum, but rather at discrete time points. It is also quite common that these measurements come with an additive error, which one would like eliminate for the statistical analysis. When the measurements for each functional datum are taken on the same grid, the underlying signal-plus-noise model can be viewed as a factor model. The signals refer to the common components of the factor model, the noise is related to the idiosyncratic components. We formulate a framework which allows to consistently recover the signal by a PCA based factor model estimation scheme. Our theoretical results hold under rather mild conditions, in particular we don't require specific smoothness assumptions for the underlying curves and allow for a certain degree of autocorrelation in the noise.

math.ST

Testing normality of spatially indexed functional data

We develop a test of normality for spatially indexed functions. The assumption of normality is common in spatial statistics, yet no significance tests, or other means of assessment, have been available for functional data. This paper aims at filling this gap in the case of functional observations on a spatial grid. Our test compares the moments of the spatial (frequency domain) principal component scores to those of a suitable Gaussian distribution. Critical values can be readily obtained from a chi-squared distribution. We provide rigorous theoretical justification for a broad class of weakly stationary functional random fields. We perform simulation studies to assess the the power of the test against various alternatives. An application to Surface Incoming Shortwave Radiation illustrates the practical value of this procedure.

stat.ME

Estimating the conditional distribution in functional regression problems

We consider the problem of consistently estimating the conditional distribution $P(Y \in A |X)$ of a functional data object $Y=(Y(t): t\in[0,1])$ given covariates $X$ in a general space, assuming that $Y$ and $X$ are related by a functional linear regression model. Two natural estimation methods are proposed, based on either bootstrapping the estimated model residuals, or fitting functional parametric models to the model residuals and estimating $P(Y \in A |X)$ via simulation. Whether either of these methods lead to consistent estimation depends on the consistency properties of the regression operator estimator, and the space within which $Y$ is viewed. We show that under general consistency conditions on the regression operator estimator, which hold for certain functional principal component based estimators, consistent estimation of the conditional distribution can be achieved, both when $Y$ is an element of a separable Hilbert space, and when $Y$ is an element of the Banach space of continuous functions. The latter results imply that sets $A$ that specify path properties of $Y$, which are of interest in applications, can be considered. The proposed methods are studied in several simulation experiments, and data analyses of electricity price and pollution curves.

math.ST

The maximum of the periodogram of Hilbert space valued time series

We are interested to detect periodic signals in Hilbert space valued time series when the length of the period is unknown. A natural test statistic is the maximum Hilbert-Schmidt norm of the periodogram operator over all fundamental frequencies. In this paper we analyze the asymptotic distribution of this test statistic. We consider the case where the noise variables are independent and then generalize our results to functional linear processes. Details for implementing the test are provided for the class of functional autoregressive processes. We illustrate the usefulness of our approach by examining air quality data from Graz, Austria. The accuracy of the asymptotic theory in finite samples is evaluated in a simulation experiment.

math.ST

Detection of periodicity in functional time series

We derive several tests for the presence of a periodic component in a time series of functions. We consider both the traditional setting in which the periodic functional signal is contaminated by functional white noise, and a more general setting of a contaminating process which is weakly dependent. Several forms of the periodic component are considered. Our tests are motivated by the likelihood principle and fall into two broad categories, which we term multivariate and fully functional. Overall, for the functional series that motivate this research, the fully functional tests exhibit a superior balance of size and power. Asymptotic null distributions of all tests are derived and their consistency is established. Their finite sample performance is examined and compared by numerical studies and application to pollution data.

stat.ME

On the CLT for discrete Fourier transforms of functional time series

We consider a strictly stationary and ergodic sequence of random elements taking values in some Hilbert space. Our target is to study the weak convergence of the discrete Fourier transforms under sharp conditions. As a side-result we obtain the regular CLT for partial sums under mild assumptions.

math.PR

Dynamic Functional Principal Component

In this paper, we address the problem of dimension reduction for time series of functional data $(X_t\colon t\in\mathbb{Z})$. Such {\it functional time series} frequently arise, e.g., when a continuous-time process is segmented into some smaller natural units, such as days. Then each~$X_t$ represents one intraday curve. We argue that functional principal component analysis (FPCA), though a key technique in the field and a benchmark for any competitor, does not provide an adequate dimension reduction in a time-series setting. FPCA indeed is a {\it static} procedure which ignores the essential information provided by the serial dependence structure of the functional data under study. Therefore, inspired by Brillinger's theory of {\it dynamic principal components}, we propose a {\it dynamic} version of FPCA, which is based on a frequency-domain approach. By means of a simulation study and an empirical illustration, we show the considerable improvement the dynamic approach entails when compared to the usual static procedure.

math.ST

A note on estimation in Hilbertian linear models

We study estimation and prediction in linear models where the response and the regressor variable both take values in some Hilbert space. Our main objective is to obtain consistency of a principal components based estimator for the regression operator under minimal assumptions. In particular, we avoid some inconvenient technical restrictions that have been used throughout the literature. We develop our theory in a time dependent setup which comprises as important special case the autoregressive Hilbertian model.

math.ST

On the prediction of stationary functional time series

This paper addresses the prediction of stationary functional time series. Existing contributions to this problem have largely focused on the special case of first-order functional autoregressive processes because of their technical tractability and the current lack of advanced functional time series methodology. It is shown here how standard multivariate prediction techniques can be utilized in this context. The connection between functional and multivariate predictions is made precise for the important case of vector and functional autoregressions. The proposed method is easy to implement, making use of existing statistical software packages, and may therefore be attractive to a broader, possibly non-academic, audience. Its practical applicability is enhanced through the introduction of a novel functional final prediction error model selection criterion that allows for an automatic determination of the lag structure and the dimensionality of the model. The usefulness of the proposed methodology is demonstrated in a simulation study and an application to environmental data, namely the prediction of daily pollution curves describing the concentration of particulate matter in ambient air. It is found that the proposed prediction method often significantly outperforms existing methods.

stat.ME

Consistency of the mean and the principal components of spatially distributed functional data

This paper develops a framework for the estimation of the functional mean and the functional principal components when the functions form a random field. More specifically, the data we study consist of curves $X(\mathbf{s}_k;t),t\in[0,T]$, observed at spatial points $\mathbf{s}_1,\mathbf{s}_2,\ldots,\mathbf{s}_N$. We establish conditions for the sample average (in space) of the $X(\mathbf{s}_k)$ to be a consistent estimator of the population mean function, and for the usual empirical covariance operator to be a consistent estimator of the population covariance operator. These conditions involve an interplay of the assumptions on an appropriately defined dependence between the functions $X(\mathbf{s}_k)$ and the assumptions on the spatial distribution of the points $\mathbf{s}_k$. The rates of convergence may be the same as for i.i.d. functional samples, but generally depend on the strength of dependence and appropriately quantified distances between the points $\mathbf{s}_k$. We also formulate conditions for the lack of consistency.

math.ST

Split invariance principles for stationary processes

The results of Komlós, Major and Tusnády give optimal Wiener approximation of partial sums of i.i.d. random variables and provide an extremely powerful tool in probability and statistical inference. Recently Wu [Ann. Probab. 35 (2007) 2294--2320] obtained Wiener approximation of a class of dependent stationary processes with finite $p$th moments, $2 0$, and Liu and Lin [Stochastic Process. Appl. 119 (2009) 249--280] removed the logarithmic factor, reaching the Komlós--Major--Tusnády bound $o(n^{1/p})$. No similar results exist for $p>4$, and in fact, no existing method for dependent approximation yields an a.s. rate better than $o(n^{1/4})$. In this paper we show that allowing a second Wiener component in the approximation, we can get rates near to $o(n^{1/p})$ for arbitrary $p>2$. This extends the scope of applications of the results essentially, as we illustrate it by proving new limit theorems for increments of stochastic processes and statistical tests for short term (epidemic) changes in stationary processes. Our method works under a general weak dependence condition covering wide classes of linear and nonlinear time series models and classical dynamical systems.

math.PR