Searcharxiv⌕ Search

arXiv subjects

Peter Hall

Publications and source records attributed to Peter Hall.

68 records · Page 4Linked to original sources

Methodology and convergence rates for functional linear regression

In functional linear regression, the slope ``parameter'' is a function. Therefore, in a nonparametric context, it is determined by an infinite number of unknowns. Its estimation involves solving an ill-posed problem and has points of contact with a range of methodologies, including statistical smoothing and deconvolution. The standard approach to estimating the slope function is based explicitly on functional principal components analysis and, consequently, on spectral decomposition in terms of eigenvalues and eigenfunctions. We discuss this approach in detail and show that in certain circumstances, optimal convergence rates are achieved by the PCA technique. An alternative approach based on quadratic regularisation is suggested and shown to have advantages from some points of view.

math.ST↗

Nonparametric estimation when data on derivatives are available

We consider settings where data are available on a nonparametric function and various partial derivatives. Such circumstances arise in practice, for example in the joint estimation of cost and input functions in economics. We show that when derivative data are available, local averages can be replaced in certain dimensions by nonlocal averages, thus reducing the nonparametric dimension of the problem. We derive optimal rates of convergence and conditions under which dimension reduction is achieved. Kernel estimators and their properties are analyzed, although other estimators, such as local polynomial, spline and nonparametric least squares, may also be used. Simulations and an application to the estimation of electricity distribution costs are included.

math.ST↗

Prediction in functional linear regression

There has been substantial recent work on methods for estimating the slope function in linear regression for functional data analysis. However, as in the case of more conventional finite-dimensional regression, much of the practical interest in the slope centers on its application for the purpose of prediction, rather than on its significance in its own right. We show that the problems of slope-function estimation, and of prediction from an estimator of the slope function, have very different characteristics. While the former is intrinsically nonparametric, the latter can be either nonparametric or semiparametric. In particular, the optimal mean-square convergence rate of predictors is $n^{-1}$, where $n$ denotes sample size, if the predictand is a sufficiently smooth function. In other cases, convergence occurs at a polynomial rate that is strictly slower than $n^{-1}$. At the boundary between these two regimes, the mean-square convergence rate is less than $n^{-1}$ by only a logarithmic factor. More generally, the rate of convergence of the predicted value of the mean response in the regression model, given a particular value of the explanatory variable, is determined by a subtle interaction among the smoothness of the predictand, of the slope function in the model, and of the autocovariance function for the distribution of explanatory variables.

math.ST↗

Nonparametric estimation of mean-squared prediction error in nested-error regression models

Nested-error regression models are widely used for analyzing clustered data. For example, they are often applied to two-stage sample surveys, and in biology and econometrics. Prediction is usually the main goal of such analyses, and mean-squared prediction error is the main way in which prediction performance is measured. In this paper we suggest a new approach to estimating mean-squared prediction error. We introduce a matched-moment, double-bootstrap algorithm, enabling the notorious underestimation of the naive mean-squared error estimator to be substantially reduced. Our approach does not require specific assumptions about the distributions of errors. Additionally, it is simple and easy to apply. This is achieved through using Monte Carlo simulation to implicitly develop formulae which, in a more conventional approach, would be derived laboriously by mathematical arguments.

math.ST↗

Properties of principal component methods for functional and longitudinal data analysis

The use of principal component methods to analyze functional data is appropriate in a wide range of different settings. In studies of ``functional data analysis,'' it has often been assumed that a sample of random functions is observed precisely, in the continuum and without noise. While this has been the traditional setting for functional data analysis, in the context of longitudinal data analysis a random function typically represents a patient, or subject, who is observed at only a small number of randomly distributed points, with nonnegligible measurement error. Nevertheless, essentially the same methods can be used in both these cases, as well as in the vast number of settings that lie between them. How is performance affected by the sampling plan? In this paper we answer that question. We show that if there is a sample of $n$ functions, or subjects, then estimation of eigenvalues is a semiparametric problem, with root-$n$ consistent estimators, even if only a few observations are made of each function, and if each observation is encumbered by noise. However, estimation of eigenfunctions becomes a nonparametric problem when observations are sparse. The optimal convergence rates in this case are those which pertain to more familiar function-estimation settings. We also describe the effects of sampling at regularly spaced points, as opposed to random points. In particular, it is shown that there are often advantages in sampling randomly. However, even in the case of noisy data there is a threshold sampling rate (depending on the number of functions treated) above which the rate of sampling (either randomly or regularly) has negligible impact on estimator performance, no matter whether eigenfunctions or eigenvectors are being estimated.

math.ST↗

Assessing extrema of empirical principal component functions

The difficulties of estimating and representing the distributions of functional data mean that principal component methods play a substantially greater role in functional data analysis than in more conventional finite-dimensional settings. Local maxima and minima in principal component functions are of direct importance; they indicate places in the domain of a random function where influence on the function value tends to be relatively strong but of opposite sign. We explore statistical properties of the relationship between extrema of empirical principal component functions, and their counterparts for the true principal component functions. It is shown that empirical principal component funcions have relatively little trouble capturing conventional extrema, but can experience difficulty distinguishing a ``shoulder'' in a curve from a small bump. For example, when the true principal component function has a shoulder, the probability that the empirical principal component function has instead a bump is approximately equal to 1/2. We suggest and describe the performance of bootstrap methods for assessing the strength of extrema. It is shown that the subsample bootstrap is more effective than the standard bootstrap in this regard. A ``bootstrap likelihood'' is proposed for measuring extremum strength. Exploratory numerical methods are suggested.

math.ST↗

Nonparametric methods for inference in the presence of instrumental variables

We suggest two nonparametric approaches, based on kernel methods and orthogonal series to estimating regression functions in the presence of instrumental variables. For the first time in this class of problems, we derive optimal convergence rates, and show that they are attained by particular estimators. In the presence of instrumental variables the relation that identifies the regression function also defines an ill-posed inverse problem, the ``difficulty'' of which depends on eigenvalues of a certain integral operator which is determined by the joint density of endogenous and instrumental variables. We delineate the role played by problem difficulty in determining both the optimal convergence rate and the appropriate choice of smoothing parameter.

math.ST↗

Approximating conditional distribution functions using dimension reduction

Motivated by applications to prediction and forecasting, we suggest methods for approximating the conditional distribution function of a random variable Y given a dependent random d-vector X. The idea is to estimate not the distribution of Y|X, but that of Y|θ^TX, where the unit vector θis selected so that the approximation is optimal under a least-squares criterion. We show that θmay be estimated root-n consistently. Furthermore, estimation of the conditional distribution function of Y, given θ^TX, has the same first-order asymptotic properties that it would enjoy if θwere known. The proposed method is illustrated using both simulated and real-data examples, showing its effectiveness for both independent datasets and data from time series. Numerical work corroborates the theoretical result that θcan be estimated particularly accurately.

math.ST↗

Testing for monotone increasing hazard rate

A test of the null hypothesis that a hazard rate is monotone nondecreasing, versus the alternative that it is not, is proposed. Both the test statistic and the means of calibrating it are new. Unlike previous approaches, neither is based on the assumption that the null distribution is exponential. Instead, empirical information is used to effectively identify and eliminate from further consideration parts of the line where the hazard rate is clearly increasing; and to confine subsequent attention only to those parts that remain. This produces a test with greater apparent power, without the excessive conservatism of exponential-based tests. Our approach to calibration borrows from ideas used in certain tests for unimodality of a density, in that a bandwidth is increased until a distribution with the desired properties is obtained. However, the test statistic does not involve any smoothing, and is, in fact, based directly on an assessment of convexity of the distribution function, using the conventional empirical distribution. The test is shown to have optimal power properties in difficult cases, where it is called upon to detect a small departure, in the form of a bump, from monotonicity. More general theoretical properties of the test and its numerical performance are explored.

math.ST↗

Bandwidth choice for nonparametric classification

It is shown that, for kernel-based classification with univariate distributions and two populations, optimal bandwidth choice has a dichotomous character. If the two densities cross at just one point, where their curvatures have the same signs, then minimum Bayes risk is achieved using bandwidths which are an order of magnitude larger than those which minimize pointwise estimation error. On the other hand, if the curvature signs are different, or if there are multiple crossing points, then bandwidths of conventional size are generally appropriate. The range of different modes of behavior is narrower in multivariate settings. There, the optimal size of bandwidth is generally the same as that which is appropriate for pointwise density estimation. These properties motivate empirical rules for bandwidth choice.

math.ST↗

Attributing a probability to the shape of a probability density

We discuss properties of two methods for ascribing probabilities to the shape of a probability distribution. One is based on the idea of counting the number of modes of a bootstrap version of a standard kernel density estimator. We argue that the simplest form of that method suffers from the same difficulties that inhibit level accuracy of Silverman's bandwidth-based test for modality: the conditional distribution of the bootstrap form of a density estimator is not a good approximation to the actual distribution of the estimator. This difficulty is less pronounced if the density estimator is oversmoothed, but the problem of selecting the extent of oversmoothing is inherently difficult. It is shown that the optimal bandwidth, in the sense of producing optimally high sensitivity, depends on the widths of putative bumps in the unknown density and is exactly as difficult to determine as those bumps are to detect. We also develop a second approach to ascribing a probability to shape, using Muller and Sawitzki's notion of excess mass. In contrast to the context just discussed, it is shown that the bootstrap distribution of empirical excess mass is a relatively good approximation to its true distribution. This leads to empirical approximations to the likelihoods of different levels of ``modal sharpness,'' or ``delineation,'' of modes of a density. The technique is illustrated numerically.

math.ST↗

Bump hunting with non-Gaussian kernels

It is well known that the number of modes of a kernel density estimator is monotone nonincreasing in the bandwidth if the kernel is a Gaussian density. There is numerical evidence of nonmonotonicity in the case of some non-Gaussian kernels, but little additional information is available. The present paper provides theoretical and numerical descriptions of the extent to which the number of modes is a nonmonotone function of bandwidth in the case of general compactly supported densities. Our results address popular kernels used in practice, for example, the Epanechnikov, biweight and triweight kernels, and show that in such cases nonmonotonicity is present with strictly positive probability for all sample sizes n\geq3. In the Epanechnikov and biweight cases the probability of nonmonotonicity equals 1 for all n\geq2. Nevertheless, in spite of the prevalence of lack of monotonicity revealed by these results, it is shown that the notion of a critical bandwidth (the smallest bandwidth above which the number of modes is guaranteed to be monotone) is still well defined. Moreover, just as in the Gaussian case, the critical bandwidth is of the same size as the bandwidth that minimises mean squared error of the density estimator. These theoretical results, and new numerical evidence, show that the main effects of nonmonotonicity occur for relatively small bandwidths, and have negligible impact on many aspects of bump hunting.

math.ST↗

Wavelet-based estimation with multiple sampling rates

We suggest an adaptive sampling rule for obtaining information from noisy signals using wavelet methods. The technique involves increasing the sampling rate when relatively high-frequency terms are incorporated into the wavelet estimator, and decreasing it when, again using thresholded terms as an empirical guide, signal complexity is judged to have decreased. Through sampling in this way the algorithm is able to accurately recover relatively complex signals without increasing the long-run average expense of sampling. It achieves this level of performance by exploiting the opportunities for near-real time sampling that are available if one uses a relatively high primary resolution level when constructing the basic wavelet estimator. In the practical problems that motivate the work, where signal to noise ratio is particularly high and the long-run average sampling rate may be several hundred thousand operations per second, high primary resolution levels are quite feasible.

math.ST↗

Exact convergence rate and leading term in central limit theorem for student's t statistic

The leading term in the normal approximation to the distribution of Student's t statistic is derived in a general setting, with the sole assumption being that the sampled distribution is in the domain of attraction of a normal law. The form of the leading term is shown to have its origin in the way in which extreme data influence properties of the Studentized sum. The leading-term approximation is used to give the exact rate of convergence in the central limit theorem up to order n^{-1/2}, where n denotes sample size. It is proved that the exact rate uniformly on the whole real line is identical to the exact rate on sets of just three points. Moreover, the exact rate is identical to that for the non-Studentized sum when the latter is normalized for scale using a truncated form of variance, but when the corresponding truncated centering constant is omitted. Examples of characterizations of convergence rates are also given. It is shown that, in some instances, their validity uniformly on the whole real line is equivalent to their validity on just two symmetric points.

math.PR↗