SearcharxivSearch

arXiv subjects

Probal Chaudhuri

Publications and source records attributed to Probal Chaudhuri.

15 recordsLinked to original sources

Quantile processes and their applications in finite populations

The weak convergence of the quantile processes, which are constructed based on different estimators of the finite population quantiles, is shown under various well-known sampling designs based on a superpopulation model. The results related to the weak convergence of these quantile processes are applied to find asymptotic distributions of the smooth $L$-estimators and the estimators of smooth functions of finite population quantiles. Based on these asymptotic distributions, confidence intervals are constructed for several finite population parameters like the median, the $\alpha$-trimmed means, the interquartile range and the quantile based measure of skewness. Comparisons of various estimators are carried out based on their asymptotic distributions. We show that the use of the auxiliary information in the construction of the estimators sometimes has an adverse effect on the performances of the smooth $L$-estimators and the estimators of smooth functions of finite population quantiles under several sampling designs. Further, the performance of each of the above-mentioned estimators sometimes becomes worse under sampling designs, which use the auxiliary information, than their performances under simple random sampling without replacement (SRSWOR).

math.ST

On estimators of the mean of infinite dimensional data in finite populations

The Horvitz-Thompson (HT), the Rao-Hartley-Cochran (RHC) and the generalized regression (GREG) estimators of the finite population mean are considered, when the observations are from an infinite dimensional space. We compare these estimators based on their asymptotic distributions under some commonly used sampling designs and some superpopulations satisfying linear regression models. We show that the GREG estimator is asymptotically at least as efficient as any of the other two estimators under different sampling designs considered in this paper. Further, we show that the use of some well known sampling designs utilizing auxiliary information may have an adverse effect on the performance of the GREG estimator, when the degree of heteroscedasticity present in linear regression models is not very large. On the other hand, the use of those sampling designs improves the performance of this estimator, when the degree of heteroscedasticity present in linear regression models is large. We develop methods for determining the degree of heteroscedasticity, which in turn determines the choice of appropriate sampling design to be used with the GREG estimator. We also investigate the consistency of the covariance operators of the above estimators. We carry out some numerical studies using real and synthetic data, and our theoretical results are supported by the results obtained from those numerical studies.

math.ST

A comparison of estimators of mean and its functions in finite populations

Several well known estimators of finite population mean and its functions are investigated under some standard sampling designs. Such functions of mean include the variance, the correlation coefficient and the regression coefficient in the population as special cases. We compare the performance of these estimators under different sampling designs based on their asymptotic distributions. Equivalence classes of estimators under different sampling designs are constructed so that estimators in the same class have equivalent performance in terms of asymptotic mean squared errors (MSEs). Estimators in different equivalence classes are then compared under some superpopulations satisfying linear models. It is shown that the pseudo empirical likelihood (PEML) estimator of the population mean under simple random sampling without replacement (SRSWOR) has the lowest asymptotic MSE among all the estimators under different sampling designs considered in this paper. It is also shown that for the variance, the correlation coefficient and the regression coefficient of the population, the plug-in estimators based on the PEML estimator have the lowest asymptotic MSEs among all the estimators considered in this paper under SRSWOR. On the other hand, for any high entropy $\pi$PS (HE$\pi$PS) sampling design, which uses the auxiliary information, the plug-in estimators of those parameters based on the H\'ajek estimator have the lowest asymptotic MSEs among all the estimators considered in this paper.

math.ST

Multi-sample Comparison Using Spatial Signs for Infinite Dimensional Data

We consider an analysis of variance type problem, where the sample observations are random elements in an infinite dimensional space. This scenario covers the case, where the observations are random functions. For such a problem, we propose a test based on spatial signs. We develop an asymptotic implementation as well as a bootstrap implementation and a permutation implementation of this test and investigate their size and power properties. We compare the performance of our test with that of several mean based tests of analysis of variance for functional data studied in the literature. Interestingly, our test not only outperforms the mean based tests in several non-Gaussian models with heavy tails or skewed distributions, but in some Gaussian models also. Further, we also compare the performance of our test with the mean based tests in several models involving contaminated probability distributions. Finally, we demonstrate the performance of these tests in three real datasets: a Canadian weather dataset, a spectrometric dataset on chemical analysis of meat samples and a dataset on orthotic measurements on volunteers.

stat.ME

Depth based inference on conditional distribution with infinite dimensional data

We develop inference and testing procedures for conditional dispersion and skewness in a nonparametric regression setup based on statistical depth functions. The methods developed can be applied in situations, where the response is multivariate and the covariate is a random element in a metric space. This includes regression with functional covariate as a special case. We construct measures of the center, the spread and the skewness of the conditional distribution of the response given the covariate using depth based nonparametric regression procedures. We establish the asymptotic consistency of those measures and develop a test for heteroscedasticity and a test for conditional skewness. We present level and power study for the tests in several simulated models. The usefulness of the methodology is also demonstrated in a real dataset. In that dataset, our responses are the nutritional contents of different meat samples measured by their protein, fat and moisture contents, and the functional covariate is the absorbance spectra of the meat samples.

stat.ME

Convergence Rates for Kernel Regression in Infinite Dimensional Spaces

We consider a nonparametric regression setup, where the covariate is a random element in a complete separable metric space, and the parameter of interest associated with the conditional distribution of the response lies in a separable Banach space. We derive the optimum convergence rate for the kernel estimate of the parameter in this setup. The small ball probability in the covariate space plays a critical role in determining the asymptotic variance of kernel estimates. Unlike the case of finite dimensional covariates, we show that the asymptotic orders of the bias and the variance of the estimate achieving the optimum convergence rate may be different for infinite dimensional covariates. Also, the bandwidth, which balances the bias and the variance, may lead to an estimate with suboptimal mean square error for infinite dimensional covariates. We describe a data-driven adaptive choice of the bandwidth, and derive the asymptotic behavior of the adaptive estimate.

math.ST

Nonparametric Depth and Quantile Regression for Functional Data

We investigate nonparametric regression methods based on spatial depth and quantiles when the response and the covariate are both functions. As in classical quantile regression for finite dimensional data, regression techniques developed here provide insight into the influence of the functional covariate on different parts, like the center as well as the tails, of the conditional distribution of the functional response. Depth and quantile based nonparametric regressions are useful to detect heteroscedasticity in functional regression. We derive the asymptotic behaviour of nonparametric depth and quantile regression estimates, which depend on the small ball probabilities in the covariate space. Our nonparametric regression procedures are used to analyse a dataset about the influence of per capita GDP on saving rates for 125 countries, and another dataset on the effects of per capita net disposable income on the sale of cigarettes in some states in the US.

stat.ME

Tests for high dimensional data based on means, spatial signs and spatial ranks

Tests based on sample mean vectors and sample spatial signs have been studied in the recent literature for high dimensional data with the dimension larger than the sample size. For suitable sequences of alternatives, we show that the powers of the mean based tests and the tests based on spatial signs and ranks tend to be same as the data dimension grows to infinity for any sample size, when the coordinate variables satisfy appropriate mixing conditions. Further, their limiting powers do not depend on the heaviness of the tails of the distributions. This is in striking contrast to the asymptotic results obtained in the classical multivariate setup. On the other hand, we show that in the presence of stronger dependence among the coordinate variables, the spatial sign and rank based tests for high dimensional data can be asymptotically more powerful than the mean based tests if in addition to the data dimension, the sample size also grows to infinity. The sizes of some mean based tests for high dimensional data studied in the recent literature are observed to be significantly different from their nominal levels. This is due to the inadequacy of the asymptotic approximations used for the distributions of those test statistics. However, our asymptotic approximations for the tests based on spatial signs and ranks are observed to work well when the tests are applied on a variety of simulated and real datasets.

math.ST

Paired sample tests in infinite dimensional spaces

The sign and the signed-rank tests for univariate data are perhaps the most popular nonparametric competitors of the t test for paired sample problems. These tests have been extended in various ways for multivariate data in finite dimensional spaces. These extensions include tests based on spatial signs and signed ranks, which have been studied extensively by Hannu Oja and his coauthors. They showed that these tests are asymptotically more powerful than Hotelling's $T^{2}$ test under several heavy tailed distributions. In this paper, we consider paired sample tests for data in infinite dimensional spaces based on notions of spatial sign and spatial signed rank in such spaces. We derive their asymptotic distributions under the null hypothesis and under sequences of shrinking location shift alternatives. We compare these tests with some mean based tests for infinite dimensional paired sample data. We show that for shrinking location shift alternatives, the proposed tests are asymptotically more powerful than the mean based tests for some heavy tailed distributions and even for some Gaussian distributions in infinite dimensional spaces. We also investigate the performance of different tests using some simulated data.

stat.ME

Comparison of multivariate distributions using quantile-quantile plots and related tests

The univariate quantile-quantile (Q-Q) plot is a well-known graphical tool for examining whether two data sets are generated from the same distribution or not. It is also used to determine how well a specified probability distribution fits a given sample. In this article, we develop and study a multivariate version of the Q-Q plot based on the spatial quantile. The usefulness of the proposed graphical device is illustrated on different real and simulated data, some of which have fairly large dimensions. We also develop certain statistical tests that are related to the proposed multivariate Q-Q plot and study their asymptotic properties. The performance of those tests are compared with that of some other well-known tests for multivariate distributions available in the literature.

math.ST

The spatial distribution in infinite dimensional spaces and related quantiles and depths

The spatial distribution has been widely used to develop various nonparametric procedures for finite dimensional multivariate data. In this paper, we investigate the concept of spatial distribution for data in infinite dimensional Banach spaces. Many technical difficulties are encountered in such spaces that are primarily due to the noncompactness of the closed unit ball. In this work, we prove some Glivenko-Cantelli and Donsker-type results for the empirical spatial distribution process in infinite dimensional spaces. The spatial quantiles in such spaces can be obtained by inverting the spatial distribution function. A Bahadur-type asymptotic linear representation and the associated weak convergence results for the sample spatial quantiles in infinite dimensional spaces are derived. A study of the asymptotic efficiency of the sample spatial median relative to the sample mean is carried out for some standard probability distributions in function spaces. The spatial distribution can be used to define the spatial depth in infinite dimensional Banach spaces, and we study the asymptotic properties of the empirical spatial depth in such spaces. We also demonstrate the spatial quantiles and the spatial depth using some real and simulated functional data.

stat.ME

A Wilcoxon-Mann-Whitney type test for infinite dimensional data

The Wilcoxon-Mann-Whitney test is a robust competitor of the t-test in the univariate setting. For finite dimensional multivariate data, several extensions of the Wilcoxon-Mann-Whitney test have been shown to have better performance than Hotelling's $T^{2}$ test for many non-Gaussian distributions of the data. In this paper, we study a Wilcoxon-Mann-Whitney type test based on spatial ranks for data in infinite dimensional spaces. We demonstrate the performance of this test using some real and simulated datasets. We also investigate the asymptotic properties of the proposed test and compare the test with a wide range of competing tests.

stat.ME

On data depth in infinite dimensional spaces

The concept of data depth leads to a center-outward ordering of multivariate data, and it has been effectively used for developing various data analytic tools. While different notions of depth were originally developed for finite dimensional data, there have been some recent attempts to develop depth functions for data in infinite dimensional spaces. In this paper, we consider some notions of depth in infinite dimensional spaces and study their properties under various stochastic models. Our analysis shows that some of the depth functions available in the literature have degenerate behaviour for some commonly used probability distributions in infinite dimensional spaces of sequences and functions. As a consequence, they are not very useful for the analysis of data satisfying such infinite dimensional probability models. However, some modified versions of those depth functions as well as an infinite dimensional extension of the spatial depth do not suffer from such degeneracy, and can be conveniently used for analyzing infinite dimensional data.

stat.ME

The deepest point for distributions in infinite dimensional spaces

Identification of the center of a data cloud is one of the basic problems in statistics. One popular choice for such a center is the median, and several versions of median in finite dimensional spaces have been studied in the literature. In particular, medians based on different notions of data depth have been extensively studied by many researchers, who defined median as the point, where the depth function attains its maximum value. In other words, the median is the deepest point in the sample space according to that definition. In this paper, we investigate the deepest point for probability distributions in infinite dimensional spaces. We show that for some well-known depth functions like the band depth and the half-region depth in function spaces, there may not be any meaningful deepest point for many well-known and commonly used probability models. On the other hand, certain modified versions of those depth functions as well as the spatial depth function, which can be defined in any Hilbert space, lead to some useful notions of the deepest point with nice geometric and statistical properties. The empirical versions of those deepest points can be conveniently computed for functional data, and we demonstrate this using some simulated and real data sets.

math.ST

Some intriguing properties of Tukey's half-space depth

For multivariate data, Tukey's half-space depth is one of the most popular depth functions available in the literature. It is conceptually simple and satisfies several desirable properties of depth functions. The Tukey median, the multivariate median associated with the half-space depth, is also a well-known measure of center for multivariate data with several interesting properties. In this article, we derive and investigate some interesting properties of half-space depth and its associated multivariate median. These properties, some of which are counterintuitive, have important statistical consequences in multivariate analysis. We also investigate a natural extension of Tukey's half-space depth and the related median for probability distributions on any Banach space (which may be finite- or infinite-dimensional) and prove some results that demonstrate anomalous behavior of half-space depth in infinite-dimensional spaces.

math.ST