SearcharxivSearch

arXiv subjects

Subhajit Dutta

Publications and source records attributed to Subhajit Dutta.

16 recordsLinked to original sources

Fast high-dimensional mean testing via logistic regression

We propose computationally efficient tests for equality of mean vectors of two or more high-dimensional populations. Central to our approach is an equivalence between equality of means and a zero population logistic regression parameter. We establish this equivalence for independently distributed observations without imposing common distributional assumptions across populations. Our procedure uses logistic Lasso to screen informative variables and an unpenalized logistic refit for inference in the reduced dimension, yielding asymptotically correct size and consistency. For a specified two-sample Gaussian submodel and sparse discriminative class, the test also attains the minimax separation rate. The framework extends to multiple populations through multi-class logistic regression. Simulations demonstrate accurate size control, strong power, and favorable computational scaling compared with existing tests under unbalanced designs and variance heterogeneity. Applications to gene-expression data with more than twenty-two thousand variables illustrate the practical scalability of the proposed procedures.

stat.ME

Uniform-over-dimension location tests for multivariate and high-dimensional data

Asymptotic methods for hypothesis testing in high-dimensional data usually require the dimension of the observations to increase to infinity, often with an additional relationship between the dimension (say, $p$) and the sample size (say, $n$). On the other hand, multivariate asymptotic testing methods are valid for fixed dimension only and their implementations typically require the sample size to be large compared to the dimension to yield desirable results. In practical scenarios, it is usually not possible to determine whether the dimension of the data conform to the conditions required for the validity of the high-dimensional asymptotic methods for hypothesis testing, or whether the sample size is large enough compared to the dimension of the data. In this work, we first describe the notion of uniform-over-$p$ convergences and subsequently, develop a uniform-over-dimension central limit theorem. An asymptotic test for the two-sample equality of locations is developed, which now holds uniformly over the dimension of the observations. Using simulated and real data, it is demonstrated that the proposed test exhibits better performance compared to several popular tests in the literature for high-dimensional data as well as the usual scaled two-sample tests for multivariate data, including the Hotelling's $T^2$ test for multivariate Gaussian data.

stat.ME

Ultrahigh-dimensional Quadratic Discriminant Analysis Using Random Projections

This paper investigates the effectiveness of using the Random Projection Ensemble (RPE) approach in Quadratic Discriminant Analysis (QDA) for ultrahigh-dimensional classification problems. Classical methods such as Linear Discriminant Analysis (LDA) and QDA are used widely, but face significant challenges in their implementation when the data dimension (say, $p$) exceeds the sample size (say, $n$). In particular, both LDA (using the Moore-Penrose inverse for covariance matrices) and QDA (even with known covariance matrices) may perform as poorly as random guessing when $p/n \to \infty$ as $n \to \infty$. The RPE method, known for addressing the curse of dimensionality, offers a fast and effective solution without relying on selective summary measures of the competing distributions. This paper demonstrates the practical advantages of employing RPE on QDA in terms of classification performance as well as computational efficiency. We establish results for limiting perfect classification in both the population and sample versions of the proposed RPE-QDA classifier, under fairly general assumptions that allow for sub-exponential growth of $p$ relative to $n$. Several simulated and gene expression data sets are analyzed to evaluate the performance of the proposed classifier in ultrahigh-dimensional~scenarios.

stat.ME

Uniform-over-dimension convergence with application to location tests for high-dimensional data

Asymptotic methods for hypothesis testing in high-dimensional data usually require the dimension of the observations to increase to infinity, often with an additional condition on its rate of increase compared to the sample size. On the other hand, multivariate asymptotic methods are valid for fixed dimension only, and their practical implementations in hypothesis testing methodology typically require the sample size to be large compared to the dimension for yielding desirable results. However, in practical scenarios, it is usually not possible to determine whether the dimension of the data at hand conform to the conditions required for the validity of the high-dimensional asymptotic methods, or whether the sample size is large enough compared to the dimension of the data. In this work, a theory of asymptotic convergence is proposed, which holds uniformly over the dimension of the random vectors. This theory attempts to unify the asymptotic results for fixed-dimensional multivariate data and high-dimensional data, and accounts for the effect of the dimension of the data on the performance of the hypothesis testing procedures. The methodology developed based on this asymptotic theory can be applied to data of any dimension. An application of this theory is demonstrated in the two-sample test for the equality of locations. The test statistic proposed is unscaled by the sample covariance, similar to usual tests for high-dimensional data. Using simulated examples, it is demonstrated that the proposed test exhibits better performance compared to several popular tests in the literature for high-dimensional data. Further, it is demonstrated in simulated models that the proposed unscaled test performs better than the usual scaled two-sample tests for multivariate data, including the Hotelling's $T^2$ test for multivariate Gaussian data.

math.ST

Bayesian Variable Selection Under High-dimensional Settings With Grouped Covariates

Consider the normal linear regression setup when the number of covariates p is much larger than the sample size n, and the covariates form correlated groups. The response variable y is not related to an entire group of covariates in all or none basis, rather the sparsity assumption persists within and between groups. We extend the traditional g-prior setup to this framework. Variable selection consistency of the proposed method is shown under fairly general conditions, assuming the covariates to be random and allowing the true model to grow with both n and p. For the purpose of implementation of the proposed g-prior method to high-dimensional setup, we propose two procedures. First, a group screening procedure, termed as group SIS (GSIS), and secondly, a novel stochastic search variable selection algorithm, termed as group informed variable selection algorithm (GiVSA), which uses the known group structure efficiently to explore the model space without discarding any covariate based on an initial screening. Screening consistency of GSIS, and theoretical mixing time of GiVSA are studied using the canonical path ensemble approach of Yang et al. (2016). Performance of the proposed prior with implementation of GSIS as well as GiVSA are validated using various simulated examples and a real data related to residential buildings.

stat.ME

Robust Classification of High-Dimensional Data using Data-Adaptive Energy Distance

Classification of high-dimensional low sample size (HDLSS) data poses a challenge in a variety of real-world situations, such as gene expression studies, cancer research, and medical imaging. This article presents the development and analysis of some classifiers that are specifically designed for HDLSS data. These classifiers are free of tuning parameters and are robust, in the sense that they are devoid of any moment conditions of the underlying data distributions. It is shown that they yield perfect classification in the HDLSS asymptotic regime, under some fairly general conditions. The comparative performance of the proposed classifiers is also investigated. Our theoretical results are supported by extensive simulation studies and real data analysis, which demonstrate promising advantages of the proposed classification techniques over several widely recognized methods.

stat.ML

On Exact Feature Screening in Ultrahigh-dimensional Binary Classification

We propose a new model-free feature screening method based on energy distances for ultrahigh-dimensional binary classification problems. With a high probability, the proposed method retains only relevant features after discarding all the noise variables. The proposed screening method is also extended to identify pairs of variables that are marginally undetectable but have differences in their joint distributions. Finally, we build a classifier that maintains coherence between the proposed feature selection criteria and discrimination method and also establish its risk consistency. An extensive numerical study with simulated and real benchmark data sets shows clear and convincing advantages of our proposed method over the state-of-the-art methods.

stat.ME

Sub-dimensional Mardia measures of multivariate skewness and kurtosis

The Mardia measures of multivariate skewness and kurtosis summarize the respective characteristics of a multivariate distribution with two numbers. However, these measures do not reflect the sub-dimensional features of the distribution. Consequently, testing procedures based on these measures may fail to detect skewness or kurtosis present in a sub-dimension of the multivariate distribution. We introduce sub-dimensional Mardia measures of multivariate skewness and kurtosis, and investigate the information they convey about all sub-dimensional distributions of some symmetric and skewed families of multivariate distributions. The maxima of the sub-dimensional Mardia measures of multivariate skewness and kurtosis are considered, as these reflect the maximum skewness and kurtosis present in the distribution, and also allow us to identify the sub-dimension bearing the highest skewness and kurtosis. Asymptotic distributions of the vectors of sub-dimensional Mardia measures of multivariate skewness and kurtosis are derived, based on which testing procedures for the presence of skewness and of deviation from Gaussian kurtosis are developed. The performances of these tests are compared with some existing tests in the literature on simulated and real datasets.

stat.ME

On Perfect Classification and Clustering for Gaussian Processes

In this paper, we propose a data based transformation for infinite-dimensional Gaussian processes and derive its limit theorem. For a classification problem, this transformation induces complete separation among the associated Gaussian processes. The misclassification probability of any simple classifier when applied on the transformed data asymptotically converges to zero. In a clustering problem using mixture models, an appropriate modification of this transformation asymptotically leads to perfect separation of the populations. Theoretical properties are studied for the usual $k$-means clustering method when used on this transformed data. Good empirical performance of the proposed methodology is demonstrated using simulated as well as benchmark data sets, when compared with some popular parametric and nonparametric methods for such functional data.

math.ST

On Generalizations of Some Distance Based Classifiers for HDLSS Data

In high dimension, low sample size (HDLSS) settings, classifiers based on Euclidean distances like the nearest neighbor classifier and the average distance classifier perform quite poorly if differences between locations of the underlying populations get masked by scale differences. To rectify this problem, several modifications of these classifiers have been proposed in the literature. However, existing methods are confined to location and scale differences only, and often fail to discriminate among populations differing outside of the first two moments. In this article, we propose some simple transformations of these classifiers resulting into improved performance even when the underlying populations have the same location and scale. We further propose a generalization of these classifiers based on the idea of grouping of variables. The high-dimensional behavior of the proposed classifiers is studied theoretically. Numerical experiments with a variety of simulated examples as well as an extensive analysis of real data sets exhibit advantages of the proposed methods.

stat.ME

Some multivariate goodness of fit tests based on data depth

Using the fact that some depth functions characterize certain family of distribution functions, and under some mild conditions, distribution of the depth is continuous, we have constructed several new multivariate goodness of fit tests based on existing univariate GoF tests. Since exact computation of depth is difficult, depth is computed with respect to a large random sample drawn from the null distribution. It has been shown that test statistic based on estimated depth is close to that based on true depth for a large random sample from the null distribution. Some two sample tests for scale difference, based on data depth are also discussed. These tests are distribution-free under the null hypothesis. Finite sample properties of the tests are studied through several numerical examples. A real data example is discussed to illustrate usefulness of the proposed tests.

math.ST

On Construction of Higher Order Kernels Using Fourier Transforms and Covariance Functions

In this paper, we show that a suitably chosen covariance function of a continuous time, second order stationary stochastic process can be viewed as a symmetric higher order kernel. This leads to the construction of a higher order kernel by choosing an appropriate covariance function. An optimal choice of the constructed higher order kernel that partially minimizes the mean integrated square error of the kernel density estimator is also discussed.

math.ST

On a Generalization of the Average Distance Classifier

In high dimension, low sample size (HDLSS)settings, the simple average distance classifier based on the Euclidean distance performs poorly if differences between the locations get masked by the scale differences. To rectify this issue, modifications to the average distance classifier was proposed by Chan and Hall (2009). However, the existing classifiers cannot discriminate when the populations differ in other aspects than locations and scales. In this article, we propose some simple transformations of the average distance classifier to tackle this issue. The resulting classifiers perform quite well even when the underlying populations have the same location and scale. The high-dimensional behaviour of the proposed classifiers is studied theoretically. Numerical experiments with a variety of simulated as well as real data sets exhibit the usefulness of the proposed methodology.

stat.ME

On Affine Invariant $L_p$ Depth Classifiers based on an Adaptive Choice of $p$

In this article, we use L$_p$ depth for classification of multivariate data, where the value of $p$ is chosen adaptively using observations from the training sample. While many depth based classifiers are constructed assuming elliptic symmetry of the underlying distributions, our proposed L$_p$ depth classifiers cater to a larger class of distributions. We establish Bayes risk consistency of these proposed classifiers under appropriate regularity conditions. Several simulated and benchmark data sets are analyzed to compare their finite sample performance with some existing parametric and nonparametric classifiers including those based on other notions of data depth.

stat.ME

Multi-scale Classification using Localized Spatial Depth

In this article, we develop and investigate a new classifier based on features extracted using spatial depth. Our construction is based on fitting a generalized additive model to the posterior probabilities of the different competing classes. To cope with possible multi-modal as well as non-elliptic population distributions, we develop a localized version of spatial depth and use that with varying degrees of localization to build the classifier. Final classification is done by aggregating several posterior probability estimates each of which is obtained using localized spatial depth with a fixed scale of localization. The proposed classifier can be conveniently used even when the dimension is larger than the sample size, and its good discriminatory power for such data has been established using theoretical as well as numerical results.

stat.ME

Some intriguing properties of Tukey's half-space depth

For multivariate data, Tukey's half-space depth is one of the most popular depth functions available in the literature. It is conceptually simple and satisfies several desirable properties of depth functions. The Tukey median, the multivariate median associated with the half-space depth, is also a well-known measure of center for multivariate data with several interesting properties. In this article, we derive and investigate some interesting properties of half-space depth and its associated multivariate median. These properties, some of which are counterintuitive, have important statistical consequences in multivariate analysis. We also investigate a natural extension of Tukey's half-space depth and the related median for probability distributions on any Banach space (which may be finite- or infinite-dimensional) and prove some results that demonstrate anomalous behavior of half-space depth in infinite-dimensional spaces.

math.ST