SearcharxivSearch

arXiv subjects

Anil K. Ghosh

Publications and source records attributed to Anil K. Ghosh.

18 recordsLinked to original sources

On High-Dimensional Change-Point Detection Based on Pairwise Distances

In change-point analysis, one aims at finding the locations of abrupt distributional changes (if any) in a sequence of multivariate observations. In this article, we propose some nonparametric methods based on averages of pairwise distances for this purpose. These distance-based methods can be conveniently used for high-dimensional data even when the dimension is much larger than the sample size (i.e., the length of the sequence). We carry out some theoretical investigations on the behaviour of these methods not only when the dimension of the data remains fixed and the sample size grows to infinity, but also in situations where the dimension diverges to infinity while the sample size may or may not grow with the dimension. Several high-dimensional datasets are analyzed to compare the empirical performance of these proposed methods against some state-of-the-art methods.

math.ST

A nonparametric test of spherical symmetry applicable to high dimensional data

We develop a test for spherical symmetry of a multivariate distribution $\Pr$ that works well even when the dimension of the data $d$ is larger than the sample size $n$. We propose a non-negative measure of spherical asymmetry $ζ(\Pr)$ such that $ζ(\Pr)=0$ if and only if $\Pr$ is spherically symmetric. We construct a consistent estimator of $ζ(\Pr)$ using the data augmentation method and investigate its large sample properties. The proposed test based on this estimator is calibrated using a novel resampling algorithm. Our test controls the type I error, and it is consistent against general alternatives. We also study its behavior for a sequence of alternatives $(1-δ_n) F+δ_n G$, where $ζ(G)=0$ but $ζ(F)>0$, and $δ_n \in [0,1]$. When $\lim\supδ_n<1$, for any $G$, the power of our test converges to unity as $n$ increases. However, if $\lim\supδ_n=1$, the asymptotic power of our test depends on $\lim n(1-δ_n)^2$. We establish this by proving the minimax rate optimality of our test over a suitable class of alternatives and showing that it is Pitman efficient when $\lim n(1-δ_n)^2>0$. Moreover, our test is provably consistent for high-dimensional data even when $d$ grows with $n$. When the center of symmetry is not specified by the null hypothesis, most of the existing tests often fail to satisfy the level property. To take care of this problem, we propose a general recipe for constructing modified tests based on pairwise differences of the observations. Our numerical results amply demonstrate the superiority of the proposed test over some state-of-the-art methods.

math.ST

User-UAV Association for Dynamic User in mmWave Communication for eMBB and URLLC

In unmanned aerial vehicle (UAV) assisted millimeter wave (mmWave) communication, appropriate user-UAV association is crucial for improving system performance. In mmWave communication, user throughput largely depends on the line of sight (LoS) connectivity with the UAV, which in turn depends on the mobility pattern of the users. Moreover, different traffic types like enhanced mobile broadband (eMBB) and ultra reliable low latency communication (URLLC) may require different types of LoS connectivity. Existing user-UAV association policies do not consider the user mobility during a time interval and different LoS requirements of different traffic types. In this paper, we consider both of them and develop a user association policy in the presence of building blockages. First, considering a simplified scenario, we have analytically established the LoS area, which is the region where users will experience seamless LoS connectivity for eMBB traffic, and the LoS radius, which is the radius of the largest circle within which the user gets uninterrupted LoS services for URLLC traffic. Then, for a more complex scenario, we present a geometric shadow polygon-based method to compute LoS area and LoS radius. Finally, we associate eMBB and URLLC users, with the UAVs from which they get the maximum average throughput based on LoS area and maximum LoS radius respectively. We show that our approach outperforms the existing discretization based and maximum throughput based approaches.

cs.NI

Exact distribution-free tests of spherical symmetry applicable to high dimensional data

We develop some graph-based tests for spherical symmetry of a multivariate distribution using a method based on data augmentation. These tests are constructed using a new notion of signs and ranks that are computed along a path obtained by optimizing an objective function based on pairwise dissimilarities among the observations in the augmented data set. The resulting tests based on these signs and ranks have the exact distribution-free property, and irrespective of the dimension of the data, the null distributions of the test statistics remain the same. These tests can be conveniently used for high-dimensional data, even when the dimension is much larger than the sample size. Under appropriate regularity conditions, we prove the consistency of these tests in high dimensional asymptotic regime, where the dimension grows to infinity while the sample size may or may not grow with the dimension. We also propose a generalization of our methods to take care of the situations, where the center of symmetry is not specified by the null hypothesis. Several simulated data sets and a real data set are analyzed to demonstrate the utility of the proposed tests.

math.ST

On high-dimensional modifications of the nearest neighbor classifier

Nearest neighbor classifier is arguably the most simple and popular nonparametric classifier available in the literature. However, due to the concentration of pairwise distances and the violation of the neighborhood structure, this classifier often suffers in high-dimension, low-sample size (HDLSS) situations, especially when the scale difference between the competing classes dominates their location difference. Several attempts have been made in the literature to take care of this problem. In this article, we discuss some of these existing methods and propose some new ones. We carry out some theoretical investigations in this regard and analyze several simulated and benchmark datasets to compare the empirical performances of proposed methods with some of the existing ones.

stat.ML

Classification Using Global and Local Mahalanobis Distances

We propose a novel semiparametric classifier based on Mahalanobis distances of an observation from the competing classes. Our tool is a generalized additive model with the logistic link function that uses these distances as features to estimate the posterior probabilities of different classes. While popular parametric classifiers like linear and quadratic discriminant analyses are mainly motivated by the normality of the underlying distributions, the proposed classifier is more flexible and free from such parametric modeling assumptions. Since the densities of elliptic distributions are functions of Mahalanobis distances, this classifier works well when the competing classes are (nearly) elliptic. In such cases, it often outperforms popular nonparametric classifiers, especially when the sample size is small compared to the dimension of the data. To cope with non-elliptic and possibly multimodal distributions, we propose a local version of the Mahalanobis distance. Subsequently, we propose another classifier based on a generalized additive model that uses the local Mahalanobis distances as features. This nonparametric classifier usually performs like the Mahalanobis distance based semiparametric classifier when the underlying distributions are elliptic, but outperforms it for several non-elliptic and multimodal distributions. We also investigate the behaviour of these two classifiers in high dimension, low sample size situations. A thorough numerical study involving several simulated and real datasets demonstrate the usefulness of the proposed classifiers in comparison to many state-of-the-art methods.

stat.ME

A Ball Divergence Based Measure For Conditional Independence Testing

In this paper we introduce a new measure of conditional dependence between two random vectors ${\boldsymbol X}$ and ${\boldsymbol Y}$ given another random vector $\boldsymbol Z$ using the ball divergence. Our measure characterizes conditional independence and does not require any moment assumptions. We propose a consistent estimator of the measure using a kernel averaging technique and derive its asymptotic distribution. Using this statistic we construct two tests for conditional independence, one in the model-${\boldsymbol X}$ framework and the other based on a novel local wild bootstrap algorithm. In the model-${\boldsymbol X}$ framework, which assumes the knowledge of the distribution of ${\boldsymbol X}|{\boldsymbol Z}$, applying the conditional randomization test we obtain a method that controls Type I error in finite samples and is asymptotically consistent, even if the distribution of ${\boldsymbol X}|{\boldsymbol Z}$ is incorrectly specified up to distance preserving transformations. More generally, in situations where ${\boldsymbol X}|{\boldsymbol Z}$ is unknown or hard to estimate, we design a double-bandwidth based local wild bootstrap algorithm that asymptotically controls both Type I error and power. We illustrate the advantage of our method, both in terms of Type I error and power, in a range of simulation settings and also in a real data example. A consequence of our theoretical results is a general framework for studying the asymptotic properties of a 2-sample conditional $V$-statistic, which is of independent interest.

math.ST

Nearest Neighbor Classification based on Imbalanced Data: A Statistical Approach

When the competing classes in a classification problem are not of comparable size, many popular classifiers exhibit a bias towards larger classes, and the nearest neighbor classifier is no exception. To take care of this problem, we develop a statistical method for nearest neighbor classification based on such imbalanced data sets. First, we construct a classifier for the binary classification problem and then extend it for classification problems involving more than two classes. Unlike the existing oversampling or undersampling methods, our proposed classifiers do not need to generate any pseudo observations or remove any existing observations, hence the results are exactly reproducible. We establish the Bayes risk consistency of these classifiers under appropriate regularity conditions. Their superior performance over the existing methods is amply demonstrated by analyzing several benchmark data sets.

stat.ME

On Exact Feature Screening in Ultrahigh-dimensional Binary Classification

We propose a new model-free feature screening method based on energy distances for ultrahigh-dimensional binary classification problems. With a high probability, the proposed method retains only relevant features after discarding all the noise variables. The proposed screening method is also extended to identify pairs of variables that are marginally undetectable but have differences in their joint distributions. Finally, we build a classifier that maintains coherence between the proposed feature selection criteria and discrimination method and also establish its risk consistency. An extensive numerical study with simulated and real benchmark data sets shows clear and convincing advantages of our proposed method over the state-of-the-art methods.

stat.ME

On High Dimensional Behaviour of Some Two-Sample Tests Based on Ball Divergence

In this article, we propose some two-sample tests based on ball divergence and investigate their high dimensional behavior. First, we study their behavior for High Dimension, Low Sample Size (HDLSS) data, and under appropriate regularity conditions, we establish their consistency in the HDLSS regime, where the dimension of the data grows to infinity while the sample sizes from the two distributions remain fixed. Further, we show that these conditions can be relaxed when the sample sizes also increase with the dimension, and in such cases, consistency can be proved even for shrinking alternatives. We use a simple example involving two normal distributions to prove that even when there are no consistent tests in the HDLSS regime, the powers of the proposed tests can converge to unity if the sample sizes increase with the dimension at an appropriate rate. This rate is obtained by establishing the minimax rate optimality of our tests over a certain class of alternatives. Several simulated and benchmark data sets are analyzed to compare the performance of these proposed tests with the state-of-the-art methods that can be used for testing the equality of two high-dimensional probability distributions.

math.ST

On Generalizations of Some Distance Based Classifiers for HDLSS Data

In high dimension, low sample size (HDLSS) settings, classifiers based on Euclidean distances like the nearest neighbor classifier and the average distance classifier perform quite poorly if differences between locations of the underlying populations get masked by scale differences. To rectify this problem, several modifications of these classifiers have been proposed in the literature. However, existing methods are confined to location and scale differences only, and often fail to discriminate among populations differing outside of the first two moments. In this article, we propose some simple transformations of these classifiers resulting into improved performance even when the underlying populations have the same location and scale. We further propose a generalization of these classifiers based on the idea of grouping of variables. The high-dimensional behavior of the proposed classifiers is studied theoretically. Numerical experiments with a variety of simulated examples as well as an extensive analysis of real data sets exhibit advantages of the proposed methods.

stat.ME

Some Clustering-based Change-point Detection Methods Applicable to High Dimension, Low Sample Size Data

Detection of change-points in a sequence of high-dimensional observations is a very challenging problem, and this becomes even more challenging when the sample size (i.e., the sequence length) is small. In this article, we propose some change-point detection methods based on clustering, which can be conveniently used in such high dimension, low sample size situations. First, we consider the single change-point problem. Using k-means clustering based on some suitable dissimilarity measures, we propose some methods for testing the existence of a change-point and estimating its location. High-dimensional behavior of these proposed methods are investigated under appropriate regularity conditions. Next, we extend our methods for detection of multiple change-points. We carry out extensive numerical studies to compare the performance of our proposed methods with some state-of-the-art methods.

stat.ME

On high-dimensional modifications of some graph-based two-sample tests

Testing for the equality of two high-dimensional distributions is a challenging problem, and this becomes even more challenging when the sample size is small. Over the last few decades, several graph-based two-sample tests have been proposed in the literature, which can be used for data of arbitrary dimensions. Most of these test statistics are computed using pairwise Euclidean distances among the observations. But, due to concentration of pairwise Euclidean distances, these tests have poor performance in many high-dimensional problems. Some of them can have powers even below the nominal level when the scale-difference between two distributions dominates the location-difference. To overcome these limitations, we introduce a new class of dissimilarity indices and use it to modify some popular graph-based tests. These modified tests use the distance concentration phenomenon to their advantage, and as a result, they outperform the corresponding tests based on the Euclidean distance in a wide variety of examples. We establish the high-dimensional consistency of these modified tests under fairly general conditions. Analyzing several simulated as well as real data sets, we demonstrate their usefulness in high dimension, low sample size situations.

stat.ME

On perfect clustering of high dimension, low sample size data

Popular clustering algorithms based on usual distance functions (e.g., Euclidean distance) often suffer in high dimension, low sample size (HDLSS) situations, where concentration of pairwise distances has adverse effects on their performance. In this article, we use a dissimilarity measure based on the data cloud, called MADD, which takes care of this problem. MADD uses the distance concentration phenomenon to its advantage, and as a result, clustering algorithms based on MADD usually perform better for high dimensional data. Using theoretical and numerical results, we amply demonstrate it in this article. We also address the problem of estimating the number of clusters. This is a very challenging problem in cluster analysis, and several algorithms have been proposed for it. We show that many of these existing algorithms have superior performance in high dimensions when MADD is used instead of the Euclidean distance. We also construct a new estimator based on penalized Dunn index and prove its consistency in the HDLSS asymptotic regime, where the sample size remains fixed and the dimension grows to infinity. Several simulated and real data sets are analyzed to demonstrate the importance of MADD for cluster analysis of high dimensional data.

stat.ME

On Affine Invariant $L_p$ Depth Classifiers based on an Adaptive Choice of $p$

In this article, we use L$_p$ depth for classification of multivariate data, where the value of $p$ is chosen adaptively using observations from the training sample. While many depth based classifiers are constructed assuming elliptic symmetry of the underlying distributions, our proposed L$_p$ depth classifiers cater to a larger class of distributions. We establish Bayes risk consistency of these proposed classifiers under appropriate regularity conditions. Several simulated and benchmark data sets are analyzed to compare their finite sample performance with some existing parametric and nonparametric classifiers including those based on other notions of data depth.

stat.ME

Multi-scale Classification using Localized Spatial Depth

In this article, we develop and investigate a new classifier based on features extracted using spatial depth. Our construction is based on fitting a generalized additive model to the posterior probabilities of the different competing classes. To cope with possible multi-modal as well as non-elliptic population distributions, we develop a localized version of spatial depth and use that with varying degrees of localization to build the classifier. Final classification is done by aggregating several posterior probability estimates each of which is obtained using localized spatial depth with a fixed scale of localization. The proposed classifier can be conveniently used even when the dimension is larger than the sample size, and its good discriminatory power for such data has been established using theoretical as well as numerical results.

stat.ME

Some intriguing properties of Tukey's half-space depth

For multivariate data, Tukey's half-space depth is one of the most popular depth functions available in the literature. It is conceptually simple and satisfies several desirable properties of depth functions. The Tukey median, the multivariate median associated with the half-space depth, is also a well-known measure of center for multivariate data with several interesting properties. In this article, we derive and investigate some interesting properties of half-space depth and its associated multivariate median. These properties, some of which are counterintuitive, have important statistical consequences in multivariate analysis. We also investigate a natural extension of Tukey's half-space depth and the related median for probability distributions on any Banach space (which may be finite- or infinite-dimensional) and prove some results that demonstrate anomalous behavior of half-space depth in infinite-dimensional spaces.

math.ST

Secure Position Verification for Wireless Sensor Networks in Noisy Channels

Position verification in wireless sensor networks (WSNs) is quite tricky in presence of attackers (malicious sensor nodes), who try to break the verification protocol by reporting their incorrect positions (locations) during the verification stage. In the literature of WSNs, most of the existing methods of position verification have used trusted verifiers, which are in fact vulnerable to attacks by malicious nodes. They also depend on some distance estimation techniques, which are not accurate in noisy channels (mediums). In this article, we propose a secure position verification scheme for WSNs in noisy channels without relying on any trusted entities. Our verification scheme detects and filters out all malicious nodes from the network with very high probability.

cs.DC