SearcharxivSearch

arXiv subjects

Bilol Banerjee

Publications and source records attributed to Bilol Banerjee.

9 recordsLinked to original sources

Quantitative Pulse Shape-Instability Analysis Using 2D-Runs FROG

We present a method for quantifying pulse-shape instability in a train of pulses using multi-shot Second-Harmonic-Generation Frequency-Resolved Optical Gating (SHG FROG). All versions of multi-shot FROG have previously shown the ability to distinguish stable from unstable pulse trains, as systematic differences appear between measured and retrieved traces when instability is present. This has proved possible because the recently introduced Retrieved-Amplitude N-grid Algorithmic (RANA) approach provides highly reliable pulse retrieval, even for unstable pulse trains and in the presence of noise, thus eliminating the possibility that algorithm stagnation, which mimics the effects of pulse-shape instability, could be confused for it. In other words, RANAs excellent performance ensures that any non-random discrepancies between measured and retrieved FROG traces reflect physical pulse-shape instability, rather than algorithmic stagnation. To begin to quantify such instability, we now introduce an instability parameter, R. It involves the use of the well-known statistical Runs test, which tests for systematic error in fits to one-dimensional (1D) data. A runs test counts the runs consecutive points in the plot of the difference between the data and fit with the same sign evaluating the goodness of the fit while minimizing the effects of random error. However, because FROG traces are functions of two variables, we must extend the usual 1D runs test to two dimensions, that is, to enumerate 2D runs hills and valleys in the difference between measured and retrieved 2D FROG traces. Many small 2D runs indicate only random noise-like differences and hence a stable pulse train, whereas few large runs reflect additional systematic error and hence pulse-shape instability.

physics.optics

On High-Dimensional Change-Point Detection Based on Pairwise Distances

In change-point analysis, one aims at finding the locations of abrupt distributional changes (if any) in a sequence of multivariate observations. In this article, we propose some nonparametric methods based on averages of pairwise distances for this purpose. These distance-based methods can be conveniently used for high-dimensional data even when the dimension is much larger than the sample size (i.e., the length of the sequence). We carry out some theoretical investigations on the behaviour of these methods not only when the dimension of the data remains fixed and the sample size grows to infinity, but also in situations where the dimension diverges to infinity while the sample size may or may not grow with the dimension. Several high-dimensional datasets are analyzed to compare the empirical performance of these proposed methods against some state-of-the-art methods.

math.ST

Conditional Independence Testing Using Exchangeable Pairs

This article considers the problem of testing conditional independence between two random vectors \(bm X\) and \(\bm Y\) given a confounding random vector \(\bm Z\). An exchangeable-pairs framework is introduced through which the conditional independence testing problem is reformulated as a two-sample testing problem. The framework is motivated by ideas from the model-X literature and is based on a fundamental exchangeability property that holds under the null hypothesis of conditional independence. An energy-distance/maximum mean discrepancy type measure is employed on the resulting exchangeable pairs to quantify departures from conditional independence. A consistent estimator of the proposed discrepancy measure is constructed and its theoretical properties are established under general assumptions. A conditional independence test is then developed using this estimator as a test statistic and is calibrated through a suitable resampling procedure. It is shown that the proposed test is consistent against fixed alternatives, possesses nontrivial asymptotic power against local contiguous alternatives, attains the minimax separation rate for detecting alternatives characterized by the proposed discrepancy measure, and remains consistent when the data dimension diverges with the sample size. The effect of estimating the conditional distribution used to generate the exchangeable pairs is also investigated, and condition under which validity and power properties are preserved is established. Extensive simulation studies demonstrate that the proposed procedure performs competitively with some state-of-the-art methods.

math.ST

Exact distribution-free tests of spherical symmetry applicable to high dimensional data

We develop some graph-based tests for spherical symmetry of a multivariate distribution using a method based on data augmentation. These tests are constructed using a new notion of signs and ranks that are computed along a path obtained by optimizing an objective function based on pairwise dissimilarities among the observations in the augmented data set. The resulting tests based on these signs and ranks have the exact distribution-free property, and irrespective of the dimension of the data, the null distributions of the test statistics remain the same. These tests can be conveniently used for high-dimensional data, even when the dimension is much larger than the sample size. Under appropriate regularity conditions, we prove the consistency of these tests in high dimensional asymptotic regime, where the dimension grows to infinity while the sample size may or may not grow with the dimension. We also propose a generalization of our methods to take care of the situations, where the center of symmetry is not specified by the null hypothesis. Several simulated data sets and a real data set are analyzed to demonstrate the utility of the proposed tests.

math.ST

A Ball Divergence Based Measure For Conditional Independence Testing

In this paper we introduce a new measure of conditional dependence between two random vectors ${\boldsymbol X}$ and ${\boldsymbol Y}$ given another random vector $\boldsymbol Z$ using the ball divergence. Our measure characterizes conditional independence and does not require any moment assumptions. We propose a consistent estimator of the measure using a kernel averaging technique and derive its asymptotic distribution. Using this statistic we construct two tests for conditional independence, one in the model-${\boldsymbol X}$ framework and the other based on a novel local wild bootstrap algorithm. In the model-${\boldsymbol X}$ framework, which assumes the knowledge of the distribution of ${\boldsymbol X}|{\boldsymbol Z}$, applying the conditional randomization test we obtain a method that controls Type I error in finite samples and is asymptotically consistent, even if the distribution of ${\boldsymbol X}|{\boldsymbol Z}$ is incorrectly specified up to distance preserving transformations. More generally, in situations where ${\boldsymbol X}|{\boldsymbol Z}$ is unknown or hard to estimate, we design a double-bandwidth based local wild bootstrap algorithm that asymptotically controls both Type I error and power. We illustrate the advantage of our method, both in terms of Type I error and power, in a range of simulation settings and also in a real data example. A consequence of our theoretical results is a general framework for studying the asymptotic properties of a 2-sample conditional $V$-statistic, which is of independent interest.

math.ST

On high-dimensional modifications of the nearest neighbor classifier

Nearest neighbor classifier is arguably the most simple and popular nonparametric classifier available in the literature. However, due to the concentration of pairwise distances and the violation of the neighborhood structure, this classifier often suffers in high-dimension, low-sample size (HDLSS) situations, especially when the scale difference between the competing classes dominates their location difference. Several attempts have been made in the literature to take care of this problem. In this article, we discuss some of these existing methods and propose some new ones. We carry out some theoretical investigations in this regard and analyze several simulated and benchmark datasets to compare the empirical performances of proposed methods with some of the existing ones.

stat.ML

A nonparametric test of spherical symmetry applicable to high dimensional data

We develop a test for spherical symmetry of a multivariate distribution $\Pr$ that works well even when the dimension of the data $d$ is larger than the sample size $n$. We propose a non-negative measure of spherical asymmetry $\zeta(\Pr)$ such that $\zeta(\Pr)=0$ if and only if $\Pr$ is spherically symmetric. We construct a consistent estimator of $\zeta(\Pr)$ using the data augmentation method and investigate its large sample properties. The proposed test based on this estimator is calibrated using a novel resampling algorithm. Our test controls the type I error, and it is consistent against general alternatives. We also study its behavior for a sequence of alternatives $(1-\delta_n) F+\delta_n G$, where $\zeta(G)=0$ but $\zeta(F)>0$, and $\delta_n \in [0,1]$. When $\lim\sup\delta_n<1$, for any $G$, the power of our test converges to unity as $n$ increases. However, if $\lim\sup\delta_n=1$, the asymptotic power of our test depends on $\lim n(1-\delta_n)^2$. We establish this by proving the minimax rate optimality of our test over a suitable class of alternatives and showing that it is Pitman efficient when $\lim n(1-\delta_n)^2>0$. Moreover, our test is provably consistent for high-dimensional data even when $d$ grows with $n$. When the center of symmetry is not specified by the null hypothesis, most of the existing tests often fail to satisfy the level property. To take care of this problem, we propose a general recipe for constructing modified tests based on pairwise differences of the observations. Our numerical results amply demonstrate the superiority of the proposed test over some state-of-the-art methods.

math.ST

Testing distributional equality for functional random variables

In this article, we present a nonparametric method for the general two-sample problem involving functional random variables modelled as elements of a separable Hilbert space ${\cal H}$. First, we present a general recipe based on linear projections to construct a measure of dissimilarity between two probability distributions on ${\cal H}$. In particular, we consider a measure based on the energy statistic and present some of its nice theoretical properties. A plug-in estimator of this measure is used as the test statistic to construct a general two-sample test. Large sample distribution of this statistic is derived both under null and alternative hypotheses. However, since the quantiles of the limiting null distribution are analytically intractable, the test is calibrated using the permutation method. We prove the large sample consistency of the resulting permutation test under fairly general assumptions. We also study the efficiency of the proposed test by establishing a new local asymptotic normality result for functional random variables. Using that result, we derive the asymptotic distribution of the permuted test statistic and the asymptotic power of the permutation test under local contiguous alternatives. This establishes that the permutation test is statistically efficient in the Pitman sense. Extensive simulation studies are carried out and a real data set is analyzed to compare the performance of our proposed test with some state-of-the-art methods.

stat.ME

On High Dimensional Behaviour of Some Two-Sample Tests Based on Ball Divergence

In this article, we propose some two-sample tests based on ball divergence and investigate their high dimensional behavior. First, we study their behavior for High Dimension, Low Sample Size (HDLSS) data, and under appropriate regularity conditions, we establish their consistency in the HDLSS regime, where the dimension of the data grows to infinity while the sample sizes from the two distributions remain fixed. Further, we show that these conditions can be relaxed when the sample sizes also increase with the dimension, and in such cases, consistency can be proved even for shrinking alternatives. We use a simple example involving two normal distributions to prove that even when there are no consistent tests in the HDLSS regime, the powers of the proposed tests can converge to unity if the sample sizes increase with the dimension at an appropriate rate. This rate is obtained by establishing the minimax rate optimality of our tests over a certain class of alternatives. Several simulated and benchmark data sets are analyzed to compare the performance of these proposed tests with the state-of-the-art methods that can be used for testing the equality of two high-dimensional probability distributions.

math.ST