SearcharxivSearch

arXiv subjects

Anirvan Chakraborty

Publications and source records attributed to Anirvan Chakraborty.

13 recordsLinked to original sources

Two-stage Ensemble Clustering of Functional Data Using Random Projections

We propose a computationally simple framework for clustering functional data based on Gaussian-process-generated random projections. In this approach, each curve is first projected onto a large collection of independent Gaussian process realizations. The resulting high-dimensional representations are clustered using the Mean Absolute Difference of Distances (MADD), a dissimilarity measure well suited for high-dimensional settings. A population-level analysis of this dissimilarity provides insight into how random projections help capture distributional differences between functional populations. We introduce a second stage of clustering to additionally leverage on data-driven projection directions. Thus, in Stage I, an initial clustering is obtained using a set of prespecified projection families. In Stage II, this partition is refined by constructing Gaussian random projections based on an estimated covariance operator that uses the first stage of cluster labels. Finally, a normalized cost function is used to select the optimal clustering among candidate solutions. The proposed clustering algorithm is broadly applicable to diverse functional data regimes including irregular and partially observed data. Through extensive simulations and real-data applications, we show that the proposed method achieves a high degree of accuracy and outperforms many of the state-of-the-art methods across a wide range of functional data settings.

stat.ME

Consistent detection and estimation of multiple structural changes in functional data: unsupervised and supervised approaches

We develop algorithms for detecting multiple changepoints in functional data when the number of changepoints is unknown (unsupervised case), when it is specified apriori (supervised case), and when certain bounds are available (semi-supervised case). These algorithms utilize the maximum mean discrepancy (MMD) measure between distributions on Hilbert spaces. We develop an oracle analysis of the changepoint detection problem which reveals an interesting relationship between the true changepoint locations and the local maxima of the oracle MMD curve. The proposed algorithms are shown to detect general distributional changes by exploiting this connection. In the unsupervised case, we test the significance of a potential changepoint and establish its consistency under the single changepoint setting. We investigate the strong consistency of the changepoint estimators in both single and multiple changepoint settings. In both supervised and semi-supervised scenarios, we include a step to merge consecutive groups that are similar to appropriately utilize the prior information about the number of changepoints. In the supervised scenario, the algorithm satisfies an order-preserving property: the estimated changepoints are contained in the true set of changepoints in the underspecified case, while they contain the true set under overspecification. We evaluate the performance of the algorithms on a variety of datasets demonstrating the superiority of the proposed algorithms compared to some of the existing methods.

stat.ME

Near-perfect Clustering Based on Recursive Binary Splitting Using Max-MMD

We develop novel clustering algorithms for functional data when the number of clusters $K$ is unknown and also when it is prefixed. These algorithms are developed based on the Maximum Mean Discrepancy (MMD) measure between two sets of observations. The algorithms recursively use a binary splitting strategy to partition the dataset into two subgroups such that they are maximally separated in terms of an appropriate weighted MMD measure. When $K$ is unknown, the proposed clustering algorithm has an additional step to check whether a group of observations obtained by the binary splitting technique consists of observations from a single population. We also obtain a bonafide estimator of $K$ using this algorithm. When $K$ is prefixed, a modification of the previous algorithm is proposed which consists of an additional step of merging subgroups which are similar in terms of the weighted MMD distance. The theoretical properties of the proposed algorithms are investigated in an oracle scenario that requires the knowledge of the empirical distributions of the observations from different populations involved. In this setting, we prove that the algorithm proposed when $K$ is unknown achieves perfect clustering while the algorithm proposed when $K$ is prefixed has the perfect order preserving (POP) property. Extensive real and simulated data analyses using a variety of models having location difference as well as scale difference show near-perfect clustering performance of both the algorithms which improve upon the state-of-the-art clustering methods for functional data.

stat.ME

Functional Registration and Local Variations: Identifiability, Rank, and Tuning

We develop theory and methodology for the problem of nonparametric registration of functional data that have been subjected to random deformation (warping) of their time scale. The separation of this phase variation ("horizontal" variation) from the amplitude variation ("vertical" variation) is crucial in order to properly conduct further analyses, which otherwise can be severely distorted. We determine precise nonparametric conditions under which the two forms of variation are identifiable. These show that the identifiability delicately depends on the underlying rank. By means of several counterexamples, we demonstrate that our conditions are sharp if one wishes a genuinely nonparametric setup; and in doing so we caution that popular remedies such as structural assumptions or roughness penalties can easily fail. We then propose a nonparametric registration method based on a "local variation measure", the main element in elucidating identifiability. A key advantage of the method is that it is free of any tuning or penalisation parameters regulating the amount of alignment, thus circumventing the problem of over/under-registration often encountered in practice. We provide asymptotic theory for the resulting estimators under the identifiable regime, but also under mild departures from identifiability, quantifying the resulting bias in terms of the amplitude variation's spectral gap.

stat.ME

Testing for the Rank of a Covariance Operator

How can we discern whether the covariance operator of a stochastic process is of reduced rank, and if so, what its precise rank is? And how can we do so at a given level of confidence? This question is central to a great deal of methods for functional data, which require low-dimensional representations whether by functional PCA or other methods. The difficulty is that the determination is to be made on the basis of i.i.d. replications of the process observed discretely and with measurement error contamination. This adds a ridge to the empirical covariance, obfuscating the underlying dimension. We build a matrix-completion inspired test statistic that circumvents this issue by measuring the best possible least square fit of the empirical covariance's off-diagonal elements, optimised over covariances of given finite rank. For a fixed grid of sufficiently large size, we determine the statistic's asymptotic null distribution as the number of replications grows. We then use it to construct a bootstrap implementation of a stepwise testing procedure controlling the family-wise error rate corresponding to the collection of hypotheses formalising the question at hand. Under minimal regularity assumptions we prove that the procedure is consistent and that its bootstrap implementation is valid. The procedure circumvents smoothing and associated smoothing parameters, is indifferent to measurement error heteroskedasticity, and does not assume a low-noise regime. An extensive simulation study reveals an excellent practical performance, stably across a wide range of settings, and the procedure is further illustrated by means of two data analyses.

stat.ME

Regression with genuinely functional errors-in-covariates

Contamination of covariates by measurement error is a classical problem in multivariate regression, where it is well known that failing to account for this contamination can result in substantial bias in the parameter estimators. The nature and degree of this effect on statistical inference is also understood to crucially depend on the specific distributional properties of the measurement error in question. When dealing with functional covariates, measurement error has thus far been modelled as additive white noise over the observation grid. Such a setting implicitly assumes that the error arises purely at the discrete sampling stage, otherwise the model can only be viewed in a weak (stochastic differential equation) sense, white noise not being a second-order process. Departing from this simple distributional setting can have serious consequences for inference, similar to the multivariate case, and current methodology will break down. In this paper, we consider the case when the additive measurement error is allowed to be a valid stochastic process. We propose a novel estimator of the slope parameter in a functional linear model, for scalar as well as functional responses, in the presence of this general measurement error specification. The proposed estimator is inspired by the multivariate regression calibration approach, but hinges on recent advances on matrix completion methods for functional data in order to handle the nontrivial (and unknown) error covariance structure. The asymptotic properties of the proposed estimators are derived. We probe the performance of the proposed estimator of slope using simulations and observe that it substantially improves upon the spectral truncation estimator based on the erroneous observations, i.e., ignoring measurement error. We also investigate the behaviour of the estimators on a real dataset on hip and knee angle curves during a gait cycle.

stat.ME

Hybrid Regularisation of Functional Linear Models

We consider the problem of estimating the slope function in a functional regression with a scalar response and a functional covariate. This central problem of functional data analysis is well known to be ill-posed, thus requiring a regularised estimation procedure. The two most commonly used approaches are based on spectral truncation or Tikhonov regularisation of the empirical covariance operator. In principle, Tikhonov regularisation is the more canonical choice. Compared to spectral truncation, it is robust to eigenvalue ties, while it attains the optimal minimax rate of convergence in the mean squared sense, and not just in a concentration probability sense. In this paper, we show that, surprisingly, one can strictly improve upon the performance of the Tikhonov estimator in finite samples by means of a linear estimator, while retaining its stability and asymptotic properties by combining it with a form of spectral truncation. Specifically, we construct an estimator that additively decomposes the functional covariate by projecting it onto two orthogonal subspaces defined via functional PCA; it then applies Tikhonov regularisation to the one component, while leaving the other component unregularised. We prove that when the covariate is Gaussian, this hybrid estimator uniformly improves upon the MSE of the Tikhonov estimator in a non-asymptotic sense, effectively rendering it inadmissible. This domination is shown to also persist under discrete observation of the covariate function. The hybrid estimator is linear, straightforward to construct in practice, and with no computational overhead relative to the standard regularisation methods. By means of simulation, it is shown to furnish sizeable gains even for modest sample sizes.

stat.ME

Tests for high dimensional data based on means, spatial signs and spatial ranks

Tests based on sample mean vectors and sample spatial signs have been studied in the recent literature for high dimensional data with the dimension larger than the sample size. For suitable sequences of alternatives, we show that the powers of the mean based tests and the tests based on spatial signs and ranks tend to be same as the data dimension grows to infinity for any sample size, when the coordinate variables satisfy appropriate mixing conditions. Further, their limiting powers do not depend on the heaviness of the tails of the distributions. This is in striking contrast to the asymptotic results obtained in the classical multivariate setup. On the other hand, we show that in the presence of stronger dependence among the coordinate variables, the spatial sign and rank based tests for high dimensional data can be asymptotically more powerful than the mean based tests if in addition to the data dimension, the sample size also grows to infinity. The sizes of some mean based tests for high dimensional data studied in the recent literature are observed to be significantly different from their nominal levels. This is due to the inadequacy of the asymptotic approximations used for the distributions of those test statistics. However, our asymptotic approximations for the tests based on spatial signs and ranks are observed to work well when the tests are applied on a variety of simulated and real datasets.

math.ST

Paired sample tests in infinite dimensional spaces

The sign and the signed-rank tests for univariate data are perhaps the most popular nonparametric competitors of the t test for paired sample problems. These tests have been extended in various ways for multivariate data in finite dimensional spaces. These extensions include tests based on spatial signs and signed ranks, which have been studied extensively by Hannu Oja and his coauthors. They showed that these tests are asymptotically more powerful than Hotelling's $T^{2}$ test under several heavy tailed distributions. In this paper, we consider paired sample tests for data in infinite dimensional spaces based on notions of spatial sign and spatial signed rank in such spaces. We derive their asymptotic distributions under the null hypothesis and under sequences of shrinking location shift alternatives. We compare these tests with some mean based tests for infinite dimensional paired sample data. We show that for shrinking location shift alternatives, the proposed tests are asymptotically more powerful than the mean based tests for some heavy tailed distributions and even for some Gaussian distributions in infinite dimensional spaces. We also investigate the performance of different tests using some simulated data.

stat.ME

The spatial distribution in infinite dimensional spaces and related quantiles and depths

The spatial distribution has been widely used to develop various nonparametric procedures for finite dimensional multivariate data. In this paper, we investigate the concept of spatial distribution for data in infinite dimensional Banach spaces. Many technical difficulties are encountered in such spaces that are primarily due to the noncompactness of the closed unit ball. In this work, we prove some Glivenko-Cantelli and Donsker-type results for the empirical spatial distribution process in infinite dimensional spaces. The spatial quantiles in such spaces can be obtained by inverting the spatial distribution function. A Bahadur-type asymptotic linear representation and the associated weak convergence results for the sample spatial quantiles in infinite dimensional spaces are derived. A study of the asymptotic efficiency of the sample spatial median relative to the sample mean is carried out for some standard probability distributions in function spaces. The spatial distribution can be used to define the spatial depth in infinite dimensional Banach spaces, and we study the asymptotic properties of the empirical spatial depth in such spaces. We also demonstrate the spatial quantiles and the spatial depth using some real and simulated functional data.

stat.ME

A Wilcoxon-Mann-Whitney type test for infinite dimensional data

The Wilcoxon-Mann-Whitney test is a robust competitor of the t-test in the univariate setting. For finite dimensional multivariate data, several extensions of the Wilcoxon-Mann-Whitney test have been shown to have better performance than Hotelling's $T^{2}$ test for many non-Gaussian distributions of the data. In this paper, we study a Wilcoxon-Mann-Whitney type test based on spatial ranks for data in infinite dimensional spaces. We demonstrate the performance of this test using some real and simulated datasets. We also investigate the asymptotic properties of the proposed test and compare the test with a wide range of competing tests.

stat.ME

On data depth in infinite dimensional spaces

The concept of data depth leads to a center-outward ordering of multivariate data, and it has been effectively used for developing various data analytic tools. While different notions of depth were originally developed for finite dimensional data, there have been some recent attempts to develop depth functions for data in infinite dimensional spaces. In this paper, we consider some notions of depth in infinite dimensional spaces and study their properties under various stochastic models. Our analysis shows that some of the depth functions available in the literature have degenerate behaviour for some commonly used probability distributions in infinite dimensional spaces of sequences and functions. As a consequence, they are not very useful for the analysis of data satisfying such infinite dimensional probability models. However, some modified versions of those depth functions as well as an infinite dimensional extension of the spatial depth do not suffer from such degeneracy, and can be conveniently used for analyzing infinite dimensional data.

stat.ME

The deepest point for distributions in infinite dimensional spaces

Identification of the center of a data cloud is one of the basic problems in statistics. One popular choice for such a center is the median, and several versions of median in finite dimensional spaces have been studied in the literature. In particular, medians based on different notions of data depth have been extensively studied by many researchers, who defined median as the point, where the depth function attains its maximum value. In other words, the median is the deepest point in the sample space according to that definition. In this paper, we investigate the deepest point for probability distributions in infinite dimensional spaces. We show that for some well-known depth functions like the band depth and the half-region depth in function spaces, there may not be any meaningful deepest point for many well-known and commonly used probability models. On the other hand, certain modified versions of those depth functions as well as the spatial depth function, which can be defined in any Hilbert space, lead to some useful notions of the deepest point with nice geometric and statistical properties. The empirical versions of those deepest points can be conveniently computed for functional data, and we demonstrate this using some simulated and real data sets.

math.ST