SearcharxivSearch

arXiv subjects

Shubhadeep Chakraborty

Publications and source records attributed to Shubhadeep Chakraborty.

6 recordsLinked to original sources

High-dimensional Change-point Detection Using Generalized Homogeneity Metrics

Change-point detection is a classical problem in statistics. We address the problem of detecting abrupt changes in the data-generating distributions of a sequence of high-dimensional observations beyond the first two moments. This problem remains less explored, especially in the high-dimensional context, compared to detecting changes in the mean or the covariance structure. To the best of our knowledge, this is one of the first attempts to detect and localize general types of distributional changes in the high-dimensional regime. We develop a distance-based method to (i) test for the existence of a change-point, and (ii) identify the change-point locations in an independent sequence of high-dimensional observations. Our approach rests upon recent distance-based tests for the homogeneity of two high-dimensional distributions. We construct a single change-point test statistic based on a cumulative sum process in an embedded Hilbert space and rigorously derive its limiting null distribution and prove asymptotic consistency under the high-dimensional medium sample size (HDMSS) framework. Subsequently, we combine our statistics with the Narrowest-Over-Threshold (NOT) strategy to recursively estimate and test for multiple change-point locations. We also study a componentwise monotone-invariant, rank-based extension; because its pseudo-observations are pooled empirical mid-ranks and are therefore dependent, we present this version as a practically useful heuristic extension supported by simulation evidence rather than as a fully proved analogue of the original statistic. The superior performance of our methodology compared to existing procedures is illustrated via extensive simulation studies and an application to U.S. stock return data during the global financial crisis. The proposed method is implemented in the R package KDist, available at https://github.com/zhangxiany-tamu/KDist.

stat.ME

Hydrodynamics of the Fermi-Pasta-Ulam-Tsingou chain

We provide a pedagogical review of the hydrodynamics of the FPUT chain. There are three hydrodynamic fields corresponding to the conservation of mass, momentum and energy. We provide physically motivated derivations of the hydrodynamic equations at the levels of Euler and then Navier-Stokes-Fourier. Next we consider examples to test as to how successful the hydrodynamic description is in predicting the observed time evolution of nonequilibrium initial conditions such as domain walls and blasts. We find that in some cases there is good agreement of microscopic simulations with predictions from the Euler equations while, in several other cases, there is significant departure from the Euler predictions suggesting that the role of dissipation and noise is important in general.

cond-mat.stat-mech

Subgroup analysis in multi level hierarchical cluster randomized trials

Cluster or group randomized trials (CRTs) are increasingly used for both behavioral and system-level interventions, where entire clusters are randomly assigned to a study condition or intervention. Apart from the assigned cluster-level analysis, investigating whether an intervention has a differential effect for specific subgroups remains an important issue, though it is often considered an afterthought in pivotal clinical trials. Determining such subgroup effects in a CRT is a challenging task due to its inherent nested cluster structure. Motivated by a real-life HIV prevention CRT, we consider a three-level cross-sectional CRT, where randomization is carried out at the highest level and subgroups may exist at different levels of the hierarchy. We employ a linear mixed-effects model to estimate the subgroup-specific effects through their maximum likelihood estimators (MLEs). Consequently, we develop a consistent test for the significance of the differential intervention effect between two subgroups at different levels of the hierarchy, which is the key methodological contribution of this work. We also derive explicit formulae for sample size determination to detect a differential intervention effect between two subgroups, aiming to achieve a given statistical power in the case of a planned confirmatory subgroup analysis. The application of our methodology is illustrated through extensive simulation studies using synthetic data, as well as with real-world data from an HIV prevention CRT in The Bahamas.

stat.ME

Nonparametric causal structure learning in high dimensions

The PC and FCI algorithms are popular constraint-based methods for learning the structure of directed acyclic graphs (DAGs) in the absence and presence of latent and selection variables, respectively. These algorithms (and their order-independent variants, PC-stable and FCI-stable) have been shown to be consistent for learning sparse high-dimensional DAGs based on partial correlations. However, inferring conditional independences from partial correlations is valid if the data are jointly Gaussian or generated from a linear structural equation model -- an assumption that may be violated in many applications. To broaden the scope of high-dimensional causal structure learning, we propose nonparametric variants of the PC-stable and FCI-stable algorithms that employ the conditional distance covariance (CdCov) to test for conditional independence relationships. As the key theoretical contribution, we prove that the high-dimensional consistency of the PC-stable and FCI-stable algorithms carry over to general distributions over DAGs when we implement CdCov-based nonparametric tests for conditional independence. Numerical studies demonstrate that our proposed algorithms perform nearly as good as the PC-stable and FCI-stable for Gaussian distributions, and offer advantages in non-Gaussian graphical models.

stat.ME

A New Framework for Distance and Kernel-based Metrics in High Dimensions

The paper presents new metrics to quantify and test for (i) the equality of distributions and (ii) the independence between two high-dimensional random vectors. We show that the energy distance based on the usual Euclidean distance cannot completely characterize the homogeneity of two high-dimensional distributions in the sense that it only detects the equality of means and the traces of covariance matrices in the high-dimensional setup. We propose a new class of metrics which inherits the desirable properties of the energy distance and maximum mean discrepancy/(generalized) distance covariance and the Hilbert-Schmidt Independence Criterion in the low-dimensional setting and is capable of detecting the homogeneity of/completely characterizing independence between the low-dimensional marginal distributions in the high dimensional setup. We further propose t-tests based on the new metrics to perform high-dimensional two-sample testing/independence testing and study their asymptotic behavior under both high dimension low sample size (HDLSS) and high dimension medium sample size (HDMSS) setups. The computational complexity of the t-tests only grows linearly with the dimension and thus is scalable to very high dimensional data. We demonstrate the superior power behavior of the proposed tests for homogeneity of distributions and independence via both simulated and real datasets.

stat.ME

Distance Metrics for Measuring Joint Dependence with Application to Causal Inference

Many statistical applications require the quantification of joint dependence among more than two random vectors. In this work, we generalize the notion of distance covariance to quantify joint dependence among d >= 2 random vectors. We introduce the high order distance covariance to measure the so-called Lancaster interaction dependence. The joint distance covariance is then defined as a linear combination of pairwise distance covariances and their higher order counterparts which together completely characterize mutual independence. We further introduce some related concepts including the distance cumulant, distance characteristic function, and rank-based distance covariance. Empirical estimators are constructed based on certain Euclidean distances between sample elements. We study the large sample properties of the estimators and propose a bootstrap procedure to approximate their sampling distributions. The asymptotic validity of the bootstrap procedure is justified under both the null and alternative hypotheses. The new metrics are employed to perform model selection in causal inference, which is based on the joint independence testing of the residuals from the fitted structural equation models. The effectiveness of the method is illustrated via both simulated and real datasets.

stat.ME