SearcharxivSearch

arXiv subjects

Soumendra Lahiri

Publications and source records attributed to Soumendra Lahiri.

11 recordsLinked to original sources

Necessary and sufficient conditions for high dimensional Central Limit Theorem under moment conditions

High dimensional central limit theorems (the CLTs) have been extensively studied in recent years under a variety of sufficient moment conditions connecting the dimension growth rate with the tail decay rate. In this article, we investigate whether the existing moment conditions are also necessary under the independence of the components. We consider four exhaustive classes, viz. when underlying random variables (I) have all polynomial moments, (II) have some polynomial moment of order higher than two, (III) have only second moment but no polynomial moment higher than two exists, and (IV) have infinite second moment, but belong to the domain of attraction of normal distribution. We find the optimal growth rate of the dimension with respect to sample size in the high dimensional CLTs over hyper-rectangles. More precisely, we derive necessary and sufficient moment conditions for the validity of the the CLT over hyper-rectangles in each of the four regimes listed above, showing that the CLT may hold under much weaker conditions compared to those considered in the existing literature.

math.PR

Bot Identification in Social Media

Escalating proliferation of inorganic accounts, commonly known as bots, within the digital ecosystem represents an ongoing and multifaceted challenge to online security, trustworthiness, and user experience. These bots, often employed for the dissemination of malicious propaganda and manipulation of public opinion, wield significant influence in social media spheres with far-reaching implications for electoral processes, political campaigns and international conflicts. Swift and accurate identification of inorganic accounts is of paramount importance in mitigating their detrimental effects. This research paper focuses on the identification of such accounts and explores various effective methods for their detection through machine learning techniques. In response to the pervasive presence of bots in the contemporary digital landscape, this study extracts temporal and semantic features from tweet behaviors and proposes a bot detection algorithm utilizing fundamental machine learning approaches, including Support Vector Machines (SVM) and k-means clustering. Furthermore, the research ranks the importance of these extracted features for each detection technique and also provides uncertainty quantification using a distribution free method, called the conformal prediction, thereby contributing to the development of effective strategies for combating the prevalence of inorganic accounts in social media platforms.

stat.AP

Fitting Sparse Markov Models to Categorical Time Series Using Convex Clustering

Higher-order Markov chains are frequently used to model categorical time series. However, a major problem with fitting such models is the exponentially growing number of parameters in the model order. A popular approach to parsimonious modeling is to use a Variable Length Markov Chain (VLMC), which determines relevant contexts (recent pasts) of variable orders and forms a context tree. A more general parsimonious modeling approach is given by Sparse Markov Models (SMMs), where all possible histories of order $m$ are partitioned such that the transition probability vectors are identical for the histories belonging to any particular group. In this paper, we develop an elegant method of fitting SMMs based on convex clustering and regularization. The regularization parameter is selected using the BIC criterion. Theoretical results establish model selection consistency of our method for large sample size. Extensive simulation results under different set-ups are presented to study finite sample performance of the method. Real data analysis on modelling and classifying disease sub-types demonstrates the applicability of our method as well.

stat.ME

A Bootstrap-based Method for Testing Network Similarity

This paper studies the matched network inference problem, where the goal is to determine if two networks, defined on a common set of nodes, exhibit a specific form of stochastic similarity. Two notions of similarity are considered: (i) equality, i.e., testing whether the networks arise from the same random graph model, and (ii) scaling, i.e., testing whether their probability matrices are proportional for some unknown scaling constant. We develop a testing framework based on a parametric bootstrap approach and a Frobenius norm-based test statistic. The proposed approach is highly versatile as it covers both the equality and scaling problems, and ensures adaptability under various model settings, including stochastic blockmodels, Chung-Lu models, and random dot product graph models. We establish theoretical consistency of the proposed tests and demonstrate their empirical performance through extensive simulations under a wide range of model classes. Our results establish the flexibility and computational efficiency of the proposed method compared to existing approaches. We also report a real-world application involving the Aarhus network dataset, which reveals meaningful sociological patterns across different communication layers.

stat.ME

Polyspectral Mean Estimation of General Nonlinear Processes

Higher-order spectra (or polyspectra), defined as the Fourier Transform of a stationary process' autocumulants, are useful in the analysis of nonlinear and non Gaussian processes. Polyspectral means are weighted averages over Fourier frequencies of the polyspectra, and estimators can be constructed from analogous weighted averages of the higher-order periodogram (a statistic computed from the data sample's discrete Fourier Transform). We derive the asymptotic distribution of a class of polyspectral mean estimators, obtaining an exact expression for the limit distribution that depends on both the given weighting function as well as on higher-order spectra. Secondly, we use bispectral means to define a new test of the linear process hypothesis. Simulations document the finite sample properties of the asymptotic results. Two applications illustrate our results' utility: we test the linear process hypothesis for a Sunspot time series, and for the Gross Domestic Product we conduct a clustering exercise based on bispectral means with different weight functions.

math.ST

Probabilistic Guarantees of Stochastic Recursive Gradient in Non-Convex Finite Sum Problems

This paper develops a new dimension-free Azuma-Hoeffding type bound on summation norm of a martingale difference sequence with random individual bounds. With this novel result, we provide high-probability bounds for the gradient norm estimator in the proposed algorithm Prob-SARAH, which is a modified version of the StochAstic Recursive grAdient algoritHm (SARAH), a state-of-art variance reduced algorithm that achieves optimal computational complexity in expectation for the finite sum problem. The in-probability complexity by Prob-SARAH matches the best in-expectation result up to logarithmic factors. Empirical experiments demonstrate the superior probabilistic performance of Prob-SARAH on real datasets compared to other popular algorithms.

stat.ML

Online Bootstrap Inference with Nonconvex Stochastic Gradient Descent Estimator

In this paper, we investigate the theoretical properties of stochastic gradient descent (SGD) for statistical inference in the context of nonconvex optimization problems, which have been relatively unexplored compared to convex settings. Our study is the first to establish provable inferential procedures using the SGD estimator for general nonconvex objective functions, which may contain multiple local minima. We propose two novel online inferential procedures that combine SGD and the multiplier bootstrap technique. The first procedure employs a consistent covariance matrix estimator, and we establish its error convergence rate. The second procedure approximates the limit distribution using bootstrap SGD estimators, yielding asymptotically valid bootstrap confidence intervals. We validate the effectiveness of both approaches through numerical experiments. Furthermore, our analysis yields an intermediate result: the in-expectation error convergence rate for the original SGD estimator in nonconvex settings, which is comparable to existing results for convex problems. We believe this novel finding holds independent interest and enriches the literature on optimization and statistical inference.

stat.ML

Fast parameter estimation of Generalized Extreme Value distribution using Neural Networks

The heavy-tailed behavior of the generalized extreme-value distribution makes it a popular choice for modeling extreme events such as floods, droughts, heatwaves, wildfires, etc. However, estimating the distribution's parameters using conventional maximum likelihood methods can be computationally intensive, even for moderate-sized datasets. To overcome this limitation, we propose a computationally efficient, likelihood-free estimation method utilizing a neural network. Through an extensive simulation study, we demonstrate that the proposed neural network-based method provides Generalized Extreme Value (GEV) distribution parameter estimates with comparable accuracy to the conventional maximum likelihood method but with a significant computational speedup. To account for estimation uncertainty, we utilize parametric bootstrapping, which is inherent in the trained network. Finally, we apply this method to 1000-year annual maximum temperature data from the Community Climate System Model version 3 (CCSM3) across North America for three atmospheric concentrations: 289 ppm $\mathrm{CO}_2$ (pre-industrial), 700 ppm $\mathrm{CO}_2$ (future conditions), and 1400 ppm $\mathrm{CO}_2$, and compare the results with those obtained using the maximum likelihood approach.

stat.ML

Conditional Randomization Rank Test

We propose a new method named the Conditional Randomization Rank Test (CRRT) for testing conditional independence of a response variable Y and a covariate variable X, conditional on the rest of the covariates Z. The new method generalizes the Conditional Randomization Test (CRT) of [CFJL18] by exploiting the knowledge of the conditional distribution of X|Z and is a conditional sampling based method that is easy to implement and interpret. In addition to guaranteeing exact type 1 error control, owing to a more flexible framework, the new method markedly outperforms the CRT in computational efficiency. We establish bounds on the probability of type 1 error in terms of total variation norm and also in terms of observed Kullback-Leibler divergence when the conditional distribution of X|Z is misspecified. We validate our theoretical results by extensive simulations and show that our new method has considerable advantages over other existing conditional sampling based methods when we take both power and computational efficiency into consideration.

stat.ME

Central Limit Theorem in High Dimensions : The Optimal Bound on Dimension Growth Rate

In this article, we try to give an answer to the simple question: ``\textit{What is the critical growth rate of the dimension $p$ as a function of the sample size $n$ for which the Central Limit Theorem holds uniformly over the collection of $p$-dimensional hyper-rectangles ?''}. Specifically, we are interested in the normal approximation of suitably scaled versions of the sum $\sum_{i=1}^{n}X_i$ in $\mathcal{R}^p$ uniformly over the class of hyper-rectangles $\mathcal{A}^{re}=\{\prod_{j=1}^{p}[a_j,b_j]\cap\mathcal{R}:-\infty\leq a_j\leq b_j \leq \infty, j=1,\ldots,p\}$, where $X_1,\dots,X_n$ are independent $p-$dimensional random vectors with each having independent and identically distributed (iid) components. We investigate the critical cut-off rate of $\log p$ below which the uniform central limit theorem (CLT) holds and above which it fails. According to some recent results of Chernozukov et al. (2017), it is well known that the CLT holds uniformly over $\mathcal{A}^{re}$ if $\log p=o\big(n^{1/7}\big)$. They also conjectured that for CLT to hold uniformly over $\mathcal{A}^{re}$, the optimal rate is $\log p = o\big(n^{1/3}\big)$. We show instead that under some conditions, the CLT holds uniformly over $\mathcal{A}^{re}$, when $\log p=o\big(n^{1/2}\big)$. More precisely, we show that if $\log p =ε\sqrt{n}$ for some sufficiently small $ε>0$, the normal approximation is valid with an error $ε$, uniformly over $\mathcal{A}^{re}$. Further, we show by an example that the uniform CLT over $\mathcal{A}^{re}$ fails if $\limsup_{t\rightarrow \infty} n^{-(1/2+δ)} \log p >0$ for some $δ>0$. Hence the critical rate of the growth of $p$ for the validity of the CLT is given by $\log p=o\big(n^{1/2}\big)$.

math.PR

Inference in partially identified models with many moment inequalities using Lasso

This paper considers inference in a partially identified moment (in)equality model with many moment inequalities. We propose a novel two-step inference procedure that combines the methods proposed by Chernozhukov, Chetverikov and Kato (2018a) (CCK18, hereafter) with a first step moment inequality selection based on the Lasso. Our method controls asymptotic size uniformly, both in underlying parameter and data distribution. Also, the power of our method compares favorably with that of the corresponding two-step method in CCK18 for large parts of the parameter space, both in theory and in simulations. Finally, we show that our Lasso-based first step can be implemented by thresholding standardized sample averages, and so it is straightforward to implement.

math.ST