SearcharxivSearch

arXiv subjects

Somabha Mukherjee

Publications and source records attributed to Somabha Mukherjee.

At least 19 recordsLinked to original sources

Scaling Limits for Ising Models on Inhomogeneous Random Graphs and Applications

In this paper, we derive quenched scaling limits for linear functionals and the empirical spin field of Ising models on inhomogeneous random graphs generated by a graphon (encompassing both dense and sparse graphs), in the high-temperature regime. We first prove a joint central limit theorem (CLT) for finite collections of linear statistics of the spin configurations, where the limiting covariance is characterized by the resolvent of the associated graphon integral operator. Building on this result, we establish functional CLTs for the average magnetization and for the spin field indexed by suitable classes of regular test functions. We further prove convergence of the full empirical spin field, viewed as a random generalized function in negative Sobolev spaces. These scaling limits provide applications to both Bayesian neural networks and causal inference. Specifically, for the former, we derive infinite-width Gaussian-process limits for two-layer Bayesian neural networks with Ising-dependent output-layer signs, while for the latter, we establish the asymptotic normality of Hájek estimators for average treatment effects under network interference.

math.PR

Ising Models on Inhomogeneous Random Graphs: Inference, Local Asymptotic Minimaxity, and Limit of Experiments

In this paper, we develop an inferential framework with sharp asymptotic optimality guarantees for Ising models on inhomogeneous random graphs in the subcritical parameter regime. We begin by characterizing the asymptotic distribution of the maximum likelihood (ML) estimate of the natural parameter, based on a single sample from the underlying model, covering both sparse and dense network regimes. Next, to overcome the computational intractability of the ML method, we propose a simple closed-form estimate obtained from a one-step approximation to the likelihood equation. We show that this estimate attains the same asymptotic distribution and variance as the ML estimate, thereby yielding a computationally efficient and asymptotically valid confidence interval for the natural parameter. We complement these inferential results by establishing a Hájek--Le Cam-type local asymptotic minimax theorem, showing that the proposed estimate achieves the smallest possible asymptotic maximum risk, both in rate and in leading constant, over shrinking neighborhoods of the true parameter. We also derive the corresponding limit of experiments. To the best of our knowledge, these are among the first sharp asymptotic optimality results for network-dependent data. Finally, we study goodness-of-fit testing for the natural parameter, deriving the local power of the likelihood ratio test and minimax detection rates. Our analysis relies on new fluctuation results for the sufficient statistic (Hamiltonian) and for the random partition function of Ising models on inhomogeneous random graphs, which are of independent interest.

math.ST

Joint Estimation in Potts Model

In this paper, we study estimation of parameters in a two-parameter Potts model with $q$ colors and coupling matrix $A_N$. We characterize concrete sufficient conditions for existence of the pseudo-likelihood estimator of the Potts model, in terms of the local magnetic fields, and give sufficient conditions for the validity of the above characterization. We then provide sufficient criteria for estimation of both parameters at the optimal rate $\sqrt{N}$. In particular, if $A_N$ is the scaled adjacency matrix of a graph $G_N$, then we show that joint estimation is possible if either $G_N$ has bounded degree or is irregular. In contrast, we give an example of a graph sequence $G_N$ which is approximately regular and dense, where no consistent estimator exists. We also show that one-parameter estimation at the optimal rate $\sqrt{N}$ holds under much milder conditions when the other parameter is known. Along the way, we develop a concentration result for mean-field Potts models using the framework of nonlinear large deviations. Compared to the Ising case, our results for the Potts case require a novel analysis across multiple colors.

math.ST

CITS: Nonparametric Statistical Causal Modeling for High-Resolution Neural Time Series

Identifying causal interactions in complex dynamical systems is a fundamental challenge across the computational sciences. Existing functional connectivity methods capture correlations but not causation. While addressing directionality, popular causal inference tools such as Granger causality and the Peter-Clark algorithm rely on restrictive assumptions that limit their applicability to high-resolution time-series data, such as the large-scale recordings now standard in neuroscience. Here, we introduce CITS (Causal Inference in Time Series), a nonparametric framework for inferring statistically causal structure from multivariate time series. CITS models dynamics using a structural causal model of arbitrary Markov order and statistical tests for lagged conditional independence. We prove consistency under mild assumptions and demonstrate superior accuracy over state-of-the-art baselines across simulated linear, nonlinear, and recurrent neural network benchmarks. Applying CITS to large-scale neuronal recordings from the mouse visual cortex, thalamus, and hippocampus, we uncover stimulus-specific causal pathways and inter-regional hierarchies that align with known anatomy while revealing new functional insights. We further highlight CITS ability in accurately identifying conditional dependencies within small inferred neuronal motifs. These results establish CITS as a theoretically grounded and empirically validated method for discovering interpretable statistically causal networks in neural time series. Beyond neuroscience, the framework is broadly applicable to causal discovery in complex temporal systems across domains.

q-bio.NC

Robust Estimation for Dependent Binary Network Data

We consider the problem of learning the interaction strength between the nodes of a network based on dependent binary observations residing on these nodes, generated from a Markov Random Field (MRF). Since these observations can possibly be corrupted/noisy in larger networks in practice, it is important to robustly estimate the parameters of the underlying true MRF to account for such inherent contamination in observed data. However, it is well-known that classical likelihood and pseudolikelihood based approaches are highly sensitive to even a small amount of data contamination. So, in this paper, we propose a density power divergence (DPD) based robust generalization of the computationally efficient maximum pseudolikelihood (MPL) estimator of the interaction strength parameter, and derive its rate of consistency under the pure model. Along the way, we establish consistency and asymptotics for a class of general $Z$-estimators, covering our proposed DPD based estimators, under flexible assumptions that hold for a substantial class of standard models. To the best of our knowledge, these are the first central limit theorems for the class of general $Z$-estimators in such settings. Moreover, we show that the gross error sensitivities of the proposed DPD based estimators are significantly smaller than that of the MPL estimator, thereby theoretically justifying the greater (local) robustness of the former under contaminated settings. Finally, we demonstrate the superior (finite sample) performance of the DPD based variants over the traditional MPL estimator in a number of synthetically generated contaminated network datasets, and apply them to learn the network interaction strength in several real datasets from diverse domains of social science, neurobiology and genomics.

stat.ME

Mixing Phases and Metastability for the Glauber Dynamics on the p-Spin Curie-Weiss Model

The Glauber dynamics for the classical $2$-spin Curie-Weiss model on $N$ nodes with inverse temperature $β$ and zero external field is known to mix in time $Θ(N\log N)$ for $β< \frac{1}{2}$, in time $Θ(N^{3/2})$ at $β= \frac{1}{2}$, and in time $\exp(Ω(N))$ for $β>\frac{1}{2}$. In this paper, we consider the $p$-spin generalization of the Curie-Weiss model with an external field $h$, and identify three disjoint regions almost exhausting the parameter space, with the corresponding Glauber dynamics exhibiting three different orders of mixing times in these regions. The construction of these disjoint regions depends on the number of local maximizers of a certain function $H_{β,h,p}$, and the behavior of the second derivative of $H_{β,h,p}$ at such a local maximizer. Specifically, we show that if $H_{β,h,p}$ has a unique local maximizer $m_*$ with $H_{β,h,p}''(m_*) < 0$ and no other stationary point, then the Glauber dynamics mixes in time $Θ(N\log N)$, and if $H_{β,h,p}$ has multiple local maximizers, then the mixing time is $\exp(Ω(N))$. Finally, if $H_{β,h,p}$ has a unique local maximizer $m_*$ with $H_{β,h,p}''(m_*) = 0$, then the mixing time is $Θ(N^{3/2})$. We provide an explicit description of the geometry of these three different phases in the parameter space, and observe that the only portion of the parameter plane that is left out by the union of these three regions, is a one-dimensional curve, on which the function $H_{β,h,p}$ has a stationary inflection point. Finding out the exact order of the mixing time on this curve remains an open question. Finally, we show that if $H_{β,h,p}$ has multiple local maximizers (metastable states), then one can create a restricted version of the original Glauber dynamics, which still mixes in time $Θ(N\log N)$.

math.PR

Interaction Order Estimation in Tensor Curie-Weiss Models

In this paper, we consider the problem of estimating the interaction parameter $p$ of a $p$-spin Curie-Weiss model at inverse temperature $β$, given a single observation from this model. We show, by a contiguity argument, that joint estimation of the parameters $β$ and $p$ is impossible, which implies that estimation of $p$ is impossible if $β$ is unknown. These impossibility results are also extended to the more general $p$-spin Erdős-Rényi Ising model. The situation is more delicate when $β$ is known. In this case, we show that there exists an increasing threshold function $β^*(p)$, such that for all $β$, consistent estimation of $p$ is impossible when $β^*(p) > β$, and for almost all $β$, consistent estimation of $p$ is possible for $β^*(p)<β$.

math.ST

High Dimensional Logistic Regression Under Network Dependence

Logistic regression is key method for modeling the probability of a binary outcome based on a collection of covariates. However, the classical formulation of logistic regression relies on the independent sampling assumption, which is often violated when the outcomes interact through an underlying network structure, such as over a temporal/spatial domain or on a social network. This necessitates the development of models that can simultaneously handle both the network `peer-effect' and the effect of high-dimensional covariates. In this paper, we develop a framework for incorporating such dependencies in a high-dimensional logistic regression model by introducing a quadratic interaction term, as in the Ising model, designed to capture the pairwise interactions from the underlying network. The resulting model can also be viewed as an Ising model, where the node-dependent external fields linearly encode the high-dimensional covariates. We propose a penalized maximum pseudo-likelihood method for estimating the network peer-effect and the effect of the covariates (the regression coefficients), which, in addition to handling the high-dimensionality of the parameters, conveniently avoids the computational intractability of the maximum likelihood approach. Under various standard regularity conditions, we show that the corresponding estimate attains the classical high-dimensional rate of consistency. Our results imply that even under network dependence it is possible to consistently estimate the model parameters at the same rate as in classical (independent) logistic regression, when the true parameter is sparse and the underlying network is not too dense. We also develop an efficient algorithm for computing the estimates and validate our theoretical results in numerical experiments. An application to selecting genes in clustering spatial transcriptomics data is also discussed.

math.ST

Fluctuations of Quadratic Chaos

In this paper we characterize all distributional limits of the random quadratic form $T_n =\sum_{1\le u< v\le n} a_{u, v} X_u X_v$, where $((a_{u, v}))_{1\le u,v\le n}$ is a $\{0, 1\}$-valued symmetric matrix with zeros on the diagonal and $X_1, X_2, \ldots, X_n$ are i.i.d.~ mean $0$ variance $1$ random variables with common distribution function $F$. In particular, we show that any distributional limit of $S_n:=T_n/\sqrt{\mathrm{Var}[T_n]}$ can be expressed as the sum of three independent components: a Gaussian, a (possibly) infinite weighted sum of independent centered chi-squares, and a Gaussian mixture with a random variance. As a consequence, we prove a fourth moment theorem for the asymptotic normality of $S_n$, which applies even when $F$ does not have finite fourth moment. More formally, we show that $S_n$ converges to $N(0, 1)$ if and only if the fourth moment of $S_n$ (appropriately truncated when $F$ does not have finite fourth moment) converges to 3 (the fourth moment of the standard normal distribution).

math.PR

Feature Selection in High-dimensional Spaces Using Graph-Based Methods

High-dimensional feature selection is a central problem in a variety of application domains such as machine learning, image analysis, and genomics. In this paper, we propose graph-based tests as a useful basis for feature selection. We describe an algorithm for selecting informative features in high-dimensional data, where each observation comes from one of $K$ different distributions. Our algorithm can be applied in a completely nonparametric setup without any distributional assumptions on the data, and it aims at outputting those features in the data, that contribute the most to the overall distributional variation. At the heart of our method is the recursive application of distribution-free graph-based tests on subsets of the feature set, located at different depths of a hierarchical clustering tree constructed from the data. Our algorithm recovers all truly contributing features with high probability, while ensuring optimal control on false-discovery. We show the superior performance of our method over other existing ones through synthetic data, and demonstrate the utility of this method on several real-life datasets from the domains of climate change and biology, wherein our algorithm is not only able to detect known features expected to be associated with the underlying process, but also discovers novel targets that can be subsequently studied.

stat.ME

Interaction Screening and Pseudolikelihood Approaches for Tensor Learning in Ising Models

In this paper, we study two well known methods of Ising structure learning, namely the pseudolikelihood approach and the interaction screening approach, in the context of tensor recovery in $k$-spin Ising models. We show that both these approaches, with proper regularization, retrieve the underlying hypernetwork structure using a sample size logarithmic in the number of network nodes, and exponential in the maximum interaction strength and maximum node-degree. We also track down the exact dependence of the rate of tensor recovery on the interaction order $k$, that is allowed to grow with the number of samples and nodes, for both the approaches. We then provide a comparative discussion of the performance of the two approaches based on simulation studies, which also demonstrates the exponential dependence of the tensor recovery rate on the maximum coupling strength. Our tensor recovery methods are then applied on gene data taken from the Curated Microarray Database (CuMiDa), where we focus on understanding the important genes related to hepatocellular carcinoma.

stat.ME

Rates of Convergence of the Magnetization in the Tensor Curie-Weiss Potts Model

In this paper, we derive distributional convergence rates for the magnetization vector and the maximum pseudolikelihood estimator of the inverse temperature parameter in the tensor Curie-Weiss Potts model. Limit theorems for the magnetization vector have been derived recently in Bhowal and Mukherjee (2023), where several phase transition phenomena in terms of the scaling of the (centered) magnetization and its asymptotic distribution were established, depending upon the position of the true parameters in the parameter space. In the current work, we establish Berry-Esseen type results for the magnetization vector, specifying its rate of convergence at these different phases. At most points in the parameter space, this rate is $N^{-1/2}$ ($N$ being the size of the Curie-Weiss network), while at some "special" points, the rate is either $N^{-1/4}$ or $N^{-1/6}$, depending upon the behavior of the fourth derivative of a certain "negative free energy function" at these special points. These results are then used to derive Berry-Esseen type bounds for the maximum pseudolikelihood estimator of the inverse temperature parameter whenever it lies above a certain criticality threshold.

math.PR

Moderate Deviation and Berry-Esseen Bounds in the $p$-Spin Curie-Weiss Model

Limit theorems for the magnetization in the $p$-spin Curie-Weiss model, for $p \geq 3$, has been derived recently by Mukherjee et al. (2021). In this paper, we strengthen these results by proving Cramér-type moderate deviation theorems and Berry-Esseen bounds for the magnetization (suitably centered and scaled). In particular, we show that the rate of convergence is $O(N^{-\frac{1}{2}})$ when the magnetization has asymptotically Gaussian fluctuations, and it is $O(N^{-\frac{1}{4}})$ when the fluctuations are non-Gaussian. As an application, we derive a Berry-Esseen bound for the maximum pseudolikelihood estimate of the inverse temperature in $p$-spin Curie-Weiss model with no external field, for all points in the parameter space where consistent estimation is possible.

math.PR

Optimal Confidence Bands for Shape-restricted Regression in Multidimensions

In this paper, we propose and study construction of confidence bands for shape-constrained regression functions when the predictor is multivariate. In particular, we consider the continuous multidimensional white noise model given by $d Y(\mathbf{t}) = n^{1/2} f(\mathbf{t}) \,d\mathbf{t} + d W(\mathbf{t})$, where $Y$ is the observed stochastic process on $[0,1]^d$ ($d\ge 1$), $W$ is the standard Brownian sheet on $[0,1]^d$, and $f$ is the unknown function of interest assumed to belong to a (shape-constrained) function class, e.g., coordinate-wise monotone functions or convex functions. The constructed confidence bands are based on local kernel averaging with bandwidth chosen automatically via a multivariate multiscale statistic. The confidence bands have guaranteed coverage for every $n$ and for every member of the underlying function class. Under monotonicity/convexity constraints on $f$, the proposed confidence bands automatically adapt (in terms of width) to the global and local (Hölder) smoothness and intrinsic dimensionality of the unknown $f$; the bands are also shown to be optimal in a certain sense. These bands have (almost) parametric ($n^{-1/2}$) widths when the underlying function has ``low-complexity'' (e.g., piecewise constant/affine).

math.ST

Inferring Causality from Time Series data based on Structural Causal Model and its application to Neural Connectomics

Inferring causation from time series data is of scientific interest in different disciplines, particularly in neural connectomics. While different approaches exist in the literature with parametric modeling assumptions, we focus on a non-parametric model for time series satisfying a Markovian structural causal model with stationary distribution and without concurrent effects. We show that the model structure can be used to its advantage to obtain an elegant algorithm for causal inference from time series based on conditional dependence tests, coined Causal Inference in Time Series (CITS) algorithm. We describe Pearson's partial correlation and Hilbert-Schmidt criterion as candidates for such conditional dependence tests that can be used in CITS for the Gaussian and non-Gaussian settings, respectively. We prove the mathematical guarantee of the CITS algorithm in recovering the true causal graph, under standard mixing conditions on the underlying time series. We also conduct a comparative evaluation of performance of CITS with other existing methodologies in simulated datasets. We then describe the utlity of the methodology in neural connectomics -- in inferring causal functional connectivity from time series of neural activity, and demonstrate its application to a real neurobiological dataset of electro-physiological recordings from the mouse visual cortex recorded by Neuropixel probes.

stat.ME

Least Squares Estimation of a Quasiconvex Regression Function

We develop a new approach for the estimation of a multivariate function based on the economic axioms of quasiconvexity (and monotonicity). On the computational side, we prove the existence of the quasiconvex constrained least squares estimator (LSE) and provide a characterization of the function space to compute the LSE via a mixed integer quadratic programme. On the theoretical side, we provide finite sample risk bounds for the LSE via a sharp oracle inequality. Our results allow for errors to depend on the covariates and to have only two finite moments. We illustrate the superior performance of the LSE against some competing estimators via simulation. Finally, we use the LSE to estimate the production function for the Japanese plywood industry and the cost function for hospitals across the US.

stat.ME

Limit Theorems and Phase Transitions in the Tensor Curie-Weiss Potts Model

In this paper, we derive results about the limiting distribution of the empirical magnetization vector and the maximum likelihood (ML) estimates of the natural parameters in the tensor Curie-Weiss Potts model. Our results reveal surprisingly new phase transition phenomena including the existence of a smooth curve in the interior of the parameter plane on which the magnetization vector and the ML estimates have mixture limiting distributions, the latter comprising of both continuous and discrete components, and a surprising superefficiency phenomenon of the ML estimates, which stipulates an $N^{-3/4}$ rate of convergence of the estimates to some non-Gaussian distribution at certain special points of one type and an $N^{-5/6}$ rate of convergence to some other non-Gaussian distribution at another special point of a different type. The last case can arise only for one particular value of the tuple of the tensor interaction order and the number of colors. These results are then used to derive asymptotic confidence intervals for the natural parameters at all points where consistent estimation is possible.

math.ST

Tensor Recovery in High-Dimensional Ising Models

The $k$-tensor Ising model is an exponential family on a $p$-dimensional binary hypercube for modeling dependent binary data, where the sufficient statistic consists of all $k$-fold products of the observations, and the parameter is an unknown $k$-fold tensor, designed to capture higher-order interactions between the binary variables. In this paper, we describe an approach based on a penalization technique that helps us recover the signed support of the tensor parameter with high probability, assuming that no entry of the true tensor is too close to zero. The method is based on an $\ell_1$-regularized node-wise logistic regression, that recovers the signed neighborhood of each node with high probability. Our analysis is carried out in the high-dimensional regime, that allows the dimension $p$ of the Ising model, as well as the interaction factor $k$ to potentially grow to $\infty$ with the sample size $n$. We show that if the minimum interaction strength is not too small, then consistent recovery of the entire signed support is possible if one takes $n = Ω((k!)^8 d^3 \log \binom{p-1}{k-1})$ samples, where $d$ denotes the maximum degree of the hypernetwork in question. Our results are validated in two simulation settings, and applied on a real neurobiological dataset consisting of multi-array electro-physiological recordings from the mouse visual cortex, to model higher-order interactions between the brain regions.

math.ST