SearcharxivSearch

arXiv subjects

Shirshendu Chatterjee

Publications and source records attributed to Shirshendu Chatterjee.

At least 19 recordsLinked to original sources

$\beta$-Skewed Maximal Spanning Forests

The Free $\mathbf{w}$-Maximal Spanning Forest (FMaxSF) is a weighted generalization of the classical Free Minimal Spanning Forest (FMSF) that is able to detect nonhyperfiniteness in percolation on nonunimodular graphs. We introduce a parameterized family of invariant random spanning forests that interpolates between these models. For every finite positive value of the parameter $\beta$, the construction retains many of the desired properties of FMaxSF while also admitting the finite-subtree forcing property of FMSF. We study local limits of these forests and, in particular, show that the small-$\beta$ limit of the wired variant coincides with FMSF if and only if $p_h=p_u$, where $p_h$ is the threshold for the existence of heavy clusters and $p_u$ is the uniqueness threshold for Bernoulli$(p)$ percolation. Finally, we show that the Free and the Wired $\mathbf{w}$-Maximal Spanning Forests may coincide even if $p_h<p_u$, providing a negative answer to a question of Terlov and Tim\'ar.

math.PR

Interface between competing random walks on a cycle

We consider a competition between two independent random walks on a cycle of length $N$. Each vertex is claimed by the walker that visits it first, and remains claimed thereafter. We prove that if the initial distance between the walkers is $d$, then the expected number of edges whose endpoints are claimed by different walkers is of order $\ln(1+N/d).$ This confirms the logarithmic dependence on $N/d$ predicted in Gomes Jr. et al. [Coloring of a one-dimensional lattice by two independent random walkers. Physica A: Statistical Mechanics and its Applications 225.1 (1996): 81-88].

math.PR

Convergence of $k$-point functions in high dimensional percolation

Consider critical Bernoulli percolation on $\mathbb{Z}^d$ for $d$ large; let $y_0, \dots, y_{k-1}$ be $k$ distinct points in $\mathbb{R}^d$. We prove that the probability that $\{\lfloor n y_i\rfloor\}_{i=0}^{k-1}$ all lie in the same open cluster, rescaled by an appropriate power of $n$, converges as $n \to \infty$ to an explicit constant. This confirms a conjecture of Aizenman and Newman.

math.PR

The effects of individual versus community-influenced isolation on SIS epidemic persistence on finite random graphs

The contact process, or the SIS epidemic model, is a continuous-time Markov process used to model the spread of a recurring infection on a graph. Each vertex is either healthy or infected, and each infected vertex independently infects each of its healthy neighbors at rate $\lambda$ and recovers at rate $1$. We study the contact process in the presence of additional intervention measures by introducing a third possible state for vertices, which we call isolated. Vertices may enter the isolated state either because of individual decisions or due to community-influenced decisions, which leads to two distinct models that we call the isolation model and the vigilance model, respectively. In the isolation model, infected vertices self-isolate at rate $\alpha$. In the vigilance model, each healthy vertex causes each of its infected neighbors to isolate at rate $\alpha$. Unlike the classical contact process, these models lack the key features of attractiveness and existence of a dual, which makes analyzing them more challenging. We study the persistence times of the infection on large, finite, random graphs with heavy-tailed degree distributions. We show that the infection in the isolation model persists for at least stretched exponential time in the size of the graph for all values of $\alpha$ and $\lambda$. By contrast, in the vigilance model, for every fixed $\alpha$ the persistence time of the infection exhibits a phase transition in $\lambda$: for small $\lambda$ the infection persists for at most a linear time in the size of the graph, while for large $\lambda$ the infection persists exponentially long. As a corollary of our main results we show that for any $\lambda, \alpha>0$ the persistence time of the SIRS model on the configuration model having $n$ vertices starting from all vertices infected is at least $\exp(n^{1-\eta})$ with high probability for any $\eta>0$.

math.PR

Limiting distribution of the chemical distance in high dimensional critical percolation

We identify the asymptotic distribution of the chemical distance and other natural metrics (including the resistance) in high-dimensional critical percolation. When rescaled by the square of the Euclidean distance, each of these metrics converges in distribution to a multiple of the hitting time $T$ of a Brownian motion to hit $\mathbf{e}_1$ conditional on $\{T < \infty\}$. We extend these results to the near-critical regime, where the limiting distribution is now the analogue of $T$ for a killed Brownian motion. These results are intended to be the foundation for an understanding of the metric space structure of high-dimensional clusters. They follow from a general theorem we find interesting in its own right, a ``law of large numbers'' for local functions summed along the backbone of a long open connection. A mixing result for open clusters \cite{CCHS} in the form of a robust convergence to the incipient infinite cluster measure plays a key role in the proofs.

math.PR

Robust construction of the incipient infinite cluster in high dimensional critical percolation

We give a new construction of the incipient infinite cluster (IIC) associated with high-dimensional percolation in a broad setting and under minimal assumptions. Our arguments differ substantially from earlier constructions of the IIC; we do not directly use the machinery of the lace expansion or similar diagrammatic expansions. We show that the IIC may be constructed by conditioning on the cluster of a vertex being infinite in the supercritical regime $p > p_c$ and then taking $p \searrow p_c$. Furthermore, at criticality, we show that the IIC may be constructed by conditioning on a connection to an arbitrary distant set $V$, generalizing previous constructions where one conditions on a connection to a single distant vertex or the boundary of a large box. The input to our proof are the asymptotics for the two-point function obtained by Hara, van der Hofstad, and Slade. Our construction thus applies in all dimensions for which those asymptotics are known, rather than an unspecified high dimension considered in previous works. The results in this paper will be instrumental in upcoming work related to structural properties and scaling limits of various objects involving high-dimensional percolation clusters at and near criticality.

math.PR

A dynamic mean-field statistical model of academic collaboration

There is empirical evidence that collaboration in academia has increased significantly during the past few decades, perhaps due to the breathtaking advancements in communication and technology during this period. Multi-author articles have become more frequent than single-author ones. Interdisciplinary collaboration is also on the rise. Although there have been several studies on the dynamical aspects of collaboration networks, systematic statistical models which theoretically explain various empirically observed features of such networks have been lacking. In this work, we propose a dynamic mean-field model and an associated estimation framework for academic collaboration networks. We primarily focus on how the degree of collaboration of a typical author, rather than the local structure of her collaboration network, changes over time. We consider several popular indices of collaboration from the literature and study their dynamics under the proposed model. In particular, we obtain exact formulae for the expectations and temporal rates of change of these indices. Through extensive simulation experiments, we demonstrate that the proposed model has enough flexibility to capture various phenomena characteristic of real-world collaboration networks. Using metadata on papers from the arXiv repository, we empirically study the mean-field collaboration dynamics in disciplines such as Computer Science, Mathematics and Physics.

stat.ME

Concentration inequalities for correlated network-valued processes with applications to community estimation and changepoint analysis

Network-valued time series are currently a common form of network data. However, the study of the aggregate behavior of network sequences generated from network-valued stochastic processes is relatively rare. Most of the existing research focuses on the simple setup where the networks are independent (or conditionally independent) across time, and all edges are updated synchronously at each time step. In this paper, we study the concentration properties of the aggregated adjacency matrix and the corresponding Laplacian matrix associated with network sequences generated from lazy network-valued stochastic processes, where edges update asynchronously, and each edge follows a lazy stochastic process for its updates independent of the other edges. We demonstrate the usefulness of these concentration results in proving consistency of standard estimators in community estimation and changepoint estimation problems. We also conduct a simulation study to demonstrate the effect of the laziness parameter, which controls the extent of temporal correlation, on the accuracy of community and changepoint estimation.

math.ST

Subcritical Connectivity and Some Exact Tail Exponents in High Dimensional Percolation

In high dimensional percolation at parameter $p < p_c$, the one-arm probability $π_p(n)$ is known to decay exponentially on scale $(p_c - p)^{-1/2}$. We show the same statement for the ratio $π_p(n) / π_{p_c}(n)$, establishing a form of a hypothesis of scaling theory. As part of our study, we provide sharp estimates (with matching upper and lower bounds) for several quantities of interest at the critical probability $p_c$. These include the tail behavior of volumes of, and chemical distances within, spanning clusters, along with the scaling of the two-point function at "mesoscopic distance" from the boundary of half-spaces. As a corollary, we obtain the tightness of the number of spanning clusters of a diameter $n$ box on scale $n^{d-6}$; this result complements a lower bound of Aizenman.

math.PR

The effect of avoiding known infected neighbors on the persistence of a recurring infection process

We study a generalization of the classical contact process (SIS epidemic model) in a directed graph $G$. Our model is a continuous-time interacting particle system in which at every time, each vertex is either healthy or infected, and each oriented edge is either active or inactive. Infected vertices become healthy at rate $1$, and pass the infection along each active outgoing edge at rate $λ$. At rate $α$, healthy individuals deactivate each incoming edge from their infected neighbors. We study the persistence time of this epidemic model on the lattice $\mathbb{Z}$, the $n$-cycle $\mathbb{Z}_n$, and the $n$-star graph. We show that on $\mathbb{Z}$, for every $α>0$, there is a phase transition in $λ$ between almost sure extinction and positive probability of indefinite survival; on $\mathbb{Z}_n$ we show that there is a phase transition between poly-logarithmic and exponential survival time as the size of the graph increases. On the star graph, we show that the survival time is $n^{Δ+o(1)}$ for an explicit function $Δ(α,λ)$ whenever $α>0$ and $λ>0$. In the cases of $\mathbb{Z}$ and $\mathbb{Z}_n$, our results qualitatively match what has been shown for the classical contact process, while in the case of the star graph, the classical contact process exhibits exponential survival for all $λ> 0$, which is qualitatively different from our result. This model presents a challenge because, unlike the classical contact process, it has not been shown to be monotonic in the infection parameter $λ$ or the initial infected set.

math.PR

Estimating the treatment effect of the juvenile stay-at-home order on SARS-CoV-2 infection spread in Saline County, Arkansas

We investigate the treatment effect of the juvenile stay-at-home order (JSAHO) adopted in Saline County, Arkansas, from April 6 to May 7, in mitigating the growth of SARS-CoV-2 infection rates. To estimate the counterfactual control outcome for Saline County, we apply Difference-in-Differences and Synthetic Control design methodologies. Both approaches show that stay-at-home order (SAHO) significantly reduced the growth rate of the infections in Saline County during the period the policy was in effect, contrary to some of the findings in the literature that cast doubt on the general causal impact of SAHO with narrower scopes.

stat.AP

Consistent detection and optimal localization of all detectable change points in piecewise stationary arbitrarily sparse network-sequences

We consider the offline change point detection and localization problem in the context of piecewise stationary networks, where the observable is a finite sequence of networks. We develop algorithms involving some suitably modified CUSUM statistics based on adaptively trimmed adjacency matrices of the observed networks for both detection and localization of single or multiple change points present in the input data. We provide rigorous theoretical analysis and finite sample estimates evaluating the performance of the proposed methods when the input (finite sequence of networks) is generated from an inhomogeneous random graph model, where the change points are characterized by the change in the mean adjacency matrix. We show that the proposed algorithms can detect (resp. localize) all change points, where the change in the expected adjacency matrix is above the minimax detectability (resp. localizability) threshold, consistently without any a priori assumption about (a) a lower bound for the sparsity of the underlying networks, (b) an upper bound for the number of change points, and (c) a lower bound for the separation between successive change points, provided either the minimum separation between successive pairs of change points or the average degree of the underlying networks goes to infinity arbitrarily slowly. We also prove that the above condition is necessary to have consistency.

stat.ME

General Community Detection with Optimal Recovery Conditions for Multi-relational Sparse Networks with Dependent Layers

Multilayer and multiplex networks are becoming common network data sets in recent times. We consider the problem of identifying the common community structure for a special type of multilayer networks called multi-relational networks. We consider extensions of the spectral clustering methods for multi-relational networks and give theoretical guarantees that the spectral clustering methods recover community structure consistently for multi-relational networks generated from multilayer versions of both stochastic and degree-corrected block models even with dependence between network layers. The methods are shown to work under optimal conditions on the degree parameter of the networks to detect both assortative and disassortative community structures with vanishing error proportions even if individual layers of the multi-relational network has the network structures below community detectability threshold. We reinforce the validity of the theoretical results via simulations too.

cs.SI

Restricted percolation critical exponents in high dimensions

Despite great progress in the study of critical percolation on $\mathbb{Z}^d$ for $d$ large, properties of critical clusters in high-dimensional fractional spaces and boxes remain poorly understood, unlike the situation in two dimensions. Closely related models such as critical branching random walk give natural conjectures for the value of the relevant high-dimensional critical exponents; see in particular the conjecture by Kozma-Nachmias that the probability that $0$ and $(n, n, n, \ldots)$ are connected within $[-n,n]^d$ scales as $n^{-2-2d}$. In this paper, we study the properties of critical clusters in high-dimensional half-spaces and boxes. In half-spaces, we show that the probability of an open connection ("arm") from $0$ to the boundary of a sidelength $n$ box scales as $n^{-3}$. We also find the scaling of the half-space two-point function (the probability of an open connection between two vertices) and the tail of the cluster size distribution. In boxes, we obtain the scaling of the two-point function between vertices which are any macroscopic distance away from the boundary.

math.PR

Spectral Clustering for Multiple Sparse Networks: I

Although much of the focus of statistical works on networks has been on static networks, multiple networks are currently becoming more common among network data sets. Usually, a number of network data sets, which share some form of connection between each other are known as multiple or multi-layer networks. We consider the problem of identifying the common community structures for multiple networks. We consider extensions of the spectral clustering methods for the multiple sparse networks, and give theoretical guarantee that the spectral clustering methods produce consistent community detection in case of both multiple stochastic block model and multiple degree-corrected block models. The methods are shown to work under sufficiently mild conditions on the number of multiple networks to detect associative community structures, even if all the individual networks are sparse and most of the individual networks are below community detectability threshold. We reinforce the validity of the theoretical results via simulations too.

stat.ME

Thresholds For Detecting An Anomalous Path From Noisy Environments

We consider the "searching for a trail in a maze" composite hypothesis testing problem, in which one attempts to detect an anomalous directed path in a lattice 2D box of side n based on observations on the nodes of the box. Under the signal hypothesis, one observes independent Gaussian variables of unit variance at all nodes, with zero, mean off the anomalous path and mean μ_n on it. Under the null hypothesis, one observes i.i.d. standard Gaussians on all nodes. Arias-Castro et al. (2008) showed that if the unknown directed path under the signal hypothesis has known the initial location, then detection is possible (in the minimax sense) if μ_n >> 1/\sqrt log n, while it is not possible if μ_n << 1/ log n\sqrt log log n. In this paper, we show that this result continues to hold even when the initial location of the unknown path is not known. As is the case with Arias-Castro et al. (2008), the upper bound here also applies when the path is undirected. The improvement is achieved by replacing the linear detection statistic used in Arias-Castro et al. (2008) with a polynomial statistic, which is obtained by employing a multi-scale analysis on a quadratic statistic to bootstrap its performance. Our analysis is motivated by ideas developed in the context of the analysis of random polymers in Lacoin (2010).

math.ST

Jigsaw percolation: What social networks can collaboratively solve a puzzle?

We introduce a new kind of percolation on finite graphs called jigsaw percolation. This model attempts to capture networks of people who innovate by merging ideas and who solve problems by piecing together solutions. Each person in a social network has a unique piece of a jigsaw puzzle. Acquainted people with compatible puzzle pieces merge their puzzle pieces. More generally, groups of people with merged puzzle pieces merge if the groups know one another and have a pair of compatible puzzle pieces. The social network solves the puzzle if it eventually merges all the puzzle pieces. For an Erdős-Rényi social network with $n$ vertices and edge probability $p_n$, we define the critical value $p_c(n)$ for a connected puzzle graph to be the $p_n$ for which the chance of solving the puzzle equals $1/2$. We prove that for the $n$-cycle (ring) puzzle, $p_c(n)=Θ(1/\log n)$, and for an arbitrary connected puzzle graph with bounded maximum degree, $p_c(n)=O(1/\log n)$ and $ω(1/n^b)$ for any $b>0$. Surprisingly, with probability tending to 1 as the network size increases to infinity, social networks with a power-law degree distribution cannot solve any bounded-degree puzzle. This model suggests a mechanism for recent empirical claims that innovation increases with social density, and it might begin to show what social networks stifle creativity and what networks collectively innovate.

math.PR

Multiple phase transitions in long-range first-passage percolation on square lattices

We consider a model of long-range first-passage percolation on the $d$ dimensional square lattice $Z^d$ in which any two distinct vertices $x, y \in Z^d$ are connected by an edge having exponentially distributed passage time with mean $||x-y||^{α+o(1)}$, where $α>0$ is a fixed parameter and $||\cdot||$ is the $\ell_1$-norm on $Z^d$. We analyze the asymptotic growth rate of the set $B_t$, which consists of all $x \in Z^d$ such that the first-passage time between the origin 0 and $x$ is at most $t$, as $t\to\infty$. We show that depending on the values of $α$ there are four growth regimes: (i) instantaneous growth for $α 2d+1$ like the nearest-neighbor first-passage percolation model corresponding to $α=\infty$.

math.PR