SearcharxivSearch

arXiv subjects

Shankar Bhamidi

Publications and source records attributed to Shankar Bhamidi.

At least 19 recordsLinked to original sources

Subcritical percolation and network archaeology on random recursive tree substrate networks

We study a network-archaeology problem for a dynamic graph whose latent substrate is a random recursive tree and whose observed topology is enriched by an independent homogeneous Erd\H{o}s-R\'enyi shortcut layer. From a single unlabeled snapshot, the goal is to construct a confidence set of deterministic size for the first vertex. Since shortcut edges create cycles, the usual tree-based arguments using Jordan centrality do not apply directly. Our method uses auxiliary subcritical bond percolation to expose a tree-like renormalized structure: retained recursive-tree clusters form heavy-tailed blobs, retained shortcuts connect these blobs through a subcritical rank-one random graph, and large components are leading backbone blobs decorated by subcritical shortcut pieces. Applying Jordan centrality inside the largest auxiliary percolation components then gives a deterministic-size root confidence set for the cyclic observed network.

math.PR

Network evolution with self-reinforcement

We study a new class of preferential attachment trees with \emph{self-reinforcement}. At each time, each vertex is assigned a weight equal to the cumulative sum over past times of an affine function of its degree. A new vertex attaches itself via a single edge to an already present vertex with a probability proportional to the current weight of that vertex. This ``integrated popularity'' rule builds long memory directly into the attachment mechanism, thereby destroying the Markov and partial-exchangeability features that underlie the classical analysis of affine preferential attachment models. More broadly, the model connects to applied-probability work on long-memory self-interacting processes (such as the elephant random walk), emphasizing how non-Markovian reinforcement reshapes asymptotic behaviour. Despite this loss of structure, we identify an explicit exponent $\phi=\phi(\delta)$ governing both local and global growth: typical degrees at time $n$ scale as $n^{1/\phi}$, and the empirical degree distribution converges to a power-law with a tail exponent $\phi+1$. We further prove Benjamini--Schramm local convergence to an infinite random rooted tree characterized via an embedded continuous-time branching process. The limiting tree is a \texttt{sin}-tree, and is \emph{not} the P\'olya-type limiting tree arising in the non-reinforced setting. Our results provide a tractable probabilistic description of a natural ``memoryful'' network-growth mechanism, and quantify precisely how reinforcement renormalizes the classical preferential-attachment exponents.

math.PR

The stochastic block model has the overlap graph property for modularity

The overlap gap property (OGP) is a statement about the geometry of near-optimal solutions. Exhibiting OGP implies failure of a class of local algorithms; and has been observed to coincide with conjectured algorithmic limits in problems with statistical computational gap. We consider the Stochastic Block Model (SBM), where the graph has a planted partition with $k$ equal-size blocks which form the `communities', and where, for parameters $p>q$, vertices within the same community connect with probability $p$, while vertices in different communities connect with probability $q$, independently across pairs of vertices. Modularity--based clustering algorithms have become ubiquitous in applications. This article studies theoretical limits of local algorithms based on the modularity score on the SBM. We establish that modularity exhibits OGP on the SBM. This rules out a class of local algorithms based on modularity for recovery in the SBM, and shows slow mixing time for a related Markov Chain. Theoretically this is one of the few instances where OGP has been established for a `planted' model, as most such analyses to date consider the `null' model. As part of our analysis, we extend a result by Bickel and Chen 2009, who established that with high probability, the modularity optimal partition of SBM is $o(n)$ local moves away from the planted partition, where $n$ is the graph size. We show that, with high probability, any partition with modularity score sufficiently near the optimal value is close to the planted partition.

math.PR

Non-equilibrium coagulation processes and subcritical percolation on evolving networks

We investigate percolation on growing networks where the evolution of connected components resembles a non-equilibrium version of the multiplicative coalescent. The supercritical $\pi> \pi_c$ regime for a host of such models was conjectured in statistical physics, and then rigorously proven in mathematics, to exhibit behavior similar to the BKT infinite-order phase transition as $\pi\searrow \pi_c$. It has further been conjectured that the entire regime $\pi<\pi_c$ for such growing networks are ''critical'' with power-law cluster size distributions having a non-universal exponent for all values of $\pi \in (0, \pi_c)$. In this paper, we study percolation on the uniform attachment model, as a concrete template in order to develop general tools based on stochastic approximation, local convergence, branching random walks and tree-graph inequalities to prove the above conjectured phenomena. For each $\pi \in (0,\pi_c)$, we show there exists an explicit $\alpha(\pi) \in (0,\tfrac{1}{2}) $ such that the maximal component size, as well as the size of the component containing any fixed vertex, all re-scaled by $n^{\alpha(\pi)}$, converge almost surely to strictly positive random variables as the network size $n \to \infty$. These dynamics lead to novel phenomena, compared to classical 'static' models, including long-range dependence and fixation of the identity of the maximal component, within finite time, among a finite number of 'early' components. Moreover, in contrast with most static network models, we show that the susceptibility, that is, the expected size of the component of a uniformly chosen vertex, remains bounded as the network grows and $\pi$ approaches $\pi_c$ from below. The general tools developed in this paper will be used in follow-up work to understand percolation for general growing network evolution models.

math.PR

Evolution of recursive trees with limited memory

Motivated by questions in social networks, distributed computing and probabilistic combinatorics, the last few years have seen increasing interest in network evolution models where new vertices entering the system need to make decisions based on a partial snapshot of the current state of the network. This paper considers a specific variant of the classical random recursive tree dynamics, where a vertex at time $n+1$ has information only on those vertices that have arrived in the interval $[j(n), n]$ for a sequence $j(n) \uparrow \infty$, and connects to vertices uniformly at random amongst this set. We consider two different regimes on the density information, termed macroscopic and mesoscopic regimes, which respectively correspond to $j(n)=\theta n$ for some $\theta \in (0,1)$, and $j(n)=n-n^{\beta}$ for some $\beta \in (0,1)$. Our main interest is in studying asymptotics of various local and global functionals of the network. We show that in the macroscopic regime, the local limit is expressed in terms of an associated continuous time branching process that depends on the parameter $\theta$, while it is a $\mathrm{Poisson}(1)$-branching process in the mesoscopic regime for any $\beta \in (0,1)$. Furthermore, the height of the macroscopic tree is logarithmic, which we prove exploiting a connection with scaled-attachment random recursive trees (SARRTs) as studied by Devroye, Fawzi and Fraiman (RSA 2011), while it is polynomial in the mesoscopic regime; our argument in this latter case relies on a differential equation approach to track the ancestor indices of late-coming vertices, together with a multiscale analysis. Further, we develop an exploration algorithm to simultaneously reveal the ancestral path of youngest vertices. Using this algorithm, we show that in the mesoscopic regime, the global structure experiences a phase transition at $\beta=1/2$.

math.PR

Finding a dense submatrix of a random matrix. Sharp bounds for online algorithms

We consider the problem of finding a dense submatrix of a matrix with i.i.d. Gaussian entries, where density is measured by average value. This problem arose from practical applications in biology and social sciences \cites{madeira-survey,shabalin2009finding} and is known to exhibit a computation-to-optimization gap between the optimal value and best values achievable by existing polynomial time algorithms. In this paper we consider the class of online algorithms, which includes the best known algorithm for this problem, and derive a tight approximation factor ${4\over 3\sqrt{2}}$ for this class. The result is established using a simple implementation of recently developed Branching-Overlap-Gap-Property \cite{huang2025tight}. We further extend our results to $(\mathbb R^n)^{\otimes p}$ tensors with i.i.d. Gaussian entries, for which the approximation factor is proven to be ${2\sqrt{p}/(1+p)}$.

math.PR

Large Deviations for Markovian Graphon Processes and Associated Dynamical Systems on Networks

We consider temporal models of rapidly evolving Markovian networks whose edge-formation and dissolution rates are determined by time-dependent spatial kernels. Equivalently, these may be viewed as Markovian networks with $O(1)$ jump rates observed over long time horizons. In this regime, paths of graphon-valued processes obtained by averaging over suitable moving time windows provide natural state descriptors. Under appropriate conditions on the jump-rate kernels, we establish laws of large numbers and large deviation principles for these window-averaged paths, both in the weak topology and in the cut metric. We also show that, without such local averaging, the rapidly oscillating graphon process does not satisfy a nontrivial path-space LDP. The resulting rate functions admit explicit and tractable representations, distinct from those arising in static random graph models and finite-horizon dynamic graph models. We further analyze the associated variational problems in several examples and apply the graphon LDP to node-valent dynamical systems driven by the evolving network.

math.PR

Functional Central limit theorems for microscopic and macroscopic functionals of inhomogeneous random graphs

We study inhomogeneous random graphs with a finite type space. For a natural generalization of the model as a dynamic network-valued process, the paper establishes the following results: (a) Functional central limit theorems for the infinite vector of microscopic type-densities and characterizations of the limits as infinite-dimensional conditionally Gaussian processes in a certain Banach space. (b) Functional (joint) central limit theorems for macroscopic observables of the giant component in the supercritical regime including size, surplus and number of vertices of various types in the giant component. As a corollary this provides central limit theorems for the size of the largest connected component, its surplus, and its type vector, for percolation on dense graphs obtained from a finite type Graphon. (c) Central limit theorem for the weight of the minimum spanning tree with random i.i.d. Exponential edge weights on dense graph sequences driven by an underlying finite type graphon.

math.PR

Critical first passage percolation on random graphs

In 1999, Zhang proved that, for first passage percolation on the square lattice $\mathbb{Z}^2$ with i.i.d. non-negative edge weights, if the probability that the passage time distribution of an edge $P(t_e = 0) =1/2 $, the critical value for bond percolation on $\mathbb{Z}^2$, then the passage time from the origin $0$ to the boundary of $[-n,n]^2$ may converge to $\infty$ or stay bounded depending on the nature of the distribution of $t_e$ close to zero. In 2017, Damron, Lam, and Wang gave an easily checkable necessary and sufficient condition for the passage time to remain bounded. Concurrently, there has been tremendous growth in the study of weak and strong disorder on random graph models. Standard first passage percolation with strictly positive edge weights provides insight in the weak disorder regime. Critical percolation on such graphs provides information on the strong disorder (namely the minimal spanning tree) regime. Here we consider the analogous problem of Zhang but now for a sequence of random graphs $\{G_n:n\geq 1\}$ generated by a supercritical configuration model with a fixed degree distribution. Let $p_c$ denote the associated critical percolation parameter, and suppose each edge $e\in E(G_n)$ has weight $t_e \sim p_c \delta_0 +(1-p_c)\delta_{F_\zeta}$ where $F_\zeta$ is the cdf of a random variable $\zeta$ supported on $(0,\infty)$. The main question of interest is: when does the passage time between two randomly chosen vertices have a limit in distribution in the large network $n\to \infty$ limit? There are interesting similarities between the answers on $\mathbb{Z}^2$ and on random graphs, but it is easier for the passage times on random graphs to stay bounded.

math.PR

Network evolution with mesoscopic delay

Owing to the influence of real-world networks both in science and society, numerous mathematical models have been developed to understand the structure and evolution of these systems, particularly in a temporal context. Recent advancements in fields like distributed cyber-security and social networks have spurred the creation of probabilistic models of evolution, where individuals make decisions based on only partial information about the network's current state. This paper seeks to explore models incorporating network delay, where new participants receive information from a time-lagged snapshot of the system. In the context of mesoscopic network delays, we develop probabilistic tools built on stochastic approximation to understand asymptotics of both local functionals, such as local neighborhoods and degree distributions, as well as global properties, such as the evolution of the degree of the network's initial founder. A companion paper explores the regime of macroscopic delays in the evolution of the network.

math.PR

Network evolution with Macroscopic Delays: asymptotics and condensation

Preferential attachment models typically assume that each arriving vertex observes the current network before choosing its connection. Motivated by distributed systems and social networks, we study network delay, where this decision uses only a time-delayed snapshot. We focus on macroscopic delays, for which the delay is proportional to the current network size and hence removes a non-vanishing fraction of the available information. We identify the local weak limit as a continuous-time branching process whose reproduction point process has memory of its entire past. Since this non-Markovian description is difficult to analyze directly, we construct a dual branching process in which edges reproduce, recovering enough independence for quantitative analysis. This yields a detailed understanding of how the delay affects features such as the tail behavior of the asymptotic degree distribution, together with necessary and sufficient conditions for condensation-the phenomenon in which a positive fraction of the degree mass escapes to infinity. We conclude by studying the impact of the delay distribution on macroscopic functionals such as the root degree.

math.PR

Local weak convergence and its applications

Motivated in part by understanding average case analysis of fundamental algorithms in computer science, and in part by the wide array of network data available over the last decade, a variety of random graph models, with corresponding processes on these objects, have been proposed over the last few years. The main goal of this paper is to give an overview of local weak convergence, which has emerged as a major technique for understanding large network asymptotics for a wide array of functionals and models. As opposed to a survey, the main goal is to try to explain some of the major concepts and their use to junior researchers in the field and indicate potential resources for further reading.

math.PR

Consistency of Lloyd's Algorithm Under Perturbations

In the context of unsupervised learning, Lloyd's algorithm is one of the most widely used clustering algorithms. It has inspired a plethora of work investigating the correctness of the algorithm under various settings with ground truth clusters. In particular, in 2016, Lu and Zhou have shown that the mis-clustering rate of Lloyd's algorithm on $n$ independent samples from a sub-Gaussian mixture is exponentially bounded after $O(\log(n))$ iterations, assuming proper initialization of the algorithm. However, in many applications, the true samples are unobserved and need to be learned from the data via pre-processing pipelines such as spectral methods on appropriate data matrices. We show that the mis-clustering rate of Lloyd's algorithm on perturbed samples from a sub-Gaussian mixture is also exponentially bounded after $O(\log(n))$ iterations under the assumptions of proper initialization and that the perturbation is small relative to the sub-Gaussian noise. In canonical settings with ground truth clusters, we derive bounds for algorithms such as $k$-means$++$ to find good initializations and thus leading to the correctness of clustering via the main result. We show the implications of the results for pipelines measuring the statistical significance of derived clusters from data such as SigClust. We use these general results to derive implications in providing theoretical guarantees on the misclustering rate for Lloyd's algorithm in a host of applications, including high-dimensional time series, multi-dimensional scaling, and community detection for sparse networks via spectral clustering.

cs.LG

Correlation networks, dynamic factor models and community detection

A dynamic factor model with a mixture distribution of the loadings is introduced and studied for multivariate, possibly high-dimensional time series. The correlation matrix of the model exhibits a block structure, reminiscent of correlation patterns for many real multivariate time series. A standard $k$-means algorithm on the loadings estimated through principal components is used to cluster component time series into communities with accompanying bounds on the misclustering rate. This is one standard method of community detection applied to correlation matrices viewed as weighted networks. This work puts a mixture model, a dynamic factor model and network community detection in one interconnected framework. Performance of the proposed methodology is illustrated on simulated and real data.

stat.ME

Dynamic factor and VARMA models: equivalent representations, dimension reduction and nonlinear matrix equations

A dynamic factor model with factor series following a VAR$(p)$ model is shown to have a VARMA$(p,p)$ model representation. Reduced-rank structures are identified for the VAR and VMA components of the resulting VARMA model. It is also shown how the VMA component parameters can be computed numerically from the original model parameters via the innovations algorithm, and connections of this approach to non-linear matrix equations are made. Some VAR models related to the resulting VARMA model are also discussed.

stat.ME

Attribute network models, stochastic approximation, and network sampling and ranking algorithms

We analyze dynamic random network models where younger vertices connect to older ones with probabilities proportional to their degrees as well as a propensity kernel governed by their attribute types. Using stochastic approximation techniques we show that, in the large network limit, such networks converge in the local weak sense to limiting infinite random trees with an explicit description in terms of randomly stopped multi-type branching processes. This allows for the derivation of asymptotics for a wide class of network functionals implying, for example, that while degree distribution tail exponents depend on the attribute type (already derived by Jordan (2013)), PageRank centrality scores have the same tail exponent across attributes. The limit results also give explicit formulae for the performance of various network sampling mechanisms. One surprising consequence is the efficacy of PageRank and walk based network sampling schemes for directed networks in the setting of rare minorities.

math.PR

Scaling limits and universality: Critical percolation on weighted graphs converging to an $L^3$ graphon

We develop a general universality technique for establishing metric scaling limits of critical random discrete structures exhibiting mean-field behavior that requires four ingredients: (i) from the barely subcritical regime to the critical window, components merge approximately like the multiplicative coalescent, (ii) asymptotics of the susceptibility functions are the same as that of the Erdos-Renyi random graph, (iii) asymptotic negligibility of the maximal component size and the diameter in the barely subcritical regime, and (iv) macroscopic averaging of distances between vertices in the barely subcritical regime. As an application of the general universality theorem, we establish, under some regularity conditions, the critical percolation scaling limit of graphs that converge, in a suitable topology, to an $L^3$ graphon. In particular, we define a notion of the critical window in this setting. The $L^3$ assumption ensures that the model is in the Erdos-Renyi universality class and that the scaling limit is Brownian. Our results do not assume any specific functional form for the graphon. As a consequence of our results on graphons, we obtain the metric scaling limit for Aldous-Pittel's RGIV model [9] inside the critical window. Our universality principle has applications in a number of other problems including in the study of noise sensitivity of critical random graphs [52]. In [10], we use our universality theorem to establish the metric scaling limit of critical bounded size rules. Our method should yield the critical metric scaling limit of Rucinski and Wormald's random graph process with degree restrictions [56] provided an additional technical condition about the barely subcritical behavior of this model can be proved.

math.PR

Co-evolving dynamic networks

We propose a general class of co-evolving tree network models driven by local exploration where new vertices attach to the current network via randomly sampling a vertex and then exploring the graph for a random number of steps in the direction of the root, connecting to the terminal vertex. Specific choices of the exploration step distribution lead to the well-studied affine preferential attachment and uniform attachment models, as well as less well understood dynamic network models with global attachment functionals such as PageRank scores [Chebolu-Melsted (2008)]. We obtain local weak limits for such networks and use them to derive asymptotics for the limiting empirical degree and PageRank distribution. We also quantify asymptotics for the degree and PageRank of fixed vertices, including the root, and the height of the network. Two distinct regimes are seen to emerge, based on the expected exploration distance of incoming vertices, which we call the `fringe' and `non-fringe' regimes. These regimes are shown to exhibit different qualitative and quantitative properties. In particular, networks in the non-fringe regime undergo `condensation' where the root degree grows at the same rate as the network size. Networks in the fringe regime do not exhibit condensation. Non-trivial phase transition phenomena are displayed for the height and the PageRank distribution, the latter connecting to the well known power-law hypothesis. In the process, we develop a general set of techniques involving local limits, infinite-dimensional urn models, related multitype branching processes and corresponding Perron-Frobenius theory, branching random walks, and in particular relating tail exponents of various functionals to the scaling exponents of quasi-stationary distributions of associated random walks. These techniques are expected to shed light on a variety of other co-evolving network models.

math.PR