SearcharxivSearch

arXiv subjects

Andressa Cerqueira

Publications and source records attributed to Andressa Cerqueira.

10 recordsLinked to original sources

Normalized Likelihood Criteria for Model Selection in the Stochastic Block Model

Estimating the number of communities is a fundamental problem in network analysis under the stochastic block model (SBM). In this paper, we study penalized estimators for this task based on normalized likelihood criteria. We show that a penalized estimator derived from the Normalized Maximum Likelihood (NML) is strongly consistent with a logarithmic penalty term, although its computation is intractable. To overcome this limitation, we consider the Normalized Maximum Complete Likelihood (NMCL) and the Decomposed Normalized Maximum Likelihood (DNML). The DNML admits an explicit formulation with cubic computational complexity in the number of nodes. We prove that the NMCL- and DNML- based estimators are strongly consistent for sparse networks in which the average node degree diverges with the network size. Empirical results show that the DNML estimator performs competitively with existing methods, particularly in unbalanced networks.

math.ST

Optimal recovery by maximum and integrated conditional likelihood in the general Stochastic Block Model

In this paper, we obtain new results on the weak and strong consistency of the maximum and integrated conditional likelihood estimators for the community detection problem in the Stochastic Block Model with $k$ communities and unknown parameters. In particular, we show that maximum conditional likelihood achieves the optimal known threshold for exact recovery in the logarithmic degree regime. For the integrated conditional likelihood, we obtain a sub-optimal constant, but still obtain strong consistency in the logarithmic degree regime. Both methods are shown to be weakly consistent in the divergent degree regime. These results fill in the gap in the theory of community detection with maximum likelihood and integrated conditional likelihood, solving open problems in the literature.

math.ST

Counting communities in weighted Stochastic Block Models via semidefinite programming

We consider the problem of estimating the number of communities in a weighted balanced Stochastic Block Model. We construct hypothesis tests based on semidefinite programming and with a statistic coming from a GOE matrix to distinguish between any two candidate numbers of communities. This is possible due to a universality result for a semidefinite programming-based function that we also prove. The tests are then used to form a sequential test to estimate the number of communities. Furthermore, we also construct estimators of the communities themselves.

math.ST

Modeling sparsity in count-weighted networks

Community detection methods have been extensively studied to recover communities structures in network data. While many models and methods focus on binary data, real-world networks also present the strength of connections, which could be considered in the network analysis. We propose a probabilistic model for generating weighted networks that allows us to control network sparsity and incorporates degree corrections for each node. We propose a community detection method based on the Variational Expectation-Maximization (VEM) algorithm. We show that the proposed method works well in practice for simulated networks. We analyze the Brazilian airport network to compare the community structures before and during the COVID-19 pandemic.

stat.ME

A pseudo-likelihood approach to community detection in weighted networks

Community structure is common in many real networks, with nodes clustered in groups sharing the same connections patterns. While many community detection methods have been developed for networks with binary edges, few of them are applicable to networks with weighted edges, which are common in practice. We propose a pseudo-likelihood community estimation algorithm derived under the weighted stochastic block model for networks with normally distributed edge weights, extending the pseudo-likelihood algorithm for binary networks, which offers some of the best combinations of accuracy and computational efficiency. We prove that the estimates obtained by the proposed method are consistent under the assumption of homogeneous networks, a weighted analogue of the planted partition model, and show that they work well in practice for both homogeneous and heterogeneous networks. We illustrate the method on simulated networks and on a fMRI dataset, where edge weights represent connectivity between brain regions and are expected to be close to normal in distribution by construction.

stat.ME

Consistent model selection for the Degree Corrected Stochastic Blockmodel

The Degree Corrected Stochastic Block Model (DCSBM) was introduced by \cite{karrer2011stochastic} as a generalization of the stochastic block model in which vertices of the same community are allowed to have distinct degree distributions. On the modelling side, this variability makes the DCSBM more suitable for real life complex networks. On the statistical side, it is more challenging due to the large number of parameters when dealing with community detection. In this paper we prove that the penalized marginal likelihood estimator is strongly consistent for the estimation of the number of communities. We consider \emph{dense} or \emph{semi-sparse} random networks, and our estimator is \emph{unbounded}, in the sense that the number of communities $k$ considered can be as big as $n$, the number of nodes in the network.

math.ST

Graphical Construction of Spatial Gibbs Random Graphs

We consider a Random Graph Model on $\mathbb{Z}^{d}$ that incorporates the interplay between the statistics of the graph and the underlying space where the vertices are located. Based on a graphical construction of the model as the invariant measure of a birth and death process, we prove the existence and uniqueness of a measure defined on graphs with vertices in $\mathbb{Z}^{d}$ which coincides with the limit along the measures over graphs with finite vertex set. As a consequence, theoretical properties such as exponential mixing of the infinite volume measure and central limit theorem for averages of a real-valued function of the graph are obtained. Moreover, a perfect simulation algorithm based on the clan of ancestors is described in order to sample a finite window of the equilibrium measure defined on $\mathbb{Z}^{d}$.

math.ST

Strong consistency of Krichevsky-Trofimov estimator for the number of communities in the Stochastic Block Model

In this paper we introduce an estimator for the number of communities in the Stochastic Block Model (SBM), based on the maximization of a penalized version of the so-called Krichevsky-Trofimov mixture distribution. We prove its eventual almost sure convergence to the underlying number of communities, without assuming a known upper bound on that quantity. Our results apply to both the dense and the sparse regimes. To our knowledge this is the first consistency result for the estimation of the number of communities in the SBM in the unbounded case, that is when the number of communities is allowed to grow with the same size.

math.ST

A note on perfect simulation for exponential random graph models

In this paper we propose a perfect simulation algorithm for the Exponential Random Graph Model, based on the Coupling From The Past method of Propp & Wilson (1996). We use a Glauber dynamics to construct the Markov Chain and we prove the monotonicity of the ERGM for a subset of the parametric space. We also obtain an upper bound on the running time of the algorithm that depends on the mixing time of the Markov chain.

stat.CO

A test of hypotheses for random graph distributions built from EEG data

The theory of random graphs is being applied in recent years to model neural interactions in the brain. While the probabilistic properties of random graphs has been extensively studied in the literature, the development of statistical inference methods for this class of objects has received less attention. In this work we propose a non-parametric test of hypotheses to test if two samples of random graphs were originated from the same probability distribution. We show how to compute efficiently the test statistic and we study its performance on simulated data. We apply the test to compare graphs of brain functional network interactions built from electroencephalographic (EEG) data collected during the visualization of point light displays depicting human locomotion.

stat.AP