SearcharxivSearch

arXiv subjects

J. van Waaij

Publications and source records attributed to J. van Waaij.

4 recordsLinked to original sources

Asymptotic uncertainty quantification for communities in sparse planted bi-section models

Posterior distributions for community structure in sparse planted bi-section models are shown to achieve exact (resp. almost-exact) recovery, with sharp bounds for the sparsity regimes where edge probabilities decrease as $O(\log(n)/n)$ (resp. $O(1/n)$). Assuming posterior recovery, one may interpret credible sets (resp. enlarged credible sets) as asymptotically consistent confidence sets; the diameters of those credible sets are controlled by the rate of posterior concentration. If credible levels are chosen to grow to one quickly enough, corresponding credible sets can be interpreted as frequentist confidence sets without conditions on posterior concentration. In the regimes with $O(1/n)$ edge sparsity, or when within-community and between-community edge probabilities are very close, credible sets may be enlarged to achieve frequentist asymptotic coverage, also without conditions on posterior concentration.

math.ST

Confidence sets in a sparse stochastic block model with two communities of unknown sizes

In a sparse stochastic block model with two communities of unequal sizes we derive two posterior concentration inequalities, that imply (1) posterior (almost-)exact recovery of the community structure under sparsity bounds comparable to well-known sharp bounds in the planted bi-section model; (2) a construction of confidence sets for the community assignment from credible sets, with finite graph sizes. The latter enables exact frequentist uncertain quantification with Bayesian credible sets at non-asymptotic graph sizes, where posteriors can be simulated well. There turns out to be no proportionality between credible and confidence levels: for given edge probabilities and a desired confidence level, there exists a critical graph size where the required credible level drops sharply from close to one to close to zero. At such graph sizes the frequentist decides to include not most of the posterior support for the construction of his confidence set, but only a small subset of community assignments containing the highest amounts of posterior probability (like the maximum-a-posteriori estimator). It is argued that for the proposed construction of confidence sets, a form of early stopping applies to MCMC sampling of the posterior, which would enable the computation of confidence sets at larger graph sizes.

math.ST

Uncertainty quantification and testing in a stochastic block model with two unequal communities

We show posterior convergence for the community structure in the planted bi-section model, for several interesting priors. Examples include where the label on each vertex is iid Bernoulli distributed, with some parameter $r\in(0,1)$. The parameter $r$ may be fixed, or equipped with a beta distribution. We do not have constraints on the class sizes, which might be as small as zero, or include all vertices, and everything in between. This enables us to test between a uniform (Erdös-Rényi) random graph with no distinguishable community or the planted bi-section model. The exact bounds for posterior convergence enable us to convert credible sets into confidence sets. Symmetric testing with posterior odds is shown to be consistent.

math.ST

Uncertainty quantification in the stochastic block model with an unknown number of classes

We study the frequentist properties of Bayesian statistical inference for the stochastic block model, with an unknown number of classes of varying sizes. We equip the space of vertex labellings with a prior on the number of classes and, conditionally, a prior on the labels. The number of classes may grow to infinity as a function of the number of vertices, depending on the sparsity of the graph. We derive non-asymptotic posterior contraction rates of the form $P_{θ_{0,n}}Π_n(B_n\mid X^n)\le ε_n$, where $X^n$ is the observed graph, generated according to $P_{θ_{0,n}}$, $B_n$ is either $\{θ_{0, n}\}$ or, in the very sparse case, a ball around $θ_{0,n}$ of known extent, and $ε_n$ is an explicit rate of convergence. These results enable conversion of credible sets to confidence sets. In the sparse case, credible tests are shown to be confidence sets. In the very sparse case, credible sets are enlarged to form confidence sets. Confidence levels are explicit, for each $n$, as a function of the credible level and the rate of convergence. Hypothesis testing between the number of classes is considered with the help of posterior odds, and is shown to be consistent. Explicit upper bounds on errors of the first and second type and an explicit lower bound on the power of the tests are given.

math.ST