Searcharxiv⌕ Search

arXiv subjects

Remco van der Hofstad

Publications and source records attributed to Remco van der Hofstad.

At least 19 recordsLinked to original sources

An Information-Theoretic Analysis of Threshold Group Testing

We study the Threshold Group Testing (TGT) problem in the noiseless and non-adaptive setting, where the objective is to exactly recover a sparse binary vector from pooled tests, using as few tests as possible. In TGT, each test applied to a subset of items returns a positive outcome if the number of 1's (defective items) in that subset meets or exceeds a specified threshold, and has a negative outcome otherwise. We investigate how the complexity of TGT compares to that of Classical Group Testing (CGT), corresponding to the special case of the threshold equal to one, and analyse the impact of increasing the threshold on the required number of tests. Our main contribution is the derivation of a sharp information-theoretic phase transition at $c_{\mathrm{inf}}^{\mathrm{TGT}}k\log(n/k)$ (non-adaptive) tests for TGT within the constant-column test design. The threshold constant $c_{\mathrm{inf}}^{\mathrm{TGT}}$ is expressed as a function of the prevalence of defectives and the threshold value. Our upper bound is derived under an analytic assumption, and we verify that this assumption is satisfied for a threshold value of 2. The value of $c_{\mathrm{inf}}^{\mathrm{TGT}}$ reveals that TGT on the constant-column design has the same information-theoretic behaviour as CGT in the low-prevalence regime. Yet, strikingly, at higher prevalences, the threshold leads to a significant reduction in the number of tests. On the other hand, we provide evidence that when the asymptotic proportion of defective items is positive, TGT actually becomes strictly harder than CGT (excluding trivial reductions).

cs.IT↗

Algorithms for Threshold Group Testing

We study the Threshold Group Testing (TGT) problem without a gap in the noiseless, non-adaptive setting, where the goal is to exactly recover a sparse binary vector from pooled test outcomes using as few tests as possible. In TGT, a test applied to a subset of items returns a positive outcome if the number of defective items in the subset reaches a prescribed threshold, and a negative outcome otherwise. Under the assumption of an analytic condition, TGT has been shown to undergo a sharp information-theoretic phase transition for exact recovery on the class of constant-column test designs. In this paper, we develop an efficient inference algorithm that achieves exact recovery with high probability using the minimum number of non-adaptive tests that are needed for the constant-column design, thereby matching the information-theoretic threshold of a natural benchmark test design. Our approach is based on a spatially coupled test design and admits a significantly simpler analysis than existing algorithms for related group testing problems. In particular, unlike previous methods for binary group testing, our algorithm does not rely on the analysis of intricate weighted sums. This leads to a more straightforward proof technique, while still allowing near-optimal performance guarantees.

cs.IT↗

Logarithmic typical distances in preferential attachment models

We prove that the typical distances in a preferential attachment model with out-degree $m\geq 2$ and strictly positive fitness parameter are close to $\log_ν{n}$, where $ν$ is the exponential growth parameter of the local limit of the preferential attachment model. The proof relies on a path-counting technique, the first- and second-moment methods, as well as a novel proof of the convergence of the spectral radius of the offspring operator under a certain truncation.

math.PR↗

Likelihood-based anomaly detection in preferential attachment networks

Preferential attachment (PA) network is a widely used model for capturing the growth dynamics of real-world networks, in which newly arriving vertices are more likely to connect to existing vertices with higher degrees. In this paper, we consider a setting in which an anomalous vertex appears at some time point and receives edges with an additional attachment advantage governed by a parameter $β$, while ordinary vertices continue to follow the PA mechanism with parameter $δ$. Detecting such anomalies is challenging due to the high variability in degree growth in PA networks and the limited information available when the anomaly arrives late in the network evolution. We propose an iterative parameter estimation procedure together with a likelihood-based detection framework. Simulation results show that the proposed procedure provides accurate estimation of the parameters $β$ and $δ$. Detection performance depends on the time of anomaly occurrence: anomalies arising at midway stages of the network evolution are detected most reliably, whereas very early and late anomalies remain challenging, particularly when the parameter $β$ is small.

cs.SI↗

Power-law hypothesis and (un)fairness of PageRank on undirected multi-type PAMs

The preferential attachment model (PAM) describes the sequential growth of a network based on the "rich-get-richer" principle. Several versions of it have become established for modeling, e.g., citation networks, capturing a power-law degree distribution. Directed versions of the preferential attachment model where the edges are directed from the new to the old vertices have been the subject of extensive research. They have been shown to exhibit remarkable properties such as heavier tails for the limiting graph-normalized PageRank than for the in-degrees. By contrast, for the undirected version, we recently showed that PageRank has similar tails as the degree. In the present paper, we discuss the PageRank asymptotics for a multi-type version of the undirected PAM (here vertices have different colors), complementing previous results of Antunes, Bhamidi, Banerjee and Pipiras on the asymptotics of PageRank on similar directed multi-type or colored PAMs. Our studies are motivated by the aim to go beyond the rigid rule of edge orientation in directed preferential attachment models. As the main result, for the case of a finite set of colors, we show that the power-law hypothesis for PageRank is fulfilled also for the colored undirected PAM, where, by contrast to the directed case, the power-law exponent is color-dependent for some choices of the initial color distribution and the attractiveness function. For the specific case of a two-type model, we discuss implications of our results on fairness in sampling underrepresented nodes from the network.

math.PR↗

Network evolution with self-reinforcement

We study a new class of preferential attachment trees with \emph{self-reinforcement}. At each time, each vertex is assigned a weight equal to the cumulative sum over past times of an affine function of its degree. A new vertex attaches itself via a single edge to an already present vertex with a probability proportional to the current weight of that vertex. This ``integrated popularity'' rule builds long memory directly into the attachment mechanism, thereby destroying the Markov and partial-exchangeability features that underlie the classical analysis of affine preferential attachment models. More broadly, the model connects to applied-probability work on long-memory self-interacting processes (such as the elephant random walk), emphasizing how non-Markovian reinforcement reshapes asymptotic behaviour. Despite this loss of structure, we identify an explicit exponent $ϕ=ϕ(δ)$ governing both local and global growth: typical degrees at time $n$ scale as $n^{1/ϕ}$, and the empirical degree distribution converges to a power-law with a tail exponent $ϕ+1$. We further prove Benjamini--Schramm local convergence to an infinite random rooted tree characterized via an embedded continuous-time branching process. The limiting tree is a \texttt{sin}-tree, and is \emph{not} the Pólya-type limiting tree arising in the non-reinforced setting. Our results provide a tractable probabilistic description of a natural ``memoryful'' network-growth mechanism, and quantify precisely how reinforcement renormalizes the classical preferential-attachment exponents.

math.PR↗

The stochastic block model has the overlap graph property for modularity

The overlap gap property (OGP) is a statement about the geometry of near-optimal solutions. Exhibiting OGP implies failure of a class of local algorithms; and has been observed to coincide with conjectured algorithmic limits in problems with statistical computational gap. We consider the Stochastic Block Model (SBM), where the graph has a planted partition with $k$ equal-size blocks which form the `communities', and where, for parameters $p>q$, vertices within the same community connect with probability $p$, while vertices in different communities connect with probability $q$, independently across pairs of vertices. Modularity--based clustering algorithms have become ubiquitous in applications. This article studies theoretical limits of local algorithms based on the modularity score on the SBM. We establish that modularity exhibits OGP on the SBM. This rules out a class of local algorithms based on modularity for recovery in the SBM, and shows slow mixing time for a related Markov Chain. Theoretically this is one of the few instances where OGP has been established for a `planted' model, as most such analyses to date consider the `null' model. As part of our analysis, we extend a result by Bickel and Chen 2009, who established that with high probability, the modularity optimal partition of SBM is $o(n)$ local moves away from the planted partition, where $n$ is the graph size. We show that, with high probability, any partition with modularity score sufficiently near the optimal value is close to the planted partition.

math.PR↗

The number and structure of connected graphs with a fixed degree sequence

We study connected graphs with a fixed degree sequence, in the sparse setting where the number of edges grows linearly in the number of vertices. Using the relation to the configuration model, we identify the number of such connected graphs up to the exponential order. We do this by viewing a connected graph with a given degree distribution as the realization of the giant component in a larger configuration model, and carefully choosing the degree distribution of the larger graph so that it is likely that its giant component has the required degree distribution. To ensure that the connected graph has exactly the correct degrees, we use a switching argument. Additionally, we obtain results on rare event probabilities and describe the local structure of a uniform connected graph with a fixed degree sequence.

math.CO↗

Clustering without geometry in sparse networks with independent edges

The coexistence of sparsity and clustering (non-vanishing average fraction of triangles per node) is one of the few structural features that, irrespective of finer details, are ubiquitously observed across large real-world networks. This fact calls for generic models producing sparse clustered graphs. Earlier results suggested that sparse random graphs with independent edges fail to reproduce clustering, unless edge probabilities are assumed to depend on underlying metric distances that, thanks to the triangle inequality, naturally favour triadic closure. This observation has opened a debate on whether clustering implies (latent) geometry in real-world networks. Alternatively, recent models of higher-order networks can replicate clustering by abandoning edge independence. In this paper, we mathematically prove, and numerically confirm, that a sparse random graph with independent edges, recently identified in the context of network renormalization as an invariant model under node aggregation, produces finite clustering without any geometric or higher-order constraint. The underlying mechanism is an infinite-mean node fitness, which also implies a power-law degree distribution. Further, as a novel phenomenon that we characterize rigorously, we observe the breakdown of self-averaging of various network properties. Therefore, as an alternative to geometry or higher-order dependencies, node aggregation invariance emerges as a basic route to realistic network properties.

math.PR↗

Universality of the local limit of preferential attachment models

We study preferential attachment models where vertices enter the network with i.i.d. random numbers of edges that we call the out-degree. We identify the local limit of such models, substantially extending the work of Berger et al.(2014). The degree distribution of this limiting random graph, which we call the random Pólya point tree, has a surprising size-biasing phenomenon. Many of the existing preferential attachment models can be viewed as special cases of our preferential attachment model with i.i.d. out-degrees. Additionally, our models incorporate negative values of the preferential attachment fitness parameter, which allows us to consider preferential attachment models with infinite-variance degrees. Our proof of local convergence consists of two main steps: a Pólya urn description of our graphs, and an explicit identification of the neighbourhoods in them. We provide a novel and explicit proof to establish a coupling between the preferential attachment model and the Pólya urn graph. Our result proves a density convergence result, for fixed ages of vertices in the local limit.

math.PR↗

Sparse random graphs with many triangles

In this paper we consider the Erdős-Rényi random graph in the sparse regime in the limit as the number of vertices $n$ tends to infinity. We are interested in what this graph looks like when it contains many triangles, in two settings. First, we derive asymptotically sharp bounds on the probability that the graph contains a large number of triangles. We show that conditionally on this event, with high probability the graph contains an almost complete subgraph, i.e., the triangles form a near-clique, and has the same local limit as the original Erdős-Rényi random graph. Second, we derive asymptotically sharp bounds on the probability that the graph contains a large number of vertices that are part of a triangle. If order $n$ vertices are in triangles, then the local limit (provided it exists) is different from that of the Erdős-Rényi random graph. Our results shed light on the challenges that arise in the description of real-world networks, which often are sparse, yet highly clustered, and on exponential random graphs, which often are used to model such networks.

math.PR↗

Non-equilibrium coagulation processes and subcritical percolation on evolving networks

We investigate percolation on growing networks where the evolution of connected components resembles a non-equilibrium version of the multiplicative coalescent. The supercritical $π> π_c$ regime for a host of such models was conjectured in statistical physics, and then rigorously proven in mathematics, to exhibit behavior similar to the BKT infinite-order phase transition as $π\searrow π_c$. It has further been conjectured that the entire regime $π<π_c$ for such growing networks are ''critical'' with power-law cluster size distributions having a non-universal exponent for all values of $π\in (0, π_c)$. In this paper, we study percolation on the uniform attachment model, as a concrete template in order to develop general tools based on stochastic approximation, local convergence, branching random walks and tree-graph inequalities to prove the above conjectured phenomena. For each $π\in (0,π_c)$, we show there exists an explicit $α(π) \in (0,\tfrac{1}{2}) $ such that the maximal component size, as well as the size of the component containing any fixed vertex, all re-scaled by $n^{α(π)}$, converge almost surely to strictly positive random variables as the network size $n \to \infty$. These dynamics lead to novel phenomena, compared to classical 'static' models, including long-range dependence and fixation of the identity of the maximal component, within finite time, among a finite number of 'early' components. Moreover, in contrast with most static network models, we show that the susceptibility, that is, the expected size of the component of a uniformly chosen vertex, remains bounded as the network grows and $π$ approaches $π_c$ from below. The general tools developed in this paper will be used in follow-up work to understand percolation for general growing network evolution models.

math.PR↗

Percolation on random graphs

Percolation is a model for random damage to a network. It is one of the simplest models that displays a phase transition: when the network is severely damaged, it falls apart in many small connected components, while if the damage is light, connectivity is hardly affected. We study the location and nature of the phase transition on random graphs. In particular, we focus on the connectivity structure close to, or below, criticality, where components display intricate scaling behaviour such that a typical connected component has bounded size, while the average and maximal connected component sizes grow like powers of the network size. We review the recent progress that has been made in two important settings: random graphs whose expected adjacency matrix is close to being rank-1, the most prominent examples being the configuration model and rank-1 inhomogeneous random graphs, and dynamic random graphs, i.e., random graphs that grow with time, such as uniform and preferential attachment models. Remarkably, these two settings behave rather differently. In all cases, the inhomogeneity of the underlying random graph on which we perform percolation is of crucial importance.

math.PR↗

Long-range first-passage percolation on the torus

We study a geometric version of first-passage percolation on the complete graph, known as long-range first-passage percolation. Here, the vertices of the complete graph $\mathcal K_n$ are embedded in the $d$-dimensional torus $\mathbb T_n^d$, and each edge $e$ is assigned an independent transmission time $T_e=\|e\|_{\mathbb T_n^d}^αE_e$, where $E_e$ is a rate-one exponential random variable associated with the edge $e$, $\|\cdot\|_{\mathbb T_n^d}$ denotes the torus-norm, and $α\geq0$ is a parameter. We are interested in the case $α\in[0,d)$, which corresponds to the instantaneous percolation regime for long-range first-passage percolation on $\mathbb Z^d$ studied by Chatterjee and Dey, and which extends first-passage percolation on the complete graph (the $α=0$ case) studied by Janson. We consider the typical distance, flooding time, and diameter of the model. Our results show a $1,2,3$-type result, akin to first-passage percolation on the complete graph as shown by Janson. The results also provide a quantitative perspective to the qualitative results observed by Chatterjee and Dey on $\mathbb Z^d$.

math.PR↗

The giant in random graphs is almost local

Local convergence techniques have become a key methodology to study sparse random graphs. However, convergence of many random graph properties does not directly follow from local convergence. A notable, and important, such random graph property is the size and uniqueness of the giant component. We provide a simple criterion that guarantees that local convergence of a random graph implies the convergence of the proportion of vertices in the maximal connected component. We further show that, when this condition holds, the local properties of the giant, as well as its complement, are also described by the local limit. We give several examples where this method gives rise to a novel law of large numbers for the giant, based on results proved in the literature. Aside from these examples, we apply our method to the classical problem of giants in the configuration model as a proof of concept, reproving a well-established result. As a side result of this proof, we give an extremely simple proof of the small-world nature of the configuration model.

math.PR↗

Tame sparse exponential random graphs

In this paper, we obtain a precise estimate of the probability that the sparse binomial random graph contains a large number of vertices in a triangle. The estimate of log of this probability is correct up to second order, and enables us to propose an exponential random graph model based on the number of vertices in a triangle. Specifically, by tuning a single parameter, we can with high probability induce any given fraction of vertices in a triangle. Moreover, in the proposed exponential random graph model we derive the large deviation principle for the number of edges. As a byproduct, we propose a consistent estimator of the tuning parameter.

math.PR↗

The asymptotic rank of adjacency matrices of weighted configuration models over arbitrary fields

We study the asymptotic rank of adjacency matrices of a large class of edge-weighted configuration models. Here, the weight of a (multi-)edge can be any fixed non-zero element from an arbitrary field, as long as it is independent of the (multi-)graph. Our main result demonstrates that the asymptotic behavior of the normalized rank of the adjacency matrix neither depends on the fixed edge-weights, nor on which field they are chosen from. Our approach relies on a novel adaptation of the component exploration method of \cite{janson2009new}, which enables the application of combinatorial techniques from \cite{coja2022rank, HofMul25}.

math.CO↗

Local limit of Prim's algorithm

We study the local evolution of Prim's algorithm on large finite weighted graphs. When performed for $n$ steps, where $n$ is the size of the graph, Prim's algorithm will construct the minimal spanning tree (MST). We assume that our graphs converge locally in probability to some limiting rooted graph. In that case, Aldous and Steele already proved that the local limit of the MST converges to a limiting object, which can be thought of as the MST on the limiting infinite rooted graph. Our aim is to investigate {\em how} the local limit of the MST is reached \textit{dynamically}. For this, we take $tn+o(n)$ steps of Prim, for $t\in[0,1]$, and, under some reasonable assumptions, show how the local structure interpolates between performing Prim's algorithm on the local limit when $t=0$, to the full local limit of the MST for $t=1$. Our proof relies on the use of the recently developed theory of {\em dynamic local convergence}. We further present several examples for which our assumptions, and thus our results, apply.

math.PR↗