SearcharxivSearch

arXiv subjects

Clara Stegehuis

Publications and source records attributed to Clara Stegehuis.

At least 19 recordsLinked to original sources

node2vec or triangle-biased random walks: stationarity, regularity & recurrence

The node2vec random walk is a non-Markovian random walk on the vertex set of a graph, widely used for network embedding and exploration. This random walk model is defined in terms of three parameters which control the probability of, respectively, backtracking moves, moves within triangles, and moves to the remaining neighboring nodes. From a mathematical standpoint, the node2vec random walk is a nontrivial generalization of the non-backtracking random walk and thus belongs to the class of second-order Markov chains. Despite its widespread use in applications, little is known about its long-run behavior. The goal of this paper is to begin exploring its fundamental properties on arbitrary graphs. To this aim, we show how lifting the node2vec random walk to the state spaces of directed edges and directed wedges yields two distinct Markovian representations which are key for its asymptotic analysis. Using these representations, we find mild sufficient conditions on the underlying finite or infinite graph to guarantee ergodicity, reversibility, recurrence and characterization of the invariant measure. As we discuss, the behavior of the node2vec random walk is drastically different compared to the non-backtracking random walk. While the latter simplifies on arbitrary graphs when using its natural edge Markovian representation thanks to bistochasticity, the former simplifies on regular graphs when using its natural wedge Markovian representation. Remarkably, this representation reveals that a graph is regular if and only if a certain weighted Eulerianity condition holds.

math.PR

Simplicity of random hypergraphs

Random hypergraphs extend the classical notion of random graphs by allowing hyperedges to join more than two vertices, making them well-suited for modeling higher-order interactions in complex systems. Despite their broad applicability, many structural properties of random hypergraphs remain less understood than in the graph setting. One such property is simplicity: the absence of self-loops, multi-hyperedges, and, in the hypergraph context, degenerate hyperedges where hyperedges contain a copy of the same vertex at least twice. While the behaviour of the number of such self-loops and multi-hyperedges is well understood for random graphs through the configuration model, analogous results for hypergraphs are comparatively sparse. In this work, we study both undirected and directed hypergraphs generated by the configuration model with prescribed vertex and hyperedge degrees. We derive exact, explicit expressions for the expected number of self-loops, multi-hyperedges and degenerate hyperedges, extending classical results from the graph setting. In addition, an asymptotical analysis shows that, under mild moment conditions on the degree distribution, the expected fraction of self-loops, multi-hyperedges and degenerate hyperedges vanishes as the number of vertices grows. Our results provide a systematic understanding of simplicity in directed and undirected hypergraph models.

math.PR

Large deviations for subgraphs in inhomogeneous random graphs

Inhomogeneous random graphs are fundamental models for real-world networks, where prescribed degrees are imposed as soft constraints. A common assumption in such models is that the degree distribution follows a power-law, capturing the heavy-tailed nature observed in many contexts. While various graph functionals have been studied in this setting, inhomogeneity makes their analysis significantly more challenging. The goal of this paper is to investigate the large deviations of subgraph counts in inhomogeneous random graphs. Rare events concerning these functionals translate into quantifying the probability that extremely large hubs appear in the graph. This can be achieved by defining a specific optimization problem that captures the most likely way to generate numerous additional subgraphs. When the expected number of subgraphs is sublinear in the graph size, polynomially large deviations are possible, and in this case, we can derive sharp results on clique counts.

math.PR

Uniformly sampling random directed hypergraphs with fixed degrees

Many complex systems show non-pairwise interactions, which can be captured by hypergraphs. In this work, we propose an edge-swapping method to sample random directed hypergraphs with fixed vertex and hyperarc degrees, which can be applied to different classes of directed hypergraphs (containing self-loops, degenerate hyperarcs and/or multi-hyperarcs). We prove that this method indeed samples uniformly from the classes with self-loops and multi-hyperarcs, and that the method may not sample uniformly from classes without self-loops, or with self-loops and degenerate hyperarcs but without multi-hyperarcs. We present a partial result on the class with self-loops, but without degenerate hyperarcs or multi-hyperarcs.

math.CO

Homophily Within and Across Groups

Homophily -- the tendency of individuals to interact with similar others -- shapes how networks form and function. Yet existing approaches typically collapse homophily to a single scale, either one parameter for the whole network or one per community, thereby detaching it from other structural features. Here, we introduce a maximum-entropy random graph model that moves beyond these limits, capturing homophily across all social scales in the network, with parameters for each group size. The framework decomposes homophily into within- and across-group contributions, recovering the stochastic block model as a special case. As an exponential-family model, it fits empirical data and enables inference of group-level variation of homophily that aggregate metrics miss. The group-dependence of homophily substantially impacts network percolation thresholds, altering predictions for epidemic spread, information diffusion, and the effectiveness of interventions. Ignoring such heterogeneity risks systematically misjudging connectivity and dynamics in complex systems.

physics.soc-ph

Stationary distribution of node2vec random walks on household models

The node2vec random walk has proven to be a key tool in network embedding algorithms. These random walks are tuneable, and their transition probabilities depend on the previous visited node and on the triangles containing the current and the previously visited node. Even though these walks are widely used in practice, most mathematical properties of node2vec walks are largely unexplored, including their stationary distribution. We study the node2vec random walk on community-structured household model graphs. We prove an explicit description of the stationary distribution of node2vec walks in terms of the walk parameters. We then show that by tuning the walk parameters, the stationary distribution can interpolate between uniform, size-biased, or the simple random walk stationary distributions, demonstrating the wide range of possible walks. We further explore these effects on some specific graph settings.

math.PR

Hypercurveball algorithm for sampling hypergraphs with fixed degrees

Comparative analysis between a network and a random graph model can uncover network properties that significantly deviate from those in random networks. The standard random graph model used for comparison uniformly samples random graphs with the same degrees as the network data, often achieved through edge-swap algorithms. However, for hypergraphs, fewer such methodologies are available. This study introduces the Hypercurveball algorithm, designed to sample random, potentially directed, hypergraphs with fixed degrees. Minor adjustments enable the sampling of hypergraphs without degenerate hyperedges, self-loops, or multi-hyperedges. For most of these algorithms, we prove whether they sample uniformly or with bias. We experimentally show that the Hypercurveball algorithm can be significantly faster or slower than the standard hyperedge-shuffling algorithm, which is the hyperedge-equivalent of the edge-swap algorithm. We present criteria on the hypergraph degree sequence that indicate when the Hypercurveball algorithm is more efficient than the standard hyperedge-shuffling method. Finally, our experimental results suggest polynomial scaling of the mixing time for both the Hypercurveball and hyperedge-shuffling algorithms.

math.CO

Maximal Cliques in Scale-Free Random Graphs

We investigate the number of maximal cliques, i.e., cliques that are not contained in any larger clique, in three network models: Erdős-Rényi random graphs, inhomogeneous random graphs (also called Chung-Lu graphs), and geometric inhomogeneous random graphs. For sparse and not-too-dense Erdős-Rényi graphs, we give linear and polynomial upper bounds on the number of maximal cliques. For the dense regime, we give super-polynomial and even exponential lower bounds. Although (geometric) inhomogeneous random graphs are sparse, we give super-polynomial lower bounds for these models. This comes from the fact that these graphs have a power-law degree distribution, which leads to a dense subgraph in which we find many maximal cliques. These lower bounds seem to contradict previous empirical evidence that (geometric) inhomogeneous random graphs have only few maximal cliques. We resolve this contradiction by providing experiments indicating that, even for large networks, the linear lower-order terms dominate, before the super-polynomial asymptotic behavior kicks in only for networks of extreme size.

math.CO

Beam-align: distributed user association for mmWave networks with multi-connectivity

Since the spectrum below 6 GHz bands is insufficient to meet the high bandwidth requirements of 5G use cases, 5G networks expand their operation to mmWave bands. However, operation at these bands has to cope with a high penetration loss and susceptibility to blocking objects. Beamforming and multi-connectivity (MC) can together mitigate these challenges. But, to design such an optimal user association scheme leveraging these two features is non-trivial and computationally expensive. Previous studies either considered a fixed MC degree for all users or overlooked beamforming. Driven by the question what is the optimal degree of MC for each user in a mmWave network, we formulate a user association scheme that maximizes throughput considering beam formation and MC. Our numerical analysis shows that there is no one-size-fits-all degree of optimal MC; it depends on the number of users, their rate requirements, locations, and the maximum number of active beams at a BS.Based on the optimal association, we design BEAM-ALIGN: an efficient heuristic with polynomial-time complexity O(|U|log|U|), where |U| is the number of users. Moreover, BEAM-ALIGN only uses local BS information - i.e. the received signal quality at the user. Differing from prior works, BEAM-ALIGN considers beamforming, multiconnectivity and line-of-sight probability. Via simulations, we show that BEAM-ALIGN performs close to optimal in terms of per-user capacity and satisfaction while it outperforms frequently-used signal-to-interference-and-noise-ratio based association schemes. We then show that BEAM-ALIGN has a robust performance under various challenging scenarios: the presence of blockers, rain, and clustered users.

cs.NI

Reconstruction of geometric random graphs with the Simple algorithm

Graph reconstruction can efficiently detect the underlying topology of massive networks such as the Internet. Given a query oracle and a set of nodes, the goal is to obtain the edge set by performing as few queries as possible. An algorithm for graph reconstruction is the Simple algorithm (Mathieu & Zhou, 2023), which reconstructs bounded-degree graphs in $\tilde{O}(n^{3/2})$ queries. We extend the use of this algorithm to the class of geometric random graphs with connection radius $r \sim n^k$, with diverging average degree. We show that for this class of graphs, the query complexity is $\tilde{O}(n^{2k+1})$ when k > 3/20. This query complexity is up to a polylog(n) term equal to the number of edges in the graph, which means that the reconstruction algorithm is almost edge-optimal. We also show that with only $n^{1+o(1)}$ queries it is already possible to reconstruct at least 75% of the non-edges of a geometric random graph, in both the sparse and dense setting. Finally, we show that the number of queries is indeed of the same order as the number of edges on the basis of simulations.

cs.DS

Optimal subgraphs in geometric scale-free random graphs

Geometric scale-free random graphs are popular models for networks that exhibit as heavy-tailed degree distributions, small-worldness and high clustering. In these models, vertices have weights that cause the heavy-tailed degrees and are embedded in a metric space so that close-by groups of vertices tend to cluster. The interplay between the vertex weights and positions heavily affects the local structure of the random graph, in particular the occurrence of subgraph patterns, but the dependencies in these structures and weights make them difficult to analyze. In this paper we investigate subgraph counts using a \textit{divide et impera} strategy: first counting the number of subgraphs in specific classes of vertices; then computing which class yields maximum contribution. Interestingly, the scaling behavior of induced and general subgraphs in such geometric heavy-tailed random graphs is closely related to the solution of a mixed-integer linear program which also shows that subgraphs appear predominantly on vertices with some prescribed degrees and inter-distances. Finally, we derive precise asymptotics for trees and Hamiltonian subgraphs.

math.PR

Large deviations for triangles in scale-free random graphs

We provide large deviations estimates for the upper tail of the number of triangles in scale-free inhomogeneous random graphs where the degrees have power law tails with index $-α, α\in (1,2)$. We show that upper tail probabilities for triangles undergo a phase transition. For $α<4/3$, the upper tail is caused by many vertices of degree of order $n$, and this probability is semi-exponential. In this regime, additional triangles consist of two hubs. For $α>4/3$ on the other hand, the upper tail is caused by one hub of a specific degree, and this probability decays polynomially in $n$, leading to additional triangles with one hub. In the intermediate case $α=4/3$, we show polynomial decay of the tail probability caused by multiple but finitely many hubs. In this case, the additional triangles contain either a single hub or two hubs. Our proofs are partly based on various concentration inequalities. In particular, we tailor concentration bounds for empirical processes to make them well-suited for analyzing heavy-tailed phenomena in nonlinear settings.

math.PR

Resource and location sharing in wireless networks

With more and more demand from devices to use wireless communication networks, there has been an increased interest in resource sharing among operators, to give a better link quality. However, in the analysis of the benefits of resource sharing among these operators, the important factor of co-location is often overlooked. Indeed, often in wireless communication networks, different operators co-locate: they place their base stations at the same locations due to cost efficiency. We therefore use stochastic geometry to investigate the effect of co-location on the benefits of resource sharing. We develop an intricate relation between the co-location factor and the optimal radius to operate the network, which shows that indeed co-location is an important factor to take into account. We also investigate the limiting behavior of the expected gains of sharing, and find that for unequal operators, sharing may not always be beneficial when taking co-location into account.

cs.NI

On the distances within cliques in a soft random geometric graph

We study the distances of edges within cliques in a soft random geometric graph on a torus, where the vertices are points of a homogeneous Poisson point process, and far-away points are less likely to be connected than nearby points. We obtain the scaling of the maximal distance between any two points within a clique of size $k$. Moreover, we show that asymptotically in all cliques with large distances, there is only one remote point and all other points are nearby. Furthermore, we prove that a re-scaled version of the maximal $k$-clique distance converges in distribution to a Fréchet distribution. Thereby, we describe the order of magnitude according to which the largest distance between two points in a clique decreases with the clique size.

math.PR

Detecting hyperbolic geometry in networks: why triangles are not enough

In the past decade, geometric network models have received vast attention in the literature. These models formalize the natural idea that similar vertices are likely to connect. Because of that, these models are able to adequately capture many common structural properties of real-world networks, such as self-invariance and high clustering. Indeed, many real-world networks can be accurately modeled by positioning vertices of a network graph in hyperbolic spaces. Nevertheless, if one observes only the network connections, the presence of geometry is not always evident. Currently, triangle counts and clustering coefficients are the standard statistics to signal the presence of geometry. In this paper we show that triangle counts or clustering coefficients are insufficient because they fail to detect geometry induced by hyperbolic spaces. We therefore introduce a novel triangle-based statistic, which weighs triangles based on their strength of evidence for geometry. We show analytically, as well as on synthetic and real-world data, that this is a powerful statistic to detect hyperbolic geometry in networks.

physics.soc-ph

On the resilience of cellular networks: how can national roaming help?

Cellular networks have become one of the critical infrastructures, as many services depend increasingly on wireless connectivity. Therefore, it is important to quantify the resilience of existing cellular network infrastructures against potential risks, ranging from natural disasters to security attacks, that might occur with a low probability but can lead to severe disruption of the services. In this paper, we combine models with public data from national bodies on mobile network operator (MNO) infrastructures, population distribution, and urbanity level to assess the coverage and capacity of a cellular network at a country scale. Our analysis offers insights on the potential weak points that need improvement to ensure a low fraction of disconnected population (FDP) and high fraction of satisfied population (FSP). As a resilience improvement approach, we investigate in which regions and to what extent each MNO can benefit from infrastructure sharing or national roaming, i.e., all MNOs act as a single national operator. As our case study, we focus on Dutch cellular infrastructure and model risks as random failures, correlated failures in a geographic region, and abrupt increase in the number of users. Our analysis shows that there is a wide performance difference across MNOs and geographic regions in terms of FDP and FSP. However, national roaming consistently offers significant benefits, e.g., up to 13% improvement in FDP and up to 55% in FSP when the networks function without any failures.

cs.NI

Scale-free graphs with many edges

We develop tail estimates for the number of edges in a Chung-Lu random graph with regularly varying weight distribution. Our results show that the most likely way to have an unusually large number of edges is through the presence of one or more hubs, i.e.\ vertices with degree of order $n$.

math.PR

Localized geometry detection in scale-free random graphs

We consider the problem of detecting whether a power-law inhomogeneous random graph contains a geometric community, and we frame this as an hypothesis testing problem. More precisely, we assume that we are given a sample from an unknown distribution on the space of graphs on n vertices. Under the null hypothesis, the sample originates from the inhomogeneous random graph with a heavy-tailed degree sequence. Under the alternative hypothesis, $k = o(n)$ vertices are given spatial locations and connect between each other following the geometric inhomogeneous random graph connection rule. The remaining $n-k$ vertices follow the inhomogeneous random graph connection rule. We propose a simple and efficient test, which is based on counting normalized triangles, to differentiate between the two hypotheses. We prove that our test correctly detects the presence of the community with high probability as $n \to \infty$, and identifies large-degree vertices of the community with high probability.

math.ST