Searcharxiv⌕ Search

arXiv subjects

Paweł Prałat

Publications and source records attributed to Paweł Prałat.

At least 19 recordsLinked to original sources

Graph Data Augmentation via Contrastive Generator Inversion ($\texttt{DCBA}$)

Graphs provide a natural representation of many complex systems, ranging from social platforms to ecosystems. However, the development of graph-based machine learning methods is often constrained by the limited availability of large and diverse graph datasets. In this paper, we introduce $\texttt{DCBA}$, a model-based approach to graph data augmentation that infers the configuration of a synthetic graph generator from an observed network. We instantiate the proposed framework using the $\texttt{ABCD}$ generator, which produces scale-free networks with community structure. Our model learns a joint representation of graphs and generator parametrisations using a multi-positive contrastive objective with soft negative weighting. The learned representation enables the prediction of an $\texttt{ABCD}$ configuration whose stochastic realisations preserve the macrostructural properties encoded by the generator. Experiments show that $\texttt{DCBA}$ recovers generator parameters more accurately and robustly than an algorithmic inverse-modelling baseline. Its downstream utility is further demonstrated in community detection, where inferred configurations used to fine-tune $\texttt{PRoCD}$ improve AMI on average by $161\%$ on synthetic and $273\%$ on real-world networks.

cs.LG↗

The critical activation density in graph bootstrap percolation

In graph bootstrap percolation, edges of an Erdős-Rényi random graph ${\mathcal G}_{n,p}$ are initially active, and activation spreads to other edges of $K_n$ via the combinatorics of a fixed graph $H$: an edge becomes active whenever it is the unique inactive edge in a copy of $H$. The process $H$-percolates if all edges of $K_n$ are eventually activated. While classical cases such as $H=K_3$ (connectivity) and $H=K_4$ (related to $2$-neighbor bootstrap percolation) have been studied extensively, general graphs $H$ can exhibit wildly different behaviors. In this work, we determine the critical $H$-percolation threshold $p_c(n,H)$ for every graph $H$, fully resolving a longstanding open question of Balogh, Bollobás, and Morris. The location of $p_c(n,H)$ is governed by a new, universal parameter $ρ(H)$, which measures the maximal efficiency of witness graphs that activate an edge. To achieve this, we introduce a novel framework based on the unfolding and refolding of witness graphs. While previous works were restricted to specific families of $H$, our approach provides a unified strategy for all $H$. Inspired by algebraic topology, we lift witness graphs to covering graphs and algorithmically embed folded versions into ${\mathcal G}_{n,p}$ via a sequence of extensions. Crucially, this allows us to incorporate highly efficient witness graphs of unbounded size, which are potentially far larger than ${\mathcal G}_{n,p}$ itself. Beyond resolving $p_c(n,H)$, our framework recovers and strengthens several existing bounds in the literature. Finally, we initiate the study of the universal density parameter $ρ(H)$ and pose central open questions regarding its computability and its exact correspondence with the sharpness of the $H$-percolation threshold.

math.PR↗

Twinning Complex Networked Systems: Data-Driven Calibration of the mABCD Synthetic Graph Generator

The increasing availability of relational data has contributed to a growing reliance on network-based representations of complex systems. Over time, these models have evolved to capture more nuanced properties, such as the heterogeneity of relationships, leading to the concept of multilayer networks. However, the analysis and evaluation of methods for these structures is often hindered by the limited availability of large-scale empirical data. As a result, graph generators are commonly used as a workaround, albeit at the cost of introducing systematic biases. In this paper, we address the inverse-generator problem by inferring the configuration parameters of a multilayer network generator, \mABCD, from a real-world system. Our goal is to identify parameter settings that enable the generator to produce synthetic networks that act as digital twins of the original structure. We propose a method for estimating matching configurations and for quantifying the associated error. Our results demonstrate that this task is non-trivial, as strong interdependencies between configuration parameters weaken independent estimation and instead favour a joint-prediction approach.

cs.SI↗

A Pragmatic Method for Comparing Clusterings with Overlaps and Outliers

Clustering algorithms are an essential part of the unsupervised data science ecosystem, and extrinsic evaluation of clustering algorithms requires a method for comparing the detected clustering to a ground truth clustering. In a general setting, the detected and ground truth clusterings may have outliers (objects belonging to no cluster), overlapping clusters (objects may belong to more than one cluster), or both, but methods for comparing these clusterings are currently undeveloped. In this note, we define a pragmatic similarity measure for comparing clusterings with overlaps and outliers, show that it has several desirable properties, and experimentally confirm that it is not subject to several common biases afflicting other clustering comparison measures.

cs.LG↗

The Needle is a Thread: Finding Planted Paths in Noisy Process Trees

Motivated by applications in cybersecurity such as finding meaningful sequences of malware-related events buried inside large amounts of computer log data, we introduce the "planted path" problem and propose an algorithm to find fuzzy matchings between two trees. This algorithm can be used as a "building block" for more complicated workflows. We demonstrate usefulness of a few of such workflows in mining synthetically generated data as well as real-world ACME cybersecurity datasets.

cs.SI↗

The Artificial Benchmark for Community Detection with Outliers and Overlapping Communities (ABCD+$o^2$)

The Artificial Benchmark for Community Detection (ABCD) graph is a random graph model with community structure and power-law distribution for both degrees and community sizes. The model generates graphs similar to the well-known LFR model but it is faster, more interpretable, and can be investigated analytically. In this paper, we use the underlying ingredients of the ABCD model, and its generalization to include outliers (ABCD+$o$), and introduce another variant that allows for overlapping communities, ABCD+$o^2$.

cs.SI↗

Multilayer Artificial Benchmark for Community Detection (mABCD)

One of the most persistent challenges in network science is the development of various synthetic graph models to support subsequent analyses. Among the most notable frameworks addressing this issue is the Artificial Benchmark for Community Detection (ABCD) model, a random graph model with community structure and power-law distribution for both degrees and community sizes. The model generates graphs similar to the well-known LFR model but it is faster, more interpretable, and can be investigated analytically. In this paper, we use the underlying ingredients of ABCD and introduce its variant, mABCD, thereby addressing the gap in models capable of generating multilayer networks. The uniqueness of the proposed approach lies in its flexibility at both levels of modelling: the internal structure of individual layers and the inter-layer dependencies, which together make the network a coherent structure rather than a collection of loosely coupled graphs. In addition to the conceptual description of the framework, we provide a comprehensive analysis of its efficient Julia implementation. Finally, we illustrate the applicability of mABCD to one of the most prominent problems in the area of complex systems: spreading phenomena analysis.

cs.SI↗

Direct Paths in the Temporal Hypercube

We consider the $n$-dimensional random temporal hypercube, i.e., the $n$-dimensional hypercube graph with its edges endowed with i.i.d. continuous random weights. We say that a vertex $w$ is accessible from another vertex $v$ if and only if there is a path from $v$ to $w$ with increasing edge weights. We study accessible "direct" paths from a fixed vertex to its antipodal point and show that as $n \to \infty$, the number of such paths converges in distribution to a mixed Poisson law with mixture given by the product of two independent exponentials with rate $1$. Our proof makes use of the Chen-Stein method, coupling arguments, as well as combinatorial arguments which show that typical pairs of accessible paths have small overlap.

math.PR↗

Label propagation on binomial random graphs

We study the behavior of a label propagation algorithm (LPA) on the Erdős-Rényi random graph $\mathcal{G}(n,p)$. Initially, given a network, each vertex starts with a random label in the interval $[0,1]$. Then, in each round of LPA, every vertex switches its label to the majority label in its neighborhood (including its own label). At the first round, ties are broken towards smaller labels, while at each of the next rounds, ties are broken uniformly at random. The algorithm terminates once all labels stay the same in two consecutive iterations. LPA is successfully used in practice for detecting communities in networks (corresponding to vertex sets with the same label after termination of the algorithm). Perhaps surprisingly, LPA's performance on dense random graphs is hard to analyze, and so far convergence to consensus was known only when $np\ge n^{3/4+\varepsilon}$, where LPA converges in three rounds. By defining an alternative label attribution procedure which converges to the label propagation algorithm after three rounds, a careful multi-stage exposure of the edges allows us to break the $n^{3/4+\varepsilon}$ barrier and show that, when $np \ge n^{5/8+\varepsilon}$, a.a.s.\ the algorithm terminates with a single label. Moreover, we show that, if $np\gg n^{2/3}$, a.a.s.\ this label is the smallest one, whereas if $n^{5/8+\varepsilon}\le np\ll n^{2/3}$, the surviving label is a.a.s.\ not the smallest one. En passant, we show a presumably new monotonicity lemma for Binomial random variables that might be of independent interest.

math.PR↗

Network Embedding Exploration Tool (NEExT)

Many real-world and artificial systems and processes can be represented as graphs. Some examples of such systems include social networks, financial transactions, supply chains, and molecular structures. In many of these cases, one needs to consider a collection of graphs, rather than a single network. This could be a collection of distinct but related graphs, such as different protein structures or graphs resulting from dynamic processes on the same network. Examples of the latter include the evolution of social networks, community-induced graphs, or ego-nets around various nodes. A significant challenge commonly encountered is the absence of ground-truth labels for graphs or nodes, necessitating the use of unsupervised techniques to analyze such systems. Moreover, even when ground-truth labels are available, many existing graph machine learning methods depend on complex deep learning models, complicating model explainability and interpretability. To address some of these challenges, we have introduced NEExT (Network Embedding Exploration Tool) for embedding collections of graphs via user-defined node features. The advantages of the framework are twofold: (i) the ability to easily define your own interpretable node-based features in view of the task at hand, and (ii) fast embedding of graphs provided by the Vectorizers library. In this paper, we demonstrate the usefulness of NEExT on collections of synthetic and real-world graphs. For supervised tasks, we demonstrate that performance in graph classification tasks could be achieved similarly to other state-of-the-art techniques while maintaining model interpretability. Furthermore, our framework can also be used to generate high-quality embeddings in an unsupervised way, where target variables are not available.

cs.LG↗

Improving community detection via community association strength scores

Community detection methods play a central role in understanding complex networks by revealing highly connected subsets of entities. However, most community detection algorithms generate partitions of the nodes, thus (i) forcing every node to be part of a community and (ii) ignoring the possibility that some nodes may be part of multiple communities. In our work, we investigate three simple community association strength (CAS) scores and their usefulness as post-processing tools given some partition of the nodes. We show that these measures can be used to improve node partitions, detect outlier nodes (not part of any community), and help find nodes with multiple community memberships.

cs.SI↗

Building Hamiltonian Cycles in the Semi-Random Graph Process in Less Than $2n$ Rounds

The semi-random graph process is an adaptive random graph process in which an online algorithm is initially presented an empty graph on $n$ vertices. In each round, a vertex $u$ is presented to the algorithm independently and uniformly at random. The algorithm then adaptively selects a vertex $v$, and adds the edge $uv$ to the graph. For a given graph property, the objective of the algorithm is to force the graph to satisfy this property asymptotically almost surely in as few rounds as possible. We focus on the property of Hamiltonicity. We present an adaptive strategy which creates a Hamiltonian cycle in $αn$ rounds, where $α< 1.81696$ is derived from the solution to a system of differential equations. We also show that achieving Hamiltonicity requires at least $βn$ rounds, where $β> 1.26575$.

math.CO↗

Counting simplicial pairs in hypergraphs

We present two ways to measure the simplicial nature of a hypergraph: the simplicial ratio and the simplicial matrix. We show that the simplicial ratio captures the frequency, as well as the rarity, of simplicial interactions in a hypergraph while the simplicial matrix provides more fine-grained details. We then compute the simplicial ratio, as well as the simplicial matrix, for 10 real-world hypergraphs and, from the data collected, hypothesize that simplicial interactions are more and more deliberate as edge size increases. We then present a new Chung-Lu model that includes a parameter controlling (in expectation) the frequency of simplicial interactions. We use this new model, as well as the real-world hypergraphs, to show that multiple stochastic processes exhibit different behaviour when performed on simplicial hypergraphs vs. non-simplicial hypergraphs.

cs.SI↗

Modularity Based Community Detection in Hypergraphs

In this paper, we propose a scalable community detection algorithm using hypergraph modularity function, h-Louvain. It is an adaptation of the classical Louvain algorithm in the context of hypergraphs. We observe that a direct application of the Louvain algorithm to optimize the hypergraph modularity function often fails to find meaningful communities. We propose a solution to this issue by adjusting the initial stage of the algorithm via carefully and dynamically tuned linear combination of the graph modularity function of the corresponding two-section graph and the desired hypergraph modularity function. The process is guided by Bayesian optimization of the hyper-parameters of the proposed procedure. Various experiments on synthetic as well as real-world networks are performed showing that this process yields improved results in various regimes.

cs.SI↗

Predicting Properties of Nodes via Community-Aware Features

This paper shows how information about the network's community structure can be used to define node features with high predictive power for classification tasks. To do so, we define a family of community-aware node features and investigate their properties. Those features are designed to ensure that they can be efficiently computed even for large graphs. We show that community-aware node features contain information that cannot be completely recovered by classical node features or node embeddings (both classical and structural) and bring value in node classification tasks. This is verified for various classification tasks on synthetic and real-life networks.

cs.SI↗

Linear Colouring of Binomial Random Graphs

We investigate the linear chromatic number $χ_{\text{lin}}(G(n,p))$ of the binomial random graph $G(n,p)$ on $n$ vertices in which each edge appears independently with probability $p=p(n)$. For dense random graphs ($np \to \infty$ as $n \to \infty$), we show that asymptotically almost surely $χ_{\text{lin}}(G(n,p)) \ge n (1 - O( (np)^{-1/2} ) ) = n(1-o(1))$. Understanding the order of the linear chromatic number for subcritical random graphs ($np < 1$) and critical ones ($np=1$) is relatively easy. However, supercritical sparse random graphs ($np = c$ for some constant $c > 1$) remain to be investigated.

math.CO↗

Power of $k$ Choices in the Semi-Random Graph Process

The semi-random graph process is a single player game in which the player is initially presented an empty graph on $n$ vertices. In each round, a vertex $u$ is presented to the player independently and uniformly at random. The player then adaptively selects a vertex $v$, and adds the edge $uv$ to the graph. For a fixed monotone graph property, the objective of the player is to force the graph to satisfy this property with high probability in as few rounds as possible. In this paper, we introduce a natural generalization of this game in which $k$ random vertices $u_1, \ldots, u_k$ are presented to the player in each round. She needs to select one of the presented vertices and connect to any vertex she wants. We focus on the following three monotone properties: minimum degree at least $\ell$, the existence of a perfect matching, and the existence of a Hamiltonian cycle.

math.CO↗

Hypergraph Artificial Benchmark for Community Detection (h-ABCD)

The Artificial Benchmark for Community Detection (ABCD) graph is a recently introduced random graph model with community structure and power-law distribution for both degrees and community sizes. The model generates graphs with similar properties as the well-known LFR one, and its main parameter can be tuned to mimic its counterpart in the LFR model, the mixing parameter. In this paper, we introduce hypergraph counterpart of the ABCD model, h-ABCD, which produces random hypergraph with distributions of ground-truth community sizes and degrees following power-law. As in the original ABCD, the new model h-ABCD can produce hypergraphs with various levels of noise. More importantly, the model is flexible and can mimic any desired level of homogeneity of hyperedges that fall into one community. As a result, it can be used as a suitable, synthetic playground for analyzing and tuning hypergraph community detection algorithms.

cs.SI↗