SearcharxivSearch

arXiv subjects

Frank W. Takes

Publications and source records attributed to Frank W. Takes.

At least 19 recordsLinked to original sources

Fuzzy k-anonymity in complex networks

With the introduction of large-scale network data, including population-scale social networks, techniques for privacy-aware sharing of network data become increasingly important. While existing $k$-anonymity approaches can model different attacker scenarios, they typically assume that attacker knowledge exactly matches the published network structure. We argue that exact knowledge is often unrealistic and introduce $ϕ$-$k$-anonymity, a fuzzy variant of $k$-anonymity in which parameter $ϕ$ captures the level of uncertainty in attacker knowledge. Across a benchmark of $39$ real-world networks, a realistic level of uncertainty ($ϕ=5\%$) renders, on average, $64\%$ of previously unique nodes anonymous. To further enhance anonymity, we apply anonymization algorithms under a $5\%$ edge modification budget. While full anonymization is often unattainable under exact $k$-anonymity, with low uncertainty ($ϕ=10\%$) our newly proposed Greedy algorithm anonymizes over $99\%$ of the nodes. Uncertainty also enables effective anonymization in otherwise difficult to anonymize dense synthetic graphs. Additionally, data utility in terms of structural properties and performance on network analysis tasks is well preserved, with most metrics changing less than $5\%$. Overall, our findings suggest that modest uncertainty assumptions yield high levels of anonymity and utility, motivating further research on uncertainty-aware privacy guarantees for network data.

cs.SI

Link Fraction Mixed Membership Reveals Community Diversity in Aggregated Social Networks

Community detection is a critical tool for understanding the mesoscopic structure of large-scale networks. However, when applied to aggregated or coarse-grained social networks, disjoint community partitions cannot capture the diverse composition of community memberships within aggregated nodes. While existing mixed membership methods alleviate this issue, they may detect communities that are highly sensitive to the aggregation resolution, not reliably reflecting the community structure of the underlying individual-level network. This paper presents the Link Fraction Mixed Membership (LFMM) method, which computes the mixed memberships of nodes in aggregated networks. Unlike existing mixed membership methods, LFMM is consistent under aggregation. Specifically, we show that it conserves community membership sums at different scales. The method is utilized to study a population-scale social network of the Netherlands, aggregated at different resolutions. Experiments reveal variation in community membership across different geographical regions and evolution over the last decade. In particular, we show how our method identifies large urban hubs that act as the melting pots of diverse, spatially remote communities.

cs.SI

The anonymization problem in social networks

This paper introduces a unified computational framework for the anonymization problem in social networks, where the objective is to maximize node anonymity through graph alterations. We define three variants of the underlying optimization problem: full, partial and budgeted anonymization. In each variant, the objective is to maximize the number of $k$-anonymous nodes, i.e., nodes for which at least $k-1$ other nodes are equivalent under a particular anonymity measure. We propose four new heuristic network anonymization algorithms and implement these in ANO-NET, a reusable computational framework. Experiments on three common graph models and 19 real-world network datasets yield three empirical findings. First, regarding the method of alteration, experiments on graph models show that random edge deletion is more effective than edge rewiring and addition. Second, we show that the choice of anonymity measure strongly affects both initial network anonymity and the difficulty of anonymization. This highlights the importance of careful measure selection, matching a realistic attacker scenario. Third, comparing the four proposed algorithms and an edge sampling baseline from the literature, we find that an approach which preferentially deletes edges affecting structurally unique nodes, consistently outperforms heuristics based solely on network structure. Overall, our best performing algorithm retains on average 14 times more edges in full anonymization. Moreover, it yields 4.8 times more anonymous nodes than the baseline in the budgeted variant. On top of that, the best performing algorithm achieves a better trade-off between anonymity and data utility. This work provides a foundation for the future development of effective network anonymization algorithms.

cs.SI

Ecological Legacies of Pre-Columbian Settlements Evident in Palm Clusters of Neotropical Mountain Forests

Ancient populations inhabited and transformed neotropical forests, yet the spatial extent of their ecological influence remains underexplored at high resolution. Here we present a deep learning and remote sensing based approach to estimate areas of pre-Columbian forest modification based on modern vegetation. We apply this method to high-resolution satellite imagery from the Sierra Nevada de Santa Marta, Colombia, as a demonstration of a scalable approach, to evaluate palm tree distributions in relation to archaeological infrastructure. Our findings document a non-random spatial association between archaeological infrastructure and contemporary palm concentrations. Palms were significantly more abundant near archaeological sites with large infrastructure investment. The extent of the largest palm cluster indicates that ancient human-managed areas linked to major infrastructure sites may be up to two orders of magnitude bigger than indicated by current archaeological evidence alone. These patterns are consistent with the hypothesis that past human activity may have influenced local palm abundance and potentially reduced the logistical costs of establishing infrastructure-heavy settlements in less accessible locations. More broadly, our results highlight the utility of palm landscape distributions as an interpretable signal within environmental and multispectral datasets for constraining predictive models of archaeological site locations.

cs.CV

Fast degree-preserving rewiring of complex networks

In this paper we introduce a new, fast, degree-preserving rewiring algorithm for altering the assortativity of complex networks, which we call \textit{Fast total link (FTL) rewiring} algorithm. Commonly used existing algorithms require a large number of iterations, in particular in the case of large dense networks. This can especially be problematic when we wish to study ensembles of networks. In this work we aim to overcome aforementioned scalability problems by performing a rewiring of all edges at once to achieve a very high assortativity value before rewiring samples of edges at once to reduce this high assortativity value to the target value. The proposed method performs better than existing methods by several orders of magnitude for a range of structurally diverse complex networks, both in terms of the number of iterations taken, and time taken to reach a given assortativity value. Here we test our proposed algorithm on networks with up to $100,000$ nodes and around $750,000$ edges and find that the relative improvements in speed remain, showing that the algorithm is both efficient and scalable.

physics.soc-ph

Individual Fairness in Community Detection: Quantitative Measure and Comparative Evaluation

Community detection is a fundamental task in complex network analysis. Fairness-aware community detection seeks to prevent biased node partitions, typically framed in terms of individual fairness, which requires similar nodes to be treated similarly, and group fairness, which aims to avoid disadvantaging specific groups of nodes. While existing literature on fair community detection has primarily focused on group fairness, we introduce a novel measure to quantify individual fairness in community detection methods. The proposed measure captures unfairness as the vectorial distance between a node's true and predicted community representations, computed using the community co-occurrence matrix. We provide a comprehensive empirical investigation of a broad set of community detection algorithms from the literature on both synthetic networks, with varying levels of community explicitness, and real-world networks. We particularly investigate the fairness-performance trade-off using standard quality metrics and compare individual fairness outcomes with existing group fairness measures. The results show that individual unfairness can occur even when group fairness or clustering accuracy is high, underscoring that individual and group fairness are not interchangeable. Moreover, fairness depends critically on the detectability of community structure. However, we find that Significance and Surprise for denser graphs, and Combo, Leiden, and SBMDL for sparser graphs result in a better trade-off between individual fairness and community quality. Overall, our findings, together with the fact that community detection is an important step in many network analysis downstream tasks, highlight the necessity of developing fairness-aware community detection methods.

cs.SI

Fragmentation of a longitudinal population-scale social network: Decreasing structural social cohesion in the Netherlands

Population-level dynamics of social cohesion and its underlying mechanisms remain difficult to study. In this paper, we propose a network approach to measure the evolution of social cohesion at the population scale and identify mechanisms driving the change. We use twelve annual snapshots (2010-2021) of a population-scale social network from the Netherlands linking all residents through family, household, work, school, and neighbor relations. Results show that over this period, social cohesion, quantified as average closure in the network, declines by more than 15%. We demonstrate that the decline is not due to changes in demographic composition, but to rewiring in individual ego networks. Statistical models confirm a decreasing overlap of social contexts and greater geographical mobility as drivers. Residential relocation, however, temporarily increases closure, suggesting that local cohesion-seeking behavior can yield global network fragmentation, with implications for policies related to housing, urban planning, and social integration.

physics.soc-ph

Can social capital remedy structural inequality? Economic mobility in a longitudinal population-scale social network

The promise of equal opportunity is a cornerstone of modern societies, yet upward economic mobility remains out of reach for many. Using a decade of population-scale social network data from the Netherlands, covering over a billion family, school, workplace, and neighborhood ties, we examine how structural inequality and social capital jointly shape economic trajectories. Parental background is a strong early predictor of economic outcomes, but its influence fades over time. In contrast, bridging social capital is what positively predicts long-term mobility, particularly for economically disadvantaged groups. Reducing the dimensionality of an individual's network composition, we identify two key dimensions: exposure to affluent contacts and socioeconomic diversity of one's network. These are sufficient to capture the core aspects of social capital that matter for economic mobility. Overall, our findings demonstrate that while inherited advantage shapes the starting point of economic trajectory, social capital can powerfully reshape it, especially for the poor.

physics.soc-ph

Freshness, Persistence and Success of Scientific Teams

Team science dominates scientific knowledge production, but what makes academic teams successful? Using temporal data on 25.2 million publications and 31.8 million authors, we propose a novel network-driven approach to identify and study the success of persistent teams. Challenging the idea that persistence alone drives success, we find that team freshness - new collaborations built on prior experience - is key to success. High impact research tends to emerge early in a team's lifespan. Analyzing complex team overlap, we find that teams open to new collaborative ties consistently produce better science. Specifically, team re-combinations that introduce new freshness impulses sustain success, while persistence impulses from experienced teams are linked to earlier impact. Together, freshness and persistence shape team success across collaboration stages.

cs.DL

A systematic comparison of measures for publishing k-anonymous social network data

Sharing or publishing social network data while accounting for privacy of individuals is a difficult task due to the interconnectedness of nodes in networks. A key question in k-anonymity, a widely studied notion of privacy, is how to measure the anonymity of an individual, as this determines the attacker scenarios one protects against. In this paper, we systematically compare the most prominent anonymity measures from the literature in terms of the completeness and reach of the structural information they take into account. We present a theoretical characterization and a distance-parametrized strictness ordering of the existing measures for k-anonymity in networks. In addition, we conduct empirical experiments on a wide range of real-world network datasets with up to millions of edges. Our findings reveal that the choice of the measure significantly impacts the measured level of anonymity and hence the effectiveness of the corresponding attacker scenario, the privacy vs. utility trade-off, and computational cost. Surprisingly, we find that the anonymity measure representing the most effective attacker scenario considers a greater node vicinity yet utilizes only limited structural information and therewith minimal computational resources. Overall, the insights provided in this work offer researchers and practitioners practical guidance for selecting appropriate anonymity measures when sharing or publishing social network data under privacy constraints.

cs.SI

Quantifying Group Fairness in Community Detection

Understanding community structures is crucial for analyzing networks, as nodes join communities that collectively shape large-scale networks. In real-world settings, the formation of communities is often impacted by several social factors, such as ethnicity, gender, wealth, or other attributes. These factors may introduce structural inequalities; for instance, real-world networks can have a few majority groups and many minority groups. Community detection algorithms, which identify communities based on network topology, may generate unfair outcomes if they fail to account for existing structural inequalities, particularly affecting underrepresented groups. In this work, we propose a set of novel group fairness metrics to assess the fairness of community detection methods. Additionally, we conduct a comparative evaluation of the most common community detection methods, analyzing the trade-off between performance and fairness. Experiments are performed on synthetic networks generated using LFR, ABCD, and HICH-BA benchmark models, as well as on real-world networks. Our results demonstrate that the fairness-performance trade-off varies widely across methods, with no single class of approaches consistently excelling in both aspects. We observe that Infomap and Significance methods are high-performing and fair with respect to different types of communities across most networks. The proposed metrics and findings provide valuable insights for designing fair and effective community detection algorithms.

cs.SI

Utility-aware Social Network Anonymization using Genetic Algorithms

Social networks may contain privacy-sensitive information about individuals. The objective of the network anonymization problem is to alter a given social network dataset such that the number of anonymous nodes in the social graph is maximized. Here, a node is anonymous if it does not have a unique surrounding network structure. At the same time, the aim is to ensure data utility, i.e., preserve topological network properties and retain good performance on downstream network analysis tasks. We propose two versions of a genetic algorithm tailored to this problem: one generic GA and a uniqueness-aware GA (UGA). The latter aims to target edges more effectively during mutation by avoiding edges connected to already anonymous nodes. After hyperparameter tuning, we compare the two GAs against two existing baseline algorithms on several real-world network datasets. Results show that the proposed genetic algorithms manage to anonymize on average 14 times more nodes than the best baseline algorithm. Additionally, data utility experiments demonstrate how the UGA requires fewer edge deletions, and how our GAs and the baselines retain performance on downstream tasks equally well. Overall, our results suggest that genetic algorithms are a promising approach for finding solutions to the network anonymization problem.

cs.SI

The relevance of higher-order ties

Higher-order networks effectively represent complex systems with group interactions. Existing methods usually overlook the relative contribution of group interactions (hyperlinks) of different sizes to the overall network structure. Yet, this has many important applications, especially when the network has meaningful node labels. In this work, we propose a comprehensive methodology to precisely measure the contribution of different orders to the overall network structure. First, we propose the order contribution measure, which quantifies the contribution of hyperlinks of different orders to the link weights (local scale), number of triangles (mesoscale) and size of the largest connected component (global scale) of the pairwise weighted network. Second, we propose the measure of order relevance, which gives insights in how hyperlinks of different orders contribute to the considered network property. Most interestingly, it enables an assessment of whether this contribution is synergistic or redundant with respect to that of hyperlinks of other orders. Third, to account for labels, we propose a metric of label group balance to assess how hyperlinks of different orders connect label-induced groups of nodes. We applied these metrics to a large-scale board interlock network and scientific collaboration network, in which node labels correspond to geographical location of the nodes. Experiments including a comparison with randomized null models reveal how from the global level perspective, we observe synergistic contributions of orders in the board interlock network, whereas in the collaboration network there is more redundancy. The findings shed new light on social scientific debates on the role of busy directors in global business networks and the connective effects of large author teams in scientific collaboration networks.

physics.soc-ph

Holding Periods: Measuring the Inverse of Money Velocity from Transaction Records

This paper defines the average holding period of money and proposes a methodology for fully disaggregated measurement. Our measure is shown to be the inverse of the transfer velocity of money under stationary conditions, which is implicitly assumed in conventional aggregate measurement. Our methodology does not require stationarity. We leverage a recent computational technique to extract empirical holding periods from micro-level transaction data as recorded by real-world payment systems. This enables novel empirical analyses of money velocity under non-stationarity conditions. We illustrate several such analyses on Sarafu, a small digital community currency in Kenya, where transaction data is available from 25 January 2020 to 15 June 2021. Our measure implies faster circulation than does the aggregate transfer velocity; we can say that 58% of Sarafu was effectively static. We also disaggregate by geography to study the heterogeneous impact of economic disruptions related to the COVID-19 pandemic on Sarafu. Finally, we consider an ad-hoc currency management operation that took place in October 2020. Measuring the average holding period of money makes it possible to track money velocity during ongoing monetary interventions and on other important occasions when conditions are not stationary.

econ.GN

Fast maximal clique enumeration in weighted temporal networks

Cliques, groups of fully connected nodes in a network, are often used to study group dynamics of complex systems. In real-world settings, group dynamics often have a temporal component. For example, conference attendees moving from one group conversation to another. Recently, maximal clique enumeration methods have been introduced that add temporal (and frequency) constraints, to account for such phenomena. These methods enumerate so called (delta,gamma)-maximal cliques. In this work, we introduce an efficient (delta,gamma)-maximal clique enumeration algorithm, that extends gamma from a frequency constraint to a more versatile weighting constraint. Additionally, we introduce a definition of (delta,gamma)-cliques, that resolves a problem of existing definitions in the temporal domain. Our approach, which was inspired by a state-of-the-art two-phase approach, introduces a more efficient initial (stretching) phase. Specifically, we reduce the time complexity of this phase to be linear with respect to the number of temporal edges. Furthermore, we introduce a new approach to the second (bulking) phase, which allows us to efficiently prune search tree branches. Consequently, in experiments we observe speed-ups, often by several order of magnitude, on various (large) real-world datasets. Our algorithm vastly outperforms the existing state-of-the-art methods for temporal networks, while also extending applicability to weighted networks.

cs.SI

From Contact to Threat: A Social Network Perspective on Perceptions of Immigration

Our perceptions are shaped by the social networks we are embedded in. Despite the acknowledged influence of close contacts on how we perceive the world, the role of the broader social environment remains opaque. Here, we leverage a unique combination of population-scale social network and survey data on perceptions of immigration. We find that both direct contacts and a wider social network exposure to migrants matter. Notably, for natives, network exposure shows a shift from positive to negative association with perceptions of immigration beyond a certain exposure threshold. The multi-layer nature of our data highlights this tipping point for next-door neighbors, with private social contexts exhibiting a positive relationship between exposure and immigration perceptions. Furthermore, it shows that contacts spanning multiple contexts also strengthen this relationship. The provided insights on the interplay between network composition and attitudes toward immigration highlight generic patterns shaping public opinion on pressing societal issues.

physics.soc-ph

Connectivity and Community Structure of Online and Register-based Social Networks

The dominance of online social media data as a source of population-scale social network studies has recently been challenged by networks constructed from government-curated register data. In this paper, we investigate how the two compare, focusing on aggregations of the Dutch online social network (OSN) Hyves and a register-based social network (RSN) of the Netherlands. First and foremost, we find that the connectivity of the two population-scale networks is strikingly similar, especially between closeby municipalities, with more long-distance ties captured by the OSN. This result holds when correcting for population density and geographical distance, notwithstanding that these two patterns appear to be the main drivers of connectivity. Second, we show that the community structure of neither network follows strict administrative geographical delineations (e.g., provinces). Instead, communities appear to either center around large metropolitan areas or, outside of the country's most urbanized area, are comprised of large blocks of interdependent municipalities. Interestingly, beyond population and distance-related patterns, communities also highlight the persistence of deeply rooted historical and sociocultural communities based on religion. The results of this study suggest that both online social networks and register-based social networks are valuable resources for insights into the social network structure of an entire population.

physics.soc-ph

Early warning signals for predicting cryptomarket vendor success using dark net forum networks

In this work we focus on identifying key players in dark net cryptomarkets that facilitate online trade of illegal goods. Law enforcement aims to disrupt criminal activity conducted through these markets by targeting key players vital to the market's existence and success. We particularly focus on detecting successful vendors responsible for the majority of illegal trade. Our methodology aims to uncover whether the task of key player identification should center around plainly measuring user and forum activity, or that it requires leveraging specific patterns of user communication. We focus on a large-scale dataset from the Evolution cryptomarket, which we model as an evolving communication network. Results indicate that user and forum activity, measured through topic engagement, is best able to identify successful vendors. Interestingly, considering users with higher betweenness centrality in the communication network further improves performance, also identifying successful vendors with moderate activity on the forum. But more importantly, analyzing the forum data over time, we find evidence that attaining a high betweenness score comes before vendor success. This suggests that the proposed network-driven approach of modelling user communication might prove useful as an early warning signal for key player identification.

cs.SI