SearcharxivSearch

arXiv subjects

Manuel Suarez-Roman

Publications and source records attributed to Manuel Suarez-Roman.

3 recordsLinked to original sources

Self-Similarity in Online Networks During Social Movements

Online platforms provide an infrastructure for social movements, leaving digital traces that can be modelled as networks to quantify how information, participation, and coordination emerge during episodes of collective action and evolve over time. In this work, we unveil the emergence of scale-invariant online interaction patterns in social movements through network analysis of three geographically and sociopolitically distinct massive mobilisation events. By constructing co-occurrence networks from Twitter (now X) hashtag data and applying a degree-thresholding renormalisation procedure, we demonstrate that these highly correlated social phenomena exhibit clear signatures of self-similarity at peak mobilisation times. These critical points are characterised by modular-to-nested transitions, both in the co-occurrence networks and the bi-partite ones, maxima in user participation, and clustering spectrum collapse across multiple network scales. Despite their geographical and sociopolitical diversity, all three movements display remarkably analogous self-similar properties. Furthermore, the results hint at the emergence of a latent metric structure that supports successful hyperbolic embedding, providing an estimate of effective social distance. Together, these findings suggest that self-similarity may constitute a universal organising principle of social movements during peak mobilisation phases, as it lays the groundwork for the rapid amplification of information across scales that is necessary for the successful coordination of collective action.

physics.soc-ph

The CTI Echo Chamber: Fragmentation, Overlap, and Vendor Specificity in Twenty Years of Cyber Threat Reporting

Despite the high volume of open-source Cyber Threat Intelligence (CTI), our understanding of long-term threat actor-victim dynamics remains fragmented due to inconsistent reporting standards and the lack of structured datasets containing comprehensive analytic information. In this paper, we present a large-scale automated analysis of open-source CTI reports spanning two decades. We develop a high-precision, LLM-based pipeline to ingest and structure 16,096 reports, extracting key entities such as attributed threat actors, motivations, victims, reporting vendors, and technical indicators (IoCs and TTPs). Our analysis quantifies the evolution of CTI information density and specialization, characterizing patterns that relate specific threat actors to motivations and victim profiles. Furthermore, we perform a meta-analysis of the CTI industry itself. We identify a fragmented ecosystem of distinct silos where vendors demonstrate significant geographic and sectoral reporting biases. Our marginal coverage analysis reveals that intelligence overlap between vendors is typically low: while a few core providers may offer broad situational awareness, additional sources yield diminishing returns. Overall, our findings characterize the structural biases inherent in the CTI ecosystem, enabling practitioners and researchers to better evaluate the completeness of their intelligence sources.

cs.CR

Hesperus is Phosphorus: Mapping Threat Actor Naming Taxonomies at Scale

This paper studies the problem of Threat Actor (TA) naming convention inconsistency across leading Cyber Threat Intelligence (CTI) vendors. The current decentralized and proprietary nomenclature creates confusion and significant obstacles for researchers, including difficulties in integrating and correlating disparate CTI reports and TA profiles. This paper introduces HiP (Hesperus is Phosphorus, a reference to the classic question about the Morning and the Evening Star), a methodology for normalizing, integrating, and clustering TA names presumably corresponding to the same entity. Using HiP, we analyze a large dataset collected from 15 sources and spanning 13,371 CTI reports, 17 vendor taxonomies, 3,287 TA names, and 8 mappings between them. Our analysis of the resulting name graph provides insights on key features of the problem, such as the concentration of aliases on a relatively small subset of TAs, the evolution of this phenomenon over the years, and the factors that could explain TA name proliferation. We also report errors in the mappings and methodological pitfalls that contribute to make certain TA name clusters larger than they should be, including the use of temporary names for activity clusters, the existence of common tools and infrastructure, and overlapping operations. We conclude with a discussion on the inherent difficulties to adopt a TA naming standard, a quest fundamentally hampered by the need to share highly-sensitive telemetry that is private to each CTI vendor.

cs.CR