SearcharxivSearch

arXiv subjects

Diego Garlaschelli

Publications and source records attributed to Diego Garlaschelli.

At least 19 recordsLinked to original sources

Spectra of random graphs with discrete scale invariance

Random graphs defined by an occurrence probability that is invariant under node aggregation have been identified recently in the context of network renormalization. The invariance property requires that edges are drawn with a specific probability that, in the annealed case, depends on a necessarily infinite-mean node fitness. The diverging mean determines many properties that are uncommon in models with independent edges, but at the same time widespread in real-world networks. Here we focus on the leading eigenvalues and eigenvectors of the adjacency matrix of the model, where the n nodes are assigned a Pareto($α$)-distributed fitness with 0 < $α$ < 1. We find that the leading eigenvalues are all of order square root of n, alternate in sign and are located at the intersection between the real axis and a logarithmic spiral in the complex plane, which we characterize analytically in terms of the Gamma function. We also calculate the associated eigenvectors, finding that they display complexvalued scaling exponents and log-periodicity, which are signatures of discrete scale invariance. In contrast with the typical finite-rank behaviour of random graphs with finite-mean variables, we find that a growing number of the leading eigenvalues emerges from the bulk, whose edge extends up to order square root of n and therefore reaches the same scale as that of the structural eigenvalues.

math.SP

Inverse generalised spin models of answers to questionnaires

Network psychometrics conceptualises psychological constructs as emergent properties of systems of interacting items. Energy-based probabilistic models have gained popularity as models of these interactions, but their psychometric application has so far been limited to binary responses, bilinear interactions, and approximated inference methods. To fill these gaps, we here infer and analyse three generalized-spin models of ordinal questionnaire data: the Ising, Blume-Capel (BC), and Blume-Emery-Griffiths (BEG) models. These are maximum-entropy models that accommodate ordinal responses on Likert-type scales with an arbitrary number of options, allowing for single-site anisotropy (BC, BEG) and bi-quadratic item interactions (BEG). We prove the concavity of the maximum likelihood estimation of their parameters, as well as the gauge invariance of the Ising and BC models. We introduce a stochastic gradient ascent algorithm for maximum likelihood inference, and apply this procedure to eleven psychometric and sociological questionnaire datasets. Evaluating the predictive ability of the inferred models reveals that the BEG model systematically outperforms Factor Analysis and the other spin models in capturing the distributions of factors and distances of subject answers to the mean, across all datasets. By leveraging a competitive interplay between the quadratic and bi-quadratic energy terms, the BEG model uniquely captures individual average-extremist response styles, alongside standard latent factor positioning. Moreover, only the spin models can account for the non-concavity and multi-modality of factor histograms in the most polarizing questionnaires. Finally, the analysis reveals other highly non-linear traits of ordinal data---such as the fat-tailed distribution of Mahalanobis distances to the mean---that escape satisfactory description by both factor and spin models.

physics.data-an

Reproducing the first and second moments of empirical degree distributions

The study of probabilistic models for the analysis of complex networks represents a flourishing research field. Among the former, Exponential Random Graphs (ERGs) have gained increasing attention over the years. So far, only linear ERGs have been extensively employed to gain insight into the structural organisation of real-world complex networks. None, however, is capable of accounting for the variance of the empirical degree distribution. To this aim, non-linear ERGs must be considered. After showing that the usual mean-field approximation forces the degree-corrected version of the two-star model to degenerate, we define a fitness-induced variant of it. Such a `softened' model is capable of reproducing the sample variance, while retaining the explanatory power of its linear counterpart, within a purely canonical framework.

physics.soc-ph

Renormalisation of Inhomogeneous Random Graphs

We consider inhomogeneous random graphs in which vertices are assigned i.i.d.\ random weights, pairs of distinct vertices are connected by an edge independently with a probability that is a bi-variate function of the weights of the vertices, and single vertices are connected to themselves by a self-loop independently with a probability that is a uni-variate function of the weight of the vertex. We apply a renormalisation transformation in which vertices are aggregated into groups of equal size according to a greedy algorithm, namely, distinct groups of aggregated vertices are connected by an aggregated edge if and only if there is at least one edge connecting two constituent vertices across the groups, while a group of aggregated vertices is connected to itself by an aggregated self-loop if and only if there is at least one self-loop at an internal vertex or one edge connecting a pair of distinct internal vertices. We analyse what happens when the renormalisation transformation is iterated. In particular, we show that, starting from appropriately scaled connection functions, the iterated renormalised graphs converge to a two-parameter family of random graphs, acting as an attractor in a universality class. We consider a light-tailed regime, for which the scaling limit is a homogeneous Erdős--Rényi random graph, and a heavy-tailed regime, for which the scaling limit is an inhomogeneous random graph with stable infinite-mean random weights and an exponential disconnection function. Different scalings are needed for the two regimes. Which of the two regimes prevails depends on the choice of the connection functions and the choice of the law of the random weights.

math.PR

Analysis of a maximum-entropy based estimator for dynamic random graph models

We study dynamic random graphs in which the set of nodes is fixed, but edges evolve over time according to an underlying stochastic mechanism. Using a maximum-entropy approach, we define a probability distribution on graph trajectories that is consistent with observed constraints, capturing the inherent uncertainty in partially observed networks. We introduce a moment-based estimator for the parameters of this distribution and establish its statistical properties, such as consistency and asymptotic normality, with explicit formulas for the covariance structure. Numerical experiments demonstrate the estimator's accuracy and robustness across various dynamic network scenarios. Our framework bridges probabilistic modeling and statistical inference in time-varying networks, providing practical tools for understanding and predicting complex edge dynamics.

math.ST

q-Exponential Random Graphs: higher-order networks from simple constraints

Exponential Random Graphs (ERGs) are among the most widely used network models, derived as principled least-bias graph ensembles that maximize Shannon entropy under constraints on the expected values of given structural properties. However, it has been recently (re)discovered that, in the absence of additional information privileging Shannon entropy, the most agnostic inferential construction should maximize the broader class of Uffink entropies. The resulting entropy-maximizing distribution changes from the exponential (Boltzmann-Gibbs) to the so-called q-exponential one. Since maximizing Shannon entropy may produce an unjustified independence between degrees of freedom, here we investigate how the most popular ERGs with independent edges (namely, the Erdos-Renyi and configuration models) generalize to higher-order q-Exponential Random Graphs with dependent edges in the non-Shannon case, while keeping their defining constraints (number of links and degree sequence, respectively) unchanged. We find features, such as a phase transition between sparse and dense regimes, that are absent in the original ERGs but typical of higher-order networks, plus novel phenomena such as richer assortativity and clustering profiles, which allow for the coexistence of link sparsity and triadic closure. These results show that higher-order networks do not necessarily require higher-order constraints, as they naturally arise from simpler ones in a framework that is even more agnostic than Shannon's.

physics.soc-ph

Community detection in subject-subject networks from psychometrics data

Identifying subgroups of respondents in psychometric data is traditionally addressed with Latent Class Analysis, which requires the number of classes to be specified a priori and can perform poorly when strong inter-item correlations violate local independence assumptions. We propose a network-theoretic alternative based on community detection in subject-subject similarity networks. To suppress the systematic artifacts induced by the factor structure of the items, the similarity is computed in a low-dimensional factor-score space and the null model for modularity maximisation is obtained by removing the leading (global) mode of the similarity matrix, rather than via the standard Newman--Girvan model. The significance of a detected partition is then assessed against a column-wise resampling null through four complementary observables: the modularity, the differential entropy of the eigenvector point cloud at two neighbourhood scales, and the overlap of the within- and between-community similarity histograms. On a synthetic benchmark with controlled mixture signal, all four metrics correctly identify the homogeneous case as null-compatible -- including the demanding regime of a dataset dominated by a single factor -- and exhibit a graded departure from the null as the cluster separation grows. Applied to 14 widely used psychometric scales, the pipeline isolates a small group of datasets supporting a genuine and directly interpretable modular structure, while the remaining scales fall either in a mixed-signal regime or in one compatible with a single homogeneous community. The significance analysis is independent of the specific community-detection algorithm and provides an operational way to test for modular subject-level structure in questionnaire data.

physics.soc-ph

GDP-Driven Structural and Dynamical Heterogeneity in the Synchronization of Chaotic Macroeconomic Networks

We investigate the emergence of synchronization in a network of coupled chaotic macroeconomic systems. Each node represents an economy characterized by three key variables savings, gross domestic product (GDP), and foreign capital inflows. These economies interact or are connected through a fitness-based probability that depends on the potential GDP of each node. This formulation allows both structural heterogeneity, arising from uneven network connectivity, and dynamical heterogeneity, due to differences in local parameters, to be explored within a unified framework. Using both numerical simulations and a mean-field approximation, by varying the coupling strength and the degree of heterogeneity of both network topology and dynamical behavior of the nodes, we analyze synchronization transitions. Our results show that the mean-field approach accurately captures the collective dynamics in homogeneous and fully connected networks even with heterogeneity within the intrinsic dynamic of the nodes but fails when strong heterogeneity in the structure of the network is introduced. In heterogeneous networks, the system exhibits partial synchronization and on--off intermittency, where coherent phases of global synchronization alternate with abrupt desynchronization bursts. The distribution of laminar phase durations follows a power-law scaling, consistent with theoretical predictions for intermittent synchronization. From an economic perspective, these results suggest that global business cycle synchronization is inherently fragile: strong integration can promote temporary coordination among economies, but structural and dynamical disparities inevitably lead to intermittent breakdowns of collective behavior.

nlin.AO

Twitter climate discourse as a signal of pro-environmental behaviors

Fostering coordinated pro-environmental behaviors at scale is a key challenge for climate mitigation. Individual actions only generate meaningful impact when they diffuse widely and become socially coordinated, yet monitoring such processes remains difficult with traditional survey-based tools alone. In this study, we examine whether large-scale online climate discourse is associated with differences in offline pro-environmental behavior across European regions. We combine geolocated Twitter data from the Climate Change Twitter Dataset (2017-2019) with survey-based measures from the 2019 Special Eurobarometer, focusing on the regional density of climate-related tweets and the average number of self-reported pro-environmental actions. We find a strong positive association between tweet density and pro-environmental behavior that remains robust to socio-economic controls, alternative spatial aggregations, and a wide range of robustness checks. To move beyond aggregate volume, we further decompose online discourse using Natural Language Processing tools that capture distinct social dimensions. While knowledge exchange shows no clear relationship with offline behavior, the prevalence of activism- and social support-related expressions is negatively associated with pro-environmental actions. Overall, our results suggest that online climate discourse can serve as an informative, attention-related signal of regional differences in pro-environmental behavior, but that different forms of online engagement relate to offline action in markedly different ways. More broadly, the study highlights the potential of integrating large-scale digital traces with survey data to investigate collective behavior in socio-environmental systems, while remaining explicitly observational in scope.

cs.SI

Network Reconstruction via Jeffreys Prior under Missing Sufficient Statistics

The modeling and reconstruction of economic networks from aggregate information has important implications for counterfactual analysis and policymaking. The traditional Fitness Model (FM) achieves good performance by using node-specific variables that are easily accessible (e.g., GDP for countries or total assets for banks or firms) and the overall link density as the only sufficient statistic. However, it often ignores additional contextual or mesoscopic features which may be more difficult to observe. In this paper, we extend the framework by incorporating block structure as in the Fitness-Corrected Block Model (FCBM), which allows for heterogeneous densities within and across blocks, but in the more challenging setting where such block-specific densities are not empirically available. Our method compensates for the absence of empirical information about the sufficient statistics by using a Jeffreys prior to average, in the most unbiased way, over all compatible solutions that are otherwise left unidentified. We evaluate the method on three international trade datasets across different product classes, including fresh products, common products, geographically specific products, and high-technology products. The underlying block structure is represented by economic regions as defined by the World Bank, and we only assume empirical knowledge of the total GDPs and the overall link density. The new method systematically outperforms the baseline Block-Agnostic FM (which uses the same input information) and sometimes even the FCBM (despite the latter uses more information), thereby suggesting reduced overfitting risk.

physics.soc-ph

Clustering without geometry in sparse networks with independent edges

The coexistence of sparsity and clustering (non-vanishing average fraction of triangles per node) is one of the few structural features that, irrespective of finer details, are ubiquitously observed across large real-world networks. This fact calls for generic models producing sparse clustered graphs. Earlier results suggested that sparse random graphs with independent edges fail to reproduce clustering, unless edge probabilities are assumed to depend on underlying metric distances that, thanks to the triangle inequality, naturally favour triadic closure. This observation has opened a debate on whether clustering implies (latent) geometry in real-world networks. Alternatively, recent models of higher-order networks can replicate clustering by abandoning edge independence. In this paper, we mathematically prove, and numerically confirm, that a sparse random graph with independent edges, recently identified in the context of network renormalization as an invariant model under node aggregation, produces finite clustering without any geometric or higher-order constraint. The underlying mechanism is an infinite-mean node fitness, which also implies a power-law degree distribution. Further, as a novel phenomenon that we characterize rigorously, we observe the breakdown of self-averaging of various network properties. Therefore, as an alternative to geometry or higher-order dependencies, node aggregation invariance emerges as a basic route to realistic network properties.

math.PR

Renormalization of Interacting Random Graph Models

Random graphs offer a useful mathematical representation of a variety of real world complex networks. Exponential random graphs, for example, are particularly suited towards generating random graphs constrained to have specified statistical moments. In this investigation, we elaborate on a generalization of the former where link probabilities are conditioned on the appearance of other links, corresponding to the introduction of interactions in an effective generalized statistical mechanical formalism. When restricted to the simplest non-trivial case of pairwise interactions, one can derive a closed form renormalization group transformation for maximum coordination number two on the corresponding line graph. Higher coordination numbers do not admit exact closed form renormalization group transformations, a feature that paraphrases the usual absence of exact transformations in two or more dimensional lattice systems. We introduce disorder and study the induced renormalization group flow on its probability assignments, highlighting its formal equivalence to time reversed anisotropic drift-diffusion on the statistical manifold associated with the effective Hamiltonian. We discuss the implications of our findings, stressing the long wavelength irrelevance of certain classes of pair-wise conditioning on random graphs, and conclude with possible applications. These include modeling the scaling behavior of preferential effects on social networks, opinion dynamics, and reinforcement effects on neural networks, as well as how our findings offer a systematic framework to deal with data limitations in inference and reconstruction problems.

cond-mat.stat-mech

Multi-Scale Node Embeddings for Graph Modeling and Generation

Lying at the interface between Network Science and Machine Learning, node embedding algorithms take a graph as input and encode its structure onto output vectors that represent nodes in an abstract geometric space, enabling various vector-based downstream tasks such as network modelling, data compression, link prediction, and community detection. Two apparently unrelated limitations affect these algorithms. On one hand, it is not clear what the basic operation defining vector spaces, i.e. the vector sum, corresponds to in terms of the original nodes in the network. On the other hand, while the same input network can be represented at multiple levels of resolution by coarse-graining the constituent nodes into arbitrary block-nodes, the relationship between node embeddings obtained at different hierarchical levels is not understood. Here, building on recent results in network renormalization theory, we address these two limitations at once and define a multiscale node embedding method that, upon arbitrary coarse-grainings, ensures statistical consistency of the embedding vector of a block-node with the sum of the embedding vectors of its constituent nodes. We illustrate the power of this approach on two economic networks that can be naturally represented at multiple resolution levels: namely, the international trade between (sets of) countries and the input-output flows among (sets of) industries in the Netherlands. We confirm the statistical consistency between networks retrieved from coarse-grained node vectors and networks retrieved from sums of fine-grained node vectors, a result that cannot be achieved by alternative methods. Several key network properties, including a large number of triangles, are successfully replicated already from embeddings of very low dimensionality, allowing for the generation of faithful replicas of the original networks at arbitrary resolution levels.

physics.soc-ph

Testing maximum entropy models with e-values

E-values have recently emerged as a robust and flexible alternative to p-values for hypothesis testing, especially under optional continuation, i.e., when additional data from further experiments are collected. In this work, we define optimal e-values for testing between maximum entropy models, both in the microcanonical (hard constraints) and canonical (soft constraints) settings. We show that, when testing between two hypotheses that are both microcanonical, the so-called growth-rate optimal e-variable admits an exact analytical expression, which also serves as a valid e-variable in the canonical case. For canonical tests, where exact solutions are typically unavailable, we introduce a microcanonical approximation and verify its excellent performance via both theoretical arguments and numerical simulations. We then consider constrained binary models, focusing on $2 \times k$ contingency tables -- an essential framework in statistics and a natural representation for various models of complex systems. Our microcanonical optimal e-variable performs well in both settings, constituting a new tool that remains effective even in the challenging case when the number $k$ of groups grows with the sample size, as in models with growing features used for the analysis of real-world heterogeneous networks and time-series.

stat.ME

Renormalizable Graph Embeddings For Multi-Scale Network Reconstruction

In machine learning, graph embedding algorithms seek low-dimensional representations of the input network data, thereby allowing for downstream tasks on compressed encodings. Recently, within the framework of network renormalization, multi-scale embeddings that remain consistent under an arbitrary aggregation of nodes onto block-nodes, and consequently under an arbitrary change of resolution of the input network data, have been proposed. Here we investigate such multi-scale graph embeddings in the modified context where the input network is not entirely observable, due to data limitations or privacy constraints. This situation is typical for financial and economic networks, where connections between individual banks or firms are hidden due to confidentiality, and one has to probabilistically reconstruct the underlying network from aggregate information. We first consider state-of-the-art network reconstruction techniques based on the maximum-entropy principle, which is designed to operate optimally at a fixed resolution level. We then discuss the limitations of these methods when they are used as graph embeddings to yield predictions across different resolution levels. Finally, we propose their natural 'renormalizable' counterparts derived from the distinct principle of scale invariance, yielding consistent graph embeddings for multi-scale network reconstruction. We illustrate these methods on national economic input-output networks and on international trade networks, which can be naturally represented at multiple levels of industrial and geographic resolution, respectively.

physics.soc-ph

Description length of canonical and microcanonical models

The (non-)equivalence of canonical and microcanonical ensembles is a fundamental question in statistical physics, concerning whether the use of soft and hard constraints in the maximum-entropy construction leads to the same description of a system. Despite the fact that maximum-entropy models are also commonly used in statistical inference, pattern detection, and hypothesis testing, a complete understanding of the effects of ensemble non-equivalence on statistical modeling is still missing. Here, we study this problem from a rigorous model selection perspective by comparing canonical and microcanonical models via the Minimum Description Length (MDL) principle, which yields a trade-off between likelihood, measuring model accuracy, and complexity, measuring model flexibility and its potential to overfit data. We compute the Normalized Maximum Likelihood (NML) of both formulations and find that: (i) microcanonical models always achieve higher likelihood but are always more complex; (ii) the optimal model choice depends on the empirical values of the constraints -- the canonical model performs best when its fit to the observed data exceeds its uniform average fit across all realizations; (iii) in the thermodynamic limit, the difference in description length per node vanishes when ensemble equivalence holds but persists otherwise, showing that non-equivalence implies extensive differences between large canonical and microcanonical models. Finally, we compare the NML approach to Bayesian methods, showing that (iv) the choice of priors, practically irrelevant in equivalent models, becomes crucial when an extensive number of constraints is enforced, possibly leading to very different outcomes.

cond-mat.stat-mech

Quantum generative modeling for financial time series with temporal correlations

Quantum generative adversarial networks (QGANs) have been investigated as a method for generating synthetic data with the goal of augmenting training data sets for neural networks. This is especially relevant for financial time series, since we only ever observe one realization of the process, namely the historical evolution of the market, which is further limited by data availability and the age of the market. However, for classical generative adversarial networks it has been shown that generated data may (often) not exhibit desired properties (also called stylized facts), such as matching a certain distribution or showing specific temporal correlations. Here, we investigate whether quantum correlations in quantum inspired models of QGANs can help in the generation of financial time series. We train QGANs, composed of a quantum generator and a classical discriminator, and investigate two approaches for simulating the quantum generator: a full simulation of the quantum circuits, and an approximate simulation using tensor network methods. We tested how the choice of hyperparameters, such as the circuit depth and bond dimensions, influenced the quality of the generated time series. The QGAN that we trained generate synthetic financial time series that not only match the target distribution but also exhibit the desired temporal correlations, with the quality of each property depending on the hyperparameters and simulation method.

quant-ph

Introduction to correlation networks: Interdisciplinary approaches beyond thresholding

Many empirical networks originate from correlational data, arising in domains as diverse as psychology, neuroscience, genomics, microbiology, finance, and climate science. Specialized algorithms and theory have been developed in different application domains for working with such networks, as well as in statistics, network science, and computer science, often with limited communication between practitioners in different fields. This leaves significant room for cross-pollination across disciplines. A central challenge is that it is not always clear how to best transform correlation matrix data into networks for the application at hand, and probably the most widespread method, i.e., thresholding on the correlation value to create either unweighted or weighted networks, suffers from multiple problems. In this article, we review various methods of constructing and analyzing correlation networks, ranging from thresholding and its improvements to weighted networks, regularization, dynamic correlation networks, threshold-free approaches, comparison with null models, and more. Finally, we propose and discuss recommended practices and a variety of key open questions currently confronting this field.

physics.soc-ph