SearcharxivSearch

arXiv subjects

Jacopo Iacovacci

Publications and source records attributed to Jacopo Iacovacci.

14 recordsLinked to original sources

Evolution as a Process of Causal Inference

Recently, the mapping of the replicator equation onto Bayes' theorem has been recognised, leading to an analogy between evolutionary dynamics and Bayesian learning. However, this analogy holds only for pure selection in infinite populations and breaks down when mutations -- a central mechanism of evolution -- are introduced. Here I propose that evolution by natural selection, at least for populations of haploid replicators in static environments, is best understood not as a learning process but as a process of causal inference. Each mutation event constitutes a natural experiment in which the parent serves as the control and the mutant offspring as the treated unit. Natural selection screens the causal effect of the mutation on fitness, retaining mutations with non-negative effects. I formalise this view within the Neyman-Rubin potential-outcomes framework. I first develop the general theory using a generic fitness outcome and show how the core identification assumptions in causal inference (Stable Unit Treatment Value Assumption, Consistency, Unconfoundedness, Positivity) map onto evolutionary biology. Using the unnormalised quasispecies equation, I prove that the intergenerational change in mean fitness decomposes exactly into a selection term -- recovering Fisher's Fundamental Theorem -- plus a mutation term that corresponds to a fitness-weighted average of the cumulated effect of all mutations over all parental genotypes. I show that this decomposition extends, under suitable assumptions, to the generalised replicator-mutator equation and that the frequencies of populations of matched parents-offspring update in proportion to the average causal effect of mutations on fitness.

q-bio.PE

Extraction and integration of genetic networks from short-profile omic datasets

Mass-spectrometry technologies are widely used in the fields of ionomics and metabolomics to simultaneously profile at the genome scale intracellular concentrations of e.g. amino acids or elements. Short profiles of molecular or sub-molecular features are intrinsically non-Gaussian and may reveal patterns of correlations that reflect the system nature of the cell biochemistry and biology. Here we introduce two profile similarity measures that enforce information from the empirical covariance matrix of the data, the Mahalanobis cosine and the hybrid-Mahalanobis cosine. We evaluate the performance of these similarity measures in the task of inferring and integrating genetic networks from omics data by analysing experimental datasets derived from the ionome and the metabolome of the model organism S. cerevisiae, and several large curated databases of genetic annotations. The proposed covariance-based similarity measures can in general recover known and predicted associations between genes better than the commonly used Pearson's correlation and the standard cosine similarity. The choice of which of the two measures to recommend depends upon whether the focus is on extracting genetic associations at a global or local genetic network scale.

q-bio.MN

Visibility graphs for image processing

The family of image visibility graphs (IVGs) have been recently introduced as simple algorithms by which scalar fields can be mapped into graphs. Here we explore the usefulness of such operator in the scenario of image processing and image classification. We demonstrate that the link architecture of the image visibility graphs encapsulates relevant information on the structure of the images and we explore their potential as image filters and compressors. We introduce several graph features, including the novel concept of Visibility Patches, and show through several examples that these features are highly informative, computationally efficient and universally applicable for general pattern recognition and image classification tasks.

physics.data-an

A combinatorial framework to quantify peak/pit asymmetries in complex dynamics

We explore a combinatorial framework which efficiently quantifies the asymmetries between minima and maxima in local fluctuations of time series. We firstly showcase its performance by applying it to a battery of synthetic cases. We find rigorous results on some canonical dynamical models (stochastic processes with and without correlations, chaotic processes) complemented by extensive numerical simulations for a range of processes which indicate that the methodology correctly distinguishes different complex dynamics and outperforms state of the art metrics in several cases. Subsequently, we apply this methodology to real-world problems emerging across several disciplines including cases in neurobiology, finance and climate science. We conclude that differences between the statistics of local maxima and local minima in time series are highly informative of the complex underlying dynamics and a graph-theoretic extraction procedure allows to use these features for statistical learning purposes.

physics.data-an

Multilink Communities of Multiplex Networks

Multiplex networks describe a large number of complex social, biological and transportation networks where a set of nodes is connected by links of different nature and connotation. Here we uncover the rich community structure of multiplex networks by associating a community to each multilink where the multilinks characterize the connections existing between any two nodes of the multiplex network. Our community detection method reveals the rich interplay between the mesoscale structure of the multiplex networks and their multiplexity. For instance some nodes can belong to many layers and few communities while others can belong to few layers but many communities. Moreover the multilink communities can be formed by a different number of relevant layers. These results point out that mesoscopically there can be large differences in the compressibility of multiplex networks.

physics.soc-ph

Centralities of Nodes and Influences of Layers in Large Multiplex Networks

We formulate and propose an algorithm (MultiRank) for the ranking of nodes and layers in large multiplex networks. MultiRank takes into account the full multiplex network structure of the data and exploits the dual nature of the network in terms of nodes and layers. The proposed centrality of the layers (influences) and the centrality of the nodes are determined by a coupled set of equations. The basic idea consists in assigning more centrality to nodes that receive links from highly influential layers and from already central nodes. The layers are more influential if highly central nodes are active in them. The algorithm applies to directed/undirected as well as to weighted/unweighted multiplex networks. We discuss the application of MultiRank to three major examples of multiplex network datasets: the European Air Transportation Multiplex Network, the Pierre Auger Multiplex Collaboration Network and the FAO Multiplex Trade Network.

physics.soc-ph

Visibility graphs of random scalar fields and spatial data

The family of visibility algorithms were recently introduced as mappings between time series and graphs. Here we extend this method to characterize spatially extended data structures by mapping scalar fields of arbitrary dimension into graphs. After introducing several possible extensions, we provide analytical results on some topological properties of these graphs associated to some types of real-valued matrices, which can be understood as the high and low disorder limits of real-valued scalar fields. In particular, we find a closed expression for the degree distribution of these graphs associated to uncorrelated random fields of generic dimension, extending a well known result in one-dimensional time series. As this result holds independently of the field's marginal distribution, we show that it directly yields a statistical randomness test, applicable in any dimension. We showcase its usefulness by discriminating spatial snapshots of two-dimensional white noise from snapshots of a two-dimensional lattice of diffusively coupled chaotic maps, a system that generates high dimensional spatio-temporal chaos. We finally discuss the range of potential applications of this combinatorial framework, which include image processing in engineering, the description of surface growth in material science, soft matter or medicine and the characterization of potential energy surfaces in chemistry, disordered systems and high energy physics. An illustration on the applicability of this method for the classification of the different stages involved in carcinogenesis is briefly discussed.

physics.data-an

Functional Multiplex PageRank

Recently it has been recognized that many complex social, technological and biological networks have a multilayer nature and can be described by multiplex networks. Multiplex networks are formed by a set of nodes connected by links having different connotations forming the different layers of the multiplex. Characterizing the centrality of the nodes in a multiplex network is a challenging task since the centrality of the node naturally depends on the importance associated to links of a certain type. Here we propose to assign to each node of a multiplex network a centrality called Functional Multiplex PageRank that is a function of the weights given to every different pattern of connections (multilinks) existent in the multiplex network between any two nodes. Since multilinks distinguish all the possible ways in which the links in different layers can overlap, the Functional Multiplex PageRank can describe important non-linear effects when large relevance or small relevance is assigned to multilinks with overlap. Here we apply the Functional Page Rank to the multiplex airport networks, to the neuronal network of the nematode c.elegans, and to social collaboration and citation networks between scientists. This analysis reveals important differences existing between the most central nodes of these networks, and the correlations between their so called "pattern to success".

physics.soc-ph

Extracting Information from Multiplex Networks

Multiplex networks are generalized network structures that are able to describe networks in which the same set of nodes are connected by links that have different connotations. Multiplex networks are ubiquitous since they describe social, financial, engineering and biological networks as well. Extending our ability to analyze complex networks to multiplex network structures increases greatly the level of information that is possible to extract from Big Data. For these reasons characterizing the centrality of nodes in multiplex networks and finding new ways to solve challenging inference problems defined on multiplex networks are fundamental questions of network science. In this paper we discuss the relevance of the Multiplex PageRank algorithm for measuring the centrality of nodes in multilayer networks and we characterize the utility of the recently introduced indicator function $\widetildeΘ^{S}$ for describing their mesoscale organization and community structure. As working examples for studying these measures we consider three multiplex network datasets coming for social science.

physics.soc-ph

Sequential motif profile of natural visibility graphs

The concept of sequential visibility graph motifs -subgraphs appearing with characteristic frequencies in the visibility graphs associated to time series- has been advanced recently along with a theoretical framework to compute analytically the motif profiles associated to Horizontal Visibility Graphs (HVGs). Here we develop a theory to compute the profile of sequential visibility graph motifs in the context of Natural Visibility Graphs (VGs). This theory gives exact results for deterministic aperiodic processes with a smooth invariant density or stochastic processes that fulfil the Markov property and have a continuous marginal distribution. The framework also allows for a linear time numerical estimation in the case of empirical time series. A comparison between the HVG and the VG case (including evaluation of their robustness for short series polluted with measurement noise) is also presented.

physics.data-an

Sequential visibility-graph motifs

Visibility algorithms transform time series into graphs and encode dynamical information in their topology, paving the way for graph-theoretical time series analysis as well as building a bridge between nonlinear dynamics and network science. In this work we introduce and study the concept of sequential visibility graph motifs, smaller substructures of n consecutive nodes that appear with characteristic frequencies. We develop a theory to compute in an exact way the motif profiles associated to general classes of deterministic and stochastic dynamics. We find that this simple property is indeed a highly informative and computationally efficient feature capable to distinguish among different dynamics and robust against noise contamination. We finally confirm that it can be used in practice to perform unsupervised learning, by extracting motif profiles from experimental heart-rate series and being able, accordingly, to disentangle meditative from other relaxation states. Applications of this general theory include the automatic classification and description of physical, biological, and financial time series.

physics.data-an

Emergence of multiplex communities in collaboration networks

Community structures in collaboration networks reflect the natural tendency of individuals to organize their work in groups in order to better achieve common goals. In most of the cases, individuals exploit their connections to introduce themselves to new areas of interests, giving rise to multifaceted collaborations which span different fields. In this paper, we analyse collaborations in science and among movie actors as multiplex networks, where the layers represent respectively research topics and movie genres, and we show that communities indeed coexist and overlap at the different layers of such systems. We then propose a model to grow multiplex networks based on two mechanisms of intra and inter-layer triadic closure which mimic the real processes by which collaborations evolve. We show that our model is able to explain the multiplex community structure observed empirically, and we infer the strength of the two underlying social mechanisms from real-world systems. Being also able to correctly reproduce the values of intra-layer and inter-layer assortativity correlations, the model contributes to a better understanding of the principles driving the evolution of social networks.

physics.soc-ph

Mesoscopic Structures Reveal the Network Between the Layers of Multiplex Datasets

Multiplex networks describe a large variety of complex systems, whose elements (nodes) can be connected by different types of interactions forming different layers (networks) of the multiplex. Multiplex networks include social networks, transportation networks or biological networks in the cell or in the brain. Extracting relevant information from these networks is of crucial importance for solving challenging inference problems and for characterizing the multiplex networks microscopic and mesoscopic structure. Here we propose an information theory method to extract the network between the layers of multiplex datasets, forming a "network of networks". We build an indicator function, based on the entropy of network ensembles, to characterize the mesoscopic similarities between the layers of a multiplex network and we use clustering techniques to characterize the communities present in this network of networks. We apply the proposed method to study the Multiplex Collaboration Network formed by scientists collaborating on different subjects and publishing in the Americal Physical Society (APS) journals. The analysis of this dataset reveals the interplay between the collaboration networks and the organization of knowledge in physics.

physics.soc-ph

Triadic closure as a basic generating mechanism of communities in complex networks

Most of the complex social, technological and biological networks have a significant community structure. Therefore the community structure of complex networks has to be considered as a universal property, together with the much explored small-world and scale-free properties of these networks. Despite the large interest in characterizing the community structures of real networks, not enough attention has been devoted to the detection of universal mechanisms able to spontaneously generate networks with communities. Triadic closure is a natural mechanism to make new connections, especially in social networks. Here we show that models of network growth based on simple triadic closure naturally lead to the emergence of community structure, together with fat-tailed distributions of node degree, high clustering coefficients. Communities emerge from the initial stochastic heterogeneity in the concentration of links, followed by a cycle of growth and fragmentation. Communities are the more pronounced, the sparser the graph, and disappear for high values of link density and randomness in the attachment procedure. By introducing a fitness-based link attractivity for the nodes, we find a novel phase transition, where communities disappear for high heterogeneity of the fitness distribution, but a new mesoscopic organization of the nodes emerges, with groups of nodes being shared between just a few superhubs, which attract most of the links of the system.

physics.soc-ph