SearcharxivSearch

arXiv subjects

Andrea Mazzolini

Publications and source records attributed to Andrea Mazzolini.

12 recordsLinked to original sources

Alternative routes to universal diversity scaling in component systems: from proteomes to large language models

Remarkably common statistical laws characterize the diversity scaling and its fluctuations across a wide range of complex "component systems". These regularities are often interpreted as signatures of an underlying innovation mechanism driving the growth of component diversity, but the basic ingredients necessary for their emergence remain poorly understood. In particular, from language and technological artifacts to genomes and gene expression patterns, the number of distinct components grows sublinearly with system size, while its variance scales approximately as the square of its mean. This behavior is consistent across diverse systems, raising the question of whether general constraints or emergent principles underlying diversity and innovation define the architectures of realizations with different numbers of components. To address this question, we derive analytical conditions for the joint emergence of these two diversity laws within a broad class of growth models, showing that they require a specific asymptotic dependence of the innovation probability on diversity and system size. We then demonstrate that the same macroscopic laws arise in a different class of models with latent heterogeneity, where quadratic fluctuation scaling always emerges asymptotically as a consequence of general statistical principles, essentially the law of total variance, without explicitly assuming an innovation mechanism or any specific rule for system assembly. We compare these predictions with empirical data from language, genomes, LEGO constructions, and texts generated by large language models. Our results show that empirical diversity scaling laws strongly constrain generative models but do not uniquely identify the mechanisms generating diversity, revealing a close correspondence between innovation-driven growth models and latent-variable descriptions.

cond-mat.stat-mech

Component systems: do null models explain everything?

Component systems - ensembles of realizations built from a shared repertoire of modular parts - are ubiquitous in biological, ecological, technological, and socio-cultural domains. From genomes to texts, cities, and software, these systems exhibit statistical regularities that often meet the "bona fide" requirements of laws in the physical sciences. Here, we argue that the generality and simplicity of those laws are often due to basic combinatorial or sampling constraints, raising the question of whether such patterns are actually revealing system-specific mechanisms and how we might move beyond them. To this end, we first present a unifying mathematical framework, which allows us to compare modular systems in different fields and highlights the common "null" trends as well as the system-specific uniqueness, which, arguably, are signatures of the underlying generative dynamics. Next, we can exploit the framework with statistical mechanics and modern machine-learning tools for a twofold objective. (i) Explaining why the general regularities emerge, highlighting the constraints between them and the general principles at their origins, and (ii) "subtracting" them from data, which will isolate the informative features for inferring hidden system-specific generative processes, mechanistic and causal aspects.

cond-mat.stat-mech

Dynamics of memory B cells and plasmablasts in healthy individuals

Our adaptive immune system relies on the persistence over long times of a diverse set of antigen-experienced B cells to encode our memories of past infections and to protect us against future ones. While longitudinal repertoire sequencing promises to track the long-term dynamics of many B cell clones simultaneously, sampling and experimental noise make it hard to draw reliable quantitative conclusions. Leveraging statistical inference, we infer the dynamics of memory B cell clonal dynamics and conversion to plasmablasts, which includes clone creation, degradation, abundance fluctuations, and differentiation. We find that memory B cell clones degrade slowly, with a half-life of 10 years. Based on the inferred parameters, we predict that it takes about 50 years to renew 50\% of the repertoire, with most observed clones surviving for a lifetime. We infer that, on average, 1 out of 100 memory B cells differentiates into a plasmablast each year, more than expected from purely antigen-stimulated differentiation, and that plasmablast clones degrade with a half-life of about one year in the absence of memory imports. Our method is general and could be applied to other longitudinal repertoire sequencing B cell subsets.

q-bio.PE

Ranking nodes in bipartite systems with a non-linear iterative map

Ranking nodes in networks according to a defined measure of importance is an extensively studied task, with applications in ecology, economic trade networks, and social networks. This paper introduces a method based on a non-linear iterative map to evaluate node relevance in bipartite networks. By tuning a single parameter $\gamma$, the method captures different concepts of node importance, including established measures like degree centrality, eigenvector centrality and the fitness-complexity ranking. The algorithm's flexibility allows for efficient ranking optimization tailored to specific tasks, outperforming state-of-the-art algorithms. We apply this method to ecological mutualistic networks, where ranking quality can be assessed by the extinction area - the rate at which the system collapses when species are removed in a certain order. The map with the optimal $\gamma$ value surpasses existing ranking methods on this task. Additionally, our method excels in evaluating nestedness, another crucial structural property of ecological systems, requiring specific node rankings. Finally, we explore theoretical aspects of the map, revealing a phase transition at a critical $\gamma$ dependent on the data structure that can be characterized analytically for random networks. Near the critical point, the map exhibits unique features and a distinctive "triangular" packing pattern of the incidence matrix.

cond-mat.stat-mech

Evolutionary stability of antigenically escaping viruses

Antigenic variation is the main immune escape mechanism for RNA viruses like influenza or SARS-CoV-2. While high mutation rates promote antigenic escape, they also induce large mutational loads and reduced fitness. It remains unclear how this cost-benefit trade-off selects the mutation rate of viruses. Using a traveling wave model for the co-evolution of viruses and host immune systems in a finite population, we investigate how immunity affects the evolution of the mutation rate and other non-antigenic traits, such as virulence. We first show that the nature of the wave depends on how cross-reactive immune systems are, reconciling previous approaches. The immune-virus system behaves like a Fisher wave at low cross-reactivities, and like a fitness wave at high cross-reactivities. These regimes predict different outcomes for the evolution of non-antigenic traits. At low cross-reactivities, the evolutionarily stable strategy is to maximize the speed of the wave, implying a higher mutation rate and increased virulence. At large cross-reactivities, where our estimates place H3N2 influenza, the stable strategy is to increase the basic reproductive number, keeping the mutation rate to a minimum and virulence low.

q-bio.PE

Inspecting the interaction between HIV and the immune system through genetic turnover

Chronic infections of the human immunodeficiency virus (HIV) create a very complex co-evolutionary process, where the virus tries to escape the continuously adapting host immune system. Quantitative details of this process are largely unknown and could help in disease treatment and vaccine development. Here we study a longitudinal dataset of ten HIV-infected people, where both the B-cell receptors and the virus are deeply sequenced. We focus on simple measures of turnover, which quantify how much the composition of the viral strains and the immune repertoire change between time points. At the single-patient level, the viral-host turnover rates do not show any statistically significant correlation, however they correlate if the information is aggregated across patients. In particular, we identify an anti-correlation: large changes in the viral pool composition come with small changes in the B-cell receptor repertoire. This result seems to contradict the naive expectation that when the virus mutates quickly, the immune repertoire needs to change to keep up. However, we show that the observed anti-correlation naturally emerges and can be understood in terms of simple population-genetics models.

q-bio.PE

NoisET: Noise learning and Expansion detection of T-cell receptors

High-throughput sequencing of T- and B-cell receptors makes it possible to track immune repertoires across time, in different tissues, in acute and chronic diseases and in healthy individuals. However quantitative comparison between repertoires is confounded by variability in the read count of each receptor clonotype due to sampling, library preparation, and expression noise. We review methods for accounting for both biological and experimental noise and present an easy-to-use python package NoisET that implements and generalizes a previously developed Bayesian method. It can be used to learn experimental noise models for repertoire sequencing from replicates, and to detect responding clones following a stimulus. We test the package on different repertoire sequencing technologies and datasets. We review how such approaches have been used to identify responding clonotypes in vaccination and disease data. Availability: NoisET is freely available to use with source code at github.com/statbiophys/NoisET.

q-bio.GN

Generosity, selfishness and exploitation as optimal greedy strategies for resource sharing

Resource sharing outside the kinship bonds is rare. Besides humans, it occurs in chimpanzee, wild dogs and hyenas as well as in vampire bats. Resource sharing is an instance of animal cooperation, where an animal gives away part of the resources that it owns for the benefit of a recipient. Taking inspiration from blood-sharing in vampire bats, here show the emergence of generosity in a Markov game, which couples the resource sharing between two players with the gathering task of that resource. At variance with the classical evolutionary models for cooperation, the optimal strategies of this game can be potentially learned by animals during their life-time. The players act greedily, that is, they try to individually maximize only their personal income. Nonetheless, the analytical solution of the model shows that three non trivial optimal behaviours emerge depending on conditions. Besides the obvious case when players are selfish in their choice of resource division, there are conditions under which both players are generous. Moreover, we also found a range of situations in which one selfish player exploits another generous individual, for the satisfaction of both players. Our results show that resource sharing is favoured by three factors: a long time horizon over which the players try to optimize their own game, the similarity among players in their ability of performing the resource-gathering task, as well as by the availability of resources in the environment. These concurrent requirements lead to identify necessary conditions for the emergence of generosity.

q-bio.PE

Heaps' law, statistics of shared components and temporal patterns from a sample-space-reducing process

Zipf's law is a hallmark of several complex systems with a modular structure, such as books composed by words or genomes composed by genes. In these component systems, Zipf's law describes the empirical power law distribution of component frequencies. Stochastic processes based on a sample-space-reducing (SSR) mechanism, in which the number of accessible states reduces as the system evolves, have been recently proposed as a simple explanation for the ubiquitous emergence of this law. However, many complex component systems are characterized by other statistical patterns beyond Zipf's law, such as a sublinear growth of the component vocabulary with the system size, known as Heap's law, and a specific statistics of shared components. This work shows, with analytical calculations and simulations, that these statistical properties can emerge jointly from a SSR mechanism, thus making it an appropriate parameter-poor representation for component systems. Several alternative (and equally simple) models, for example based on the preferential attachment mechanism, can also reproduce Heaps' and Zipf's laws, suggesting that additional statistical properties should be taken into account to select the most-likely generative process for a specific system. Along this line, we will show that the temporal component distribution predicted by the SSR model is markedly different from the one emerging from the popular rich-gets-richer mechanism. A comparison with empirical data from natural language indicates that the SSR process can be chosen as a better candidate model for text generation based on this statistical property. Finally, a limitation of the SSR model in reproducing the empirical "burstiness" of word appearances in texts will be pointed out, thus indicating a possible direction for extensions of the basic SSR process.

cond-mat.stat-mech

Statistics of shared components in complex component systems

Many complex systems are modular. Such systems can be represented as "component systems", i.e., sets of elementary components, such as LEGO bricks in LEGO sets. The bricks found in a LEGO set reflect a target architecture, which can be built following a set-specific list of instructions. In other component systems, instead, the underlying functional design and constraints are not obvious a priori, and their detection is often a challenge of both scientific and practical importance, requiring a clear understanding of component statistics. Importantly, some quantitative invariants appear to be common to many component systems, most notably a common broad distribution of component abundances, which often resembles the well-known Zipf's law. Such "laws" affect in a general and non-trivial way the component statistics, potentially hindering the identification of system-specific functional constraints or generative processes. Here, we specifically focus on the statistics of shared components, i.e., the distribution of the number of components shared by different system-realizations, such as the common bricks found in different LEGO sets. To account for the effects of component heterogeneity, we consider a simple null model, which builds system-realizations by random draws from a universe of possible components. Under general assumptions on abundance heterogeneity, we provide analytical estimates of component occurrence, which quantify exhaustively the statistics of shared components. Surprisingly, this simple null model can positively explain important features of empirical component-occurrence distributions obtained from data on bacterial genomes, LEGO sets, and book chapters. Specific architectural features and functional constraints can be detected from occurrence patterns as deviations from these null predictions, as we show for the illustrative case of the "core" genome in bacteria.

q-bio.GN

Zipf and Heaps laws from dependency structures in component systems

Complex natural and technological systems can be considered, on a coarse-grained level, as assemblies of elementary components: for example, genomes as sets of genes, or texts as sets of words. On one hand, the joint occurrence of components emerges from architectural and specific constraints in such systems. On the other hand, general regularities may unify different systems, such as the broadly studied Zipf and Heaps laws, respectively concerning the distribution of component frequencies and their number as a function of system size. Dependency structures (i.e., directed networks encoding the dependency relations between the components in a system) were proposed recently as a possible organizing principles underlying some of the regularities observed. However, the consequences of this assumption were explored only in binary component systems, where solely the presence or absence of components is considered, and multiple copies of the same component are not allowed. Here, we consider a simple model that generates, from a given ensemble of dependency structures, a statistical ensemble of sets of components, allowing for components to appear with any multiplicity. Our model is a minimal extension that is memoryless, and therefore accessible to analytical calculations. A mean-field analytical approach (analogous to the "Zipfian ensemble" in the linguistics literature) captures the relevant laws describing the component statistics as we show by comparison with numerical computations. In particular, we recover a power-law Zipf rank plot, with a set of core components, and a Heaps law displaying three consecutive regimes (linear, sub-linear and saturating) that we characterize quantitatively.

physics.soc-ph

Adaptation and irreversibility in microevolution

Within the framework of population genetics we consider the evolution of an asexual haploid population under the effect of a rapidly varying natural selection (microevolution). We focus on the case in which the environment exerting selection changes stochastically. We derive the effective genotype and fitness dynamics on the slower time-scales at which the relevant genetic modifications take place. We find that, despite the fast environmental switches, the population manages to adapt on the fast time-scales yielding a finite positive contribution to the fitness. However, such contribution is balanced by the continuous loss in fitness due to the varying selection so that the statistics of the global fitness can be described neglecting the details of the fast environmental process. The occurrence of adaptation on fast time-scales would be undetectable if one were to consider only the effective genotype and fitness dynamics on the slow time-scales. We therefore propose an experimental observable to detect it.

q-bio.PE