SearcharxivSearch

arXiv subjects

Matteo Osella

Publications and source records attributed to Matteo Osella.

18 recordsLinked to original sources

Alternative routes to universal diversity scaling in component systems: from proteomes to large language models

Remarkably common statistical laws characterize the diversity scaling and its fluctuations across a wide range of complex "component systems". These regularities are often interpreted as signatures of an underlying innovation mechanism driving the growth of component diversity, but the basic ingredients necessary for their emergence remain poorly understood. In particular, from language and technological artifacts to genomes and gene expression patterns, the number of distinct components grows sublinearly with system size, while its variance scales approximately as the square of its mean. This behavior is consistent across diverse systems, raising the question of whether general constraints or emergent principles underlying diversity and innovation define the architectures of realizations with different numbers of components. To address this question, we derive analytical conditions for the joint emergence of these two diversity laws within a broad class of growth models, showing that they require a specific asymptotic dependence of the innovation probability on diversity and system size. We then demonstrate that the same macroscopic laws arise in a different class of models with latent heterogeneity, where quadratic fluctuation scaling always emerges asymptotically as a consequence of general statistical principles, essentially the law of total variance, without explicitly assuming an innovation mechanism or any specific rule for system assembly. We compare these predictions with empirical data from language, genomes, LEGO constructions, and texts generated by large language models. Our results show that empirical diversity scaling laws strongly constrain generative models but do not uniquely identify the mechanisms generating diversity, revealing a close correspondence between innovation-driven growth models and latent-variable descriptions.

cond-mat.stat-mech

Component systems: do null models explain everything?

Component systems - ensembles of realizations built from a shared repertoire of modular parts - are ubiquitous in biological, ecological, technological, and socio-cultural domains. From genomes to texts, cities, and software, these systems exhibit statistical regularities that often meet the "bona fide" requirements of laws in the physical sciences. Here, we argue that the generality and simplicity of those laws are often due to basic combinatorial or sampling constraints, raising the question of whether such patterns are actually revealing system-specific mechanisms and how we might move beyond them. To this end, we first present a unifying mathematical framework, which allows us to compare modular systems in different fields and highlights the common "null" trends as well as the system-specific uniqueness, which, arguably, are signatures of the underlying generative dynamics. Next, we can exploit the framework with statistical mechanics and modern machine-learning tools for a twofold objective. (i) Explaining why the general regularities emerge, highlighting the constraints between them and the general principles at their origins, and (ii) "subtracting" them from data, which will isolate the informative features for inferring hidden system-specific generative processes, mechanistic and causal aspects.

cond-mat.stat-mech

The Advantage of Fine-Grained Training

In classification problems, models are trained to predict a class label based on the input data features. However, class labels are organized hierarchically in many datasets. While a classification task is often defined at a specific level of this hierarchy, training can utilize a finer granularity of labels. Empirical evidence suggests that such fine-grained training can enhance performance. In this work, we investigate the generality of this observation and explore its underlying causes using both real and synthetic datasets. We show that training on fine-grained labels does not universally improve classification accuracy. Instead, the effectiveness of this strategy depends on the geometric structure of the data and its relations with the label hierarchy. Specifically, we show that the advantage of fine-grained training crucially depends on the degree of alignment between the decision boundaries required for the fine- and coarse-grained tasks, a property that we term boundary redundancy. Additionally, factors such as dataset size and model capacity significantly influence whether fine-grained labels provide a performance benefit. Indeed, we identify a transition, whose location is largely controlled by the degree of overparameterization, separating regimes where fine-grained training improves performance from those where direct coarse-grained training is preferable.

cs.LG

Ranking nodes in bipartite systems with a non-linear iterative map

Ranking nodes in networks according to a defined measure of importance is an extensively studied task, with applications in ecology, economic trade networks, and social networks. This paper introduces a method based on a non-linear iterative map to evaluate node relevance in bipartite networks. By tuning a single parameter $\gamma$, the method captures different concepts of node importance, including established measures like degree centrality, eigenvector centrality and the fitness-complexity ranking. The algorithm's flexibility allows for efficient ranking optimization tailored to specific tasks, outperforming state-of-the-art algorithms. We apply this method to ecological mutualistic networks, where ranking quality can be assessed by the extinction area - the rate at which the system collapses when species are removed in a certain order. The map with the optimal $\gamma$ value surpasses existing ranking methods on this task. Additionally, our method excels in evaluating nestedness, another crucial structural property of ecological systems, requiring specific node rankings. Finally, we explore theoretical aspects of the map, revealing a phase transition at a critical $\gamma$ dependent on the data structure that can be characterized analytically for random networks. Near the critical point, the map exhibits unique features and a distinctive "triangular" packing pattern of the incidence matrix.

cond-mat.stat-mech

Inversion dynamics of class manifolds in deep learning reveals tradeoffs underlying generalisation

To achieve near-zero training error in a classification problem, the layers of a feed-forward network have to disentangle the manifolds of data points with different labels, to facilitate the discrimination. However, excessive class separation can bring to overfitting since good generalisation requires learning invariant features, which involve some level of entanglement. We report on numerical experiments showing how the optimisation dynamics finds representations that balance these opposing tendencies with a non-monotonic trend. After a fast segregation phase, a slower rearrangement (conserved across data sets and architectures) increases the class entanglement.The training error at the inversion is stable under subsampling, and across network initialisations and optimisers, which characterises it as a property solely of the data structure and (very weakly) of the architecture. The inversion is the manifestation of tradeoffs elicited by well-defined and maximally stable elements of the training set, coined ``stragglers'', particularly influential for generalisation.

cs.LG

The effect of a linear feedback mechanism in a homeostasis model

Feedback loops are essential for regulating cell proliferation and maintaining the delicate balance between cell division and cell death. Thanks to the exact solution of a few simple models of cell growth it is by now clear that stochastic fluctuations play a central role in this process and that cell growth (and in particular the robustness and stability of homeostasis) can be properly addressed only as a stochastic process. Using epidermal homeostasis as a prototypical example, we show that it is possible to discriminate among different feedback strategies which turn out to be characterized by different, experimentally testable, behaviours. In particular, we focus on the so-called Dynamical Heterogeneity model, an epidermal homeostasis model that takes into account two well known cellular features: the plasticity of the cells and their adaptability to face environmental stimuli. We show that specific choices of the parameter on which the feedback is applied may decrease the fluctuations of the homeostatic population level and improve the recovery of the system after an external perturbation.

q-bio.PE

Initial cell density encodes proliferative potential in cancer cell populations

Individual cells exhibit specific proliferative responses to changes in microenvironmental conditions. Whether such potential is constrained by the cell density throughout the growth process is however unclear. Here, we identify a theoretical framework that captures how the information encoded in the initial density of cancer cell populations impacts their growth profile. By following the growth of hundreds of populations of cancer cells, we found that the time they need to adapt to the environment decreases as the initial cell density increases. Moreover, the population growth rate shows a maximum at intermediate initial densities. With the support of a mathematical model, we show that the observed interdependence of adaptation time and growth rate is significantly at odds both with standard logistic growth models and with the Monod-like function that governs the dependence of the growth rate on nutrient levels. Our results (i) uncover and quantify a previously unnoticed heterogeneity in the growth dynamics of cancer cell populations; (ii) unveil how population growth may be affected by single-cell adaptation times; (iii) contribute to our understanding of the clinically-observed dependence of the primary and metastatic tumor take rates on the initial density of implanted cancer cells.

q-bio.QM

Modelling the evolution of transcription factor binding preferences in complex eukaryotes

Transcription factors (TFs) exert their regulatory action by binding to DNA with specific sequence preferences. However, different TFs can partially share their binding sequences due to their common evolutionary origin. This `redundancy' of binding defines a way of organizing TFs in `motif families' by grouping TFs with similar binding preferences. Since these ultimately define the TF target genes, the motif family organization entails information about the structure of transcriptional regulation as it has been shaped by evolution. Focusing on the human TF repertoire, we show that a one-parameter evolutionary model of the Birth-Death-Innovation type can explain the TF empirical ripartition in motif families, and allows to highlight the relevant evolutionary forces at the origin of this organization. Moreover, the model allows to pinpoint few deviations from the neutral scenario it assumes: three over-expanded families (including HOX and FOX genes), a set of `singleton' TFs for which duplication seems to be selected against, and a higher-than-average rate of diversification of the binding preferences of TFs with a Zinc Finger DNA binding domain. Finally, a comparison of the TF motif family organization in different eukaryotic species suggests an increase of redundancy of binding with organism complexity.

q-bio.GN

Heaps' law, statistics of shared components and temporal patterns from a sample-space-reducing process

Zipf's law is a hallmark of several complex systems with a modular structure, such as books composed by words or genomes composed by genes. In these component systems, Zipf's law describes the empirical power law distribution of component frequencies. Stochastic processes based on a sample-space-reducing (SSR) mechanism, in which the number of accessible states reduces as the system evolves, have been recently proposed as a simple explanation for the ubiquitous emergence of this law. However, many complex component systems are characterized by other statistical patterns beyond Zipf's law, such as a sublinear growth of the component vocabulary with the system size, known as Heap's law, and a specific statistics of shared components. This work shows, with analytical calculations and simulations, that these statistical properties can emerge jointly from a SSR mechanism, thus making it an appropriate parameter-poor representation for component systems. Several alternative (and equally simple) models, for example based on the preferential attachment mechanism, can also reproduce Heaps' and Zipf's laws, suggesting that additional statistical properties should be taken into account to select the most-likely generative process for a specific system. Along this line, we will show that the temporal component distribution predicted by the SSR model is markedly different from the one emerging from the popular rich-gets-richer mechanism. A comparison with empirical data from natural language indicates that the SSR process can be chosen as a better candidate model for text generation based on this statistical property. Finally, a limitation of the SSR model in reproducing the empirical "burstiness" of word appearances in texts will be pointed out, thus indicating a possible direction for extensions of the basic SSR process.

cond-mat.stat-mech

Statistics of shared components in complex component systems

Many complex systems are modular. Such systems can be represented as "component systems", i.e., sets of elementary components, such as LEGO bricks in LEGO sets. The bricks found in a LEGO set reflect a target architecture, which can be built following a set-specific list of instructions. In other component systems, instead, the underlying functional design and constraints are not obvious a priori, and their detection is often a challenge of both scientific and practical importance, requiring a clear understanding of component statistics. Importantly, some quantitative invariants appear to be common to many component systems, most notably a common broad distribution of component abundances, which often resembles the well-known Zipf's law. Such "laws" affect in a general and non-trivial way the component statistics, potentially hindering the identification of system-specific functional constraints or generative processes. Here, we specifically focus on the statistics of shared components, i.e., the distribution of the number of components shared by different system-realizations, such as the common bricks found in different LEGO sets. To account for the effects of component heterogeneity, we consider a simple null model, which builds system-realizations by random draws from a universe of possible components. Under general assumptions on abundance heterogeneity, we provide analytical estimates of component occurrence, which quantify exhaustively the statistics of shared components. Surprisingly, this simple null model can positively explain important features of empirical component-occurrence distributions obtained from data on bacterial genomes, LEGO sets, and book chapters. Specific architectural features and functional constraints can be detected from occurrence patterns as deviations from these null predictions, as we show for the illustrative case of the "core" genome in bacteria.

q-bio.GN

Zipf and Heaps laws from dependency structures in component systems

Complex natural and technological systems can be considered, on a coarse-grained level, as assemblies of elementary components: for example, genomes as sets of genes, or texts as sets of words. On one hand, the joint occurrence of components emerges from architectural and specific constraints in such systems. On the other hand, general regularities may unify different systems, such as the broadly studied Zipf and Heaps laws, respectively concerning the distribution of component frequencies and their number as a function of system size. Dependency structures (i.e., directed networks encoding the dependency relations between the components in a system) were proposed recently as a possible organizing principles underlying some of the regularities observed. However, the consequences of this assumption were explored only in binary component systems, where solely the presence or absence of components is considered, and multiple copies of the same component are not allowed. Here, we consider a simple model that generates, from a given ensemble of dependency structures, a statistical ensemble of sets of components, allowing for components to appear with any multiplicity. Our model is a minimal extension that is memoryless, and therefore accessible to analytical calculations. A mean-field analytical approach (analogous to the "Zipfian ensemble" in the linguistics literature) captures the relevant laws describing the component statistics as we show by comparison with numerical computations. In particular, we recover a power-law Zipf rank plot, with a set of core components, and a Heaps law displaying three consecutive regimes (linear, sub-linear and saturating) that we characterize quantitatively.

physics.soc-ph

Stochastic timing in gene expression for simple regulatory strategies

Timing is essential for many cellular processes, from cellular responses to external stimuli to the cell cycle and circadian clocks. Many of these processes are based on gene expression. For example, an activated gene may be required to reach in a precise time a threshold level of expression that triggers a specific downstream process. However, gene expression is subject to stochastic fluctuations, naturally inducing an uncertainty in this threshold-crossing time with potential consequences on biological functions and phenotypes. Here, we consider such "timing fluctuations", and we ask how they can be controlled. Our analytical estimates and simulations show that, for an induced gene, timing variability is minimal if the threshold level of expression is approximately half of the steady-state level. Timing fuctuations can be reduced by increasing the transcription rate, while they are insensitive to the translation rate. In presence of self-regulatory strategies, we show that self-repression reduces timing noise for threshold levels that have to be reached quickly, while selfactivation is optimal at long times. These results lay a framework for understanding stochasticity of endogenous systems such as the cell cycle, as well as for the design of synthetic trigger circuits.

q-bio.MN

Relevant parameters in models of cell division control

A recent burst of dynamic single-cell growth-division data makes it possible to characterize the stochastic dynamics of cell division control in bacteria. Different modeling frameworks were used to infer specific mechanisms from such data, but the links between frameworks are poorly explored, with relevant consequences for how well any particular mechanism can be supported by the data. Here, we describe a simple and generic framework in which two common formalisms can be used interchangeably: (i) a continuous-time division process described by a hazard function and (ii) a discrete-time equation describing cell size across generations (where the unit of time is a cell cycle). In our framework, this second process is a discrete-time Langevin equation with a simple physical analogue. By perturbative expansion around the mean initial size (or inter-division time), we show explicitly how this framework describes a wide range of division control mechanisms, including combinations of time and size control, as well as the constant added size mechanism recently found to capture several aspects of the cell division behavior of different bacteria. As we show by analytical estimates and numerical simulation, the available data are characterized with great precision by the first-order approximation of this expansion. Hence, a single dimensionless parameter defines the strength and the action of the division control. However, this parameter may emerge from several mechanisms, which are distinguished only by higher-order terms in our perturbative expansion. An analytical estimate of the sample size needed to distinguish between second-order effects shows that this is larger than what is available in the current datasets. These results provide a unified framework for future studies and clarify the relevant parameters at play in the control of cell division.

q-bio.CB

Individuality and universality in the growth-division laws of single E. coli cells

The mean size of exponentially dividing E. coli cells cultured in different nutrient conditions is known to depend on the mean growth rate only. However, the joint fluctuations relating cell size, doubling time and individual growth rate are only starting to be characterized. Recent studies in bacteria (i) revealed the near constancy of the size extension in a single cell cycle (adder mechanism), and (ii) reported a universal trend where the spread in both size and doubling times is a linear function of the population means of these variables. Here, we combine experiments and theory and use scaling concepts to elucidate the constraints posed by the second observation on the division control mechanism and on the joint fluctuations of sizes and doubling times. We found that scaling relations based on the means both collapse size and doubling-time distributions across different conditions, and explain how the shape of their joint fluctuations deviates from the means. Our data on these joint fluctuations highlight the importance of cell individuality: single cells do not follow the dependence observed for the means between size and either growth rate or inverse doubling time. Our calculations show that these results emerge from a broad class of division control mechanisms (including the adder mechanism as a particular case) requiring a certain scaling form of the so-called "division hazard rate function", which defines the probability rate of dividing as a function of measurable parameters. This gives a rationale for the universal body-size distributions observed in microbial ecosystems across many microbial species, presumably dividing with multiple mechanisms. Additionally, our experiments show a crossover between fast and slow growth in the relation between individual-cell growth rate and division time, which can be understood in terms of different regimes of genome replication control.

q-bio.CB

Growth-rate-dependent dynamics of a bacterial genetic oscillator

Gene networks exhibiting oscillatory dynamics are widespread in biology. The minimal regulatory designs giving rise to oscillations have been implemented synthetically and studied by mathematical modeling. However, most of the available analyses generally neglect the coupling of regulatory circuits with the cellular "chassis" in which the circuits are embedded. For example, the intracellular macromolecular composition of fast-growing bacteria changes with growth rate. As a consequence, important parameters of gene expression, such as ribosome concentration or cell volume, are growth-rate dependent, ultimately coupling the dynamics of genetic circuits with cell physiology. This work addresses the effects of growth rate on the dynamics of a paradigmatic example of genetic oscillator, the repressilator. Making use of empirical growth-rate dependences of parameters in bacteria, we show that the repressilator dynamics can switch between oscillations and convergence to a fixed point depending on the cellular state of growth, and thus on the nutrients it is fed. The physical support of the circuit (type of plasmid or gene positions on the chromosome) also plays an important role in determining the oscillation stability and the growth-rate dependence of period and amplitude. This analysis has potential application in the field of synthetic biology, and suggests that the coupling between endogenous genetic oscillators and cell physiology can have substantial consequences for their functionality.

q-bio.MN

Speed of evolution in large asexual populations with diminishing returns

The adaptive evolution of large asexual populations is generally characterized by competition between clones carrying different beneficial mutations. This interference phenomenon slows down the adaptation speed and makes the theoretical description of the dynamics more complex with respect to the successional occurrence and fixation of beneficial mutations typical of small populations. A simplified modeling framework considering multiple beneficial mutations with equal and constant fitness advantage captures some of the essential features of the actual complex dynamics, and some key predictions from this model are verified in laboratory evolution experiments. However, in these experiments the relative advantage of a beneficial mutation is generally dependent on the genetic background. In particular, the general pattern is that, as mutations in different loci accumulate, the relative advantage of new mutations decreases, trend often referred to as "diminishing return" epistasis. In this paper, we propose a phenomenological model that generalizes the fixed-advantage framework to include in a simple way this feature. To evaluate the quantitative consequences of diminishing returns on the evolutionary dynamics, we approach the model analytically as well as with direct simulations. Finally, we show how the model parameters can be matched with data from evolutionary experiments in order to infer the mean effect of epistasis and derive order-of-magnitude estimates of the rate of beneficial mutations. Applying this procedure to two experimental data sets gives values of the beneficial mutation rate within the range of previous measurements.

q-bio.PE

Gene autoregulation via intronic microRNAs and its functions

Background: MicroRNAs, post-transcriptional repressors of gene expression, play a pivotal role in gene regulatory networks. They are involved in core cellular processes and their dysregulation is associated to a broad range of human diseases. This paper focus on a minimal microRNA-mediated regulatory circuit, in which a protein-coding gene (host gene) is targeted by a microRNA located inside one of its introns. Results: Autoregulation via intronic microRNAs is widespread in the human regulatory network, as confirmed by our bioinformatic analysis, and can perform several regulatory tasks despite its simple topology. Our analysis, based on analytical calculations and simulations, indicates that this circuitry alters the dynamics of the host gene expression, can induce complex responses implementing adaptation and Weber's law, and efficiently filters fluctuations propagating from the upstream network to the host gene. A fine-tuning of the circuit parameters can optimize each of these functions. Interestingly, they are all related to gene expression homeostasis, in agreement with the increasing evidence suggesting a role of microRNA regulation in conferring robustness to biological processes. In addition to model analysis, we present a list of bioinformatically predicted candidate circuits in human for future experimental tests. Conclusions: The results presented here suggest a potentially relevant functional role for negative self-regulation via intronic microRNAs, in particular as a homeostatic control mechanism of gene expression. Moreover, the map of circuit functions in terms of experimentally measurable parameters, resulting from our analysis, can be a useful guideline for possible applications in synthetic biology.

q-bio.MN

The role of incoherent microRNA-mediated feedforward loops in noise buffering

MicroRNAs are endogenous non-coding RNAs which negatively regulate the expression of protein-coding genes in plants and animals. They are known to play an important role in several biological processes and, together with transcription factors, form a complex and highly interconnected regulatory network. Looking at the structure of this network it is possible to recognize a few overrepresented motifs which are expected to perform important elementary regulatory functions. Among them a special role is played by the microRNA-mediated feedforward loop in which a master transcription factor regulates a microRNA and, together with it, a set of target genes. In this paper we show analytically and through simulations that the incoherent version of this motif can couple the fine-tuning of a target protein level with an efficient noise control, thus conferring precision and stability to the overall gene expression program, especially in the presence of fluctuations in upstream regulators. Among the other results, a nontrivial prediction of our model is that the optimal attenuation of fluctuations coincides with a modest repression of the target expression. This feature is coherent with the expected fine-tuning function and in agreement with experimental observations of the actual impact of a wide class of microRNAs on the protein output of their targets. Finally we describe the impact on noise-buffering efficiency of the cross-talk between microRNA targets that can naturally arise if the microRNA-mediated circuit is not considered as isolated, but embedded in a larger network of regulations.

q-bio.MN