SearcharxivSearch

arXiv subjects

Johannes Berg

Publications and source records attributed to Johannes Berg.

At least 19 recordsLinked to original sources

Branch length statistics in phylogenetic trees under constant-rate birth-death dynamics

Phylogenetic trees represent the evolutionary relationships between extant lineages, where extinct or non-sampled lineages are omitted. Extending the work of Stadler and collaborators, this paper focuses on the branch lengths in phylogenetic trees arising under a constant-rate birth-death model. We derive branch length distributions of phylogenetic branches with and without random sampling of individuals of the extant population under two distinct statistical scenarios: a fixed age of the birth-death process and a fixed number of individuals at the time of observation. We find that branches connected to the tree leaves (pendant branches) and branches in the interior of the tree behave very differently under sampling; pendant branches grow longer without limit as the sampling probability is decreased, whereas the interior branch lengths quickly reach an asymptotic distribution that does not depend on the sampling probability.

q-bio.PE

Power-laws in phylogenetic trees and the preferential coalescent

Phylogenetic trees capture evolutionary relationships among species and reflect the forces that shaped them. While many studies rely on branch length information, the topology of phylogenetic trees (particularly their degree of imbalance) offers a robust framework for inferring evolutionary dynamics when timing data is uncertain. Classical metrics, such as the Colless and Sackin indices, quantify tree imbalance and have been extensively used to characterize phylogenies. Empirical phylogenies typically show intermediate imbalance, falling between perfectly balanced and highly skewed trees. This regime is marked by a power-law relationship between subtree sizes and their cumulative sizes, governed by a characteristic exponent. Although a recent niche-size model replicates this scaling, its mathematical origin and the exponent's value remain unclear. We present a generative model inspired by Kingman's coalescent that incorporates niche-like dynamics through preferential node coalescence. This process maps to Smoluchowski's coagulation kinetics and is described by a generalized Smoluchowski equation. Our model produces imbalanced trees with power-law exponents matching empirical and numerical observations, revealing the mathematical basis of observed scaling laws and offering new tools to interpret tree imbalance in evolutionary contexts.

q-bio.PE

Inferring stochastic regulatory networks from perturbations of the non-equilibrium steady state

Regulatory networks describe the interactions between molecular or cellular regulators, like transcription factors and genes in gene regulatory networks, kinases and their receptors in signalling networks, or neurons in neural networks. A long-standing aim of quantitative biology is to reconstruct such networks on the basis of large-scale data. Our aim is to leverage fluctuations around the non-equilibrium steady state for network inference. To this end, we use a stochastic model of gene regulation or neural dynamics and solve it approximately within a Gaussian mean-field theory. We develop a likelihood estimate based on this stochastic theory to infer regulatory interactions from perturbation data on the network nodes. We apply this approach to artificial perturbation data as well as to phospho-proteomic data from cell-line experiments and compare our results to inference schemes restricted to mean activities in the steady state.

q-bio.MN

Stochastic clonal dynamics and genetic turnover in exponentially growing populations

We consider an exponentially growing population of cells undergoing mutations and ask about the effect of reproductive fluctuations (genetic drift) on its long-term evolution. We combine first step analysis with the stochastic dynamics of a birth-death process to analytically calculate the probability that the parent of a given genotype will go extinct. We compare the results with numerical simulations and show how this turnover of genetic clones can be used to infer the rates underlying the population dynamics. Our work is motivated by growing populations of tumour cells, the epidemic spread of viruses, and bacterial growth.

q-bio.PE

Switching off: the phenotypic transition to the uninduced state of the lactose uptake pathway

The lactose uptake-pathway of E. coli is a paradigmatic example of multistability in gene-regulatory circuits. In the induced state of the lac-pathway, the genes comprising the lac-operon are transcribed, leading to the production of proteins which import and metabolize lactose. In the uninduced state, a stable repressor-DNA loop frequently blocks the transcription of the lac-genes. Transitions from one phenotypic state to the other are driven by fluctuations, which arise from the random timing of the binding of ligands and proteins. This stochasticity affects transcription and translation, and ultimately molecular copy numbers. Our aim is to understand the transition from the induced to the uninduced state of the lac-operon. We use a detailed computational model to show that repressor-operator binding/unbinding, fluctuations in the total number of repressors, and inducer-repressor binding/unbinding all play a role in this transition. Based on the timescales on which these processes operate, we construct a minimal model of the transition to the uninduced state and compare the results with simulations and experimental observations. The induced state turns out to be very stable, with a transition rate to the uninduced state lower than $2 \times 10^{-9}$ per minute. In contrast to the transition to the induced state, the transition to the uninduced state is well described in terms of a 2D diffusive system crossing a barrier, with the diffusion rates emerging from a model of repressor unbinding.

q-bio.MN

Inferring the parameters of a Markov process from snapshots of the steady state

We seek to infer the parameters of an ergodic Markov process from samples taken independently from the steady state. Our focus is on non-equilibrium processes, where the steady state is not described by the Boltzmann measure, but is generally unknown and hard to compute, which prevents the application of established equilibrium inference methods. We propose a quantity we call propagator likelihood, which takes on the role of the likelihood in equilibrium processes. This propagator likelihood is based on fictitious transitions between those configurations of the system which occur in the samples. The propagator likelihood can be derived by minimising the relative entropy between the empirical distribution and a distribution generated by propagating the empirical distribution forward in time. Maximising the propagator likelihood leads to an efficient reconstruction of the parameters of the underlying model in different systems, both with discrete configurations and with continuous configurations. We apply the method to non-equilibrium models from statistical physics and theoretical biology, including the asymmetric simple exclusion process (ASEP), the kinetic Ising model, and replicator dynamics.

cond-mat.stat-mech

Inverse statistical problems: from the inverse Ising problem to data science

Inverse problems in statistical physics are motivated by the challenges of `big data' in different fields, in particular high-throughput experiments in biology. In inverse problems, the usual procedure of statistical physics needs to be reversed: Instead of calculating observables on the basis of model parameters, we seek to infer parameters of a model based on observations. In this review, we focus on the inverse Ising problem and closely related problems, namely how to infer the coupling strengths between spins given observed spin correlations, magnetisations, or other data. We review applications of the inverse Ising problem, including the reconstruction of neural connections, protein structure determination, and the inference of gene regulatory networks. For the inverse Ising problem in equilibrium, a number of controlled and uncontrolled approximate solutions have been developed in the statistical mechanics community. A particularly strong method, pseudolikelihood, stems from statistics. We also review the inverse Ising problem in the non-equilibrium case, where the model parameters must be reconstructed based on non-equilibrium statistics.

cond-mat.dis-nn

Statistical mechanics of the inverse Ising problem and the optimal objective function

The inverse Ising problem seeks to reconstruct the parameters of an Ising Hamiltonian on the basis of spin configurations sampled from the Boltzmann measure. Over the last decade, many applications of the inverse Ising problem have arisen, driven by the advent of large-scale data across different scientific disciplines. Recently, strategies to solve the inverse Ising problem based on convex optimisation have proven to be very successful. These approaches maximise particular objective functions with respect to the model parameters. Examples are the pseudolikelihood method and interaction screening. In this paper, we establish a link between approaches to the inverse Ising problem based on convex optimisation and the statistical physics of disordered systems. We characterise the performance of an arbitrary objective function and calculate the objective function which optimally reconstructs the model parameters. We evaluate the optimal objective function within a replica-symmetric ansatz and compare the results of the optimal objective function with other reconstruction methods. Apart from giving a theoretical underpinning to solving the inverse Ising problem by convex optimisation, the optimal objective function outperforms state-of-the-art methods, albeit by a small margin.

cond-mat.dis-nn

Network inference in the non-equilibrium steady state

Non-equilibrium systems lack an explicit characterisation of their steady state like the Boltzmann distribution for equilibrium systems. This has drastic consequences for the inference of parameters of a model when its dynamics lacks detailed balance. Such non-equilibrium systems occur naturally in applications like neural networks or gene regulatory networks. Here, we focus on the paradigmatic asymmetric Ising model and show that we can learn its parameters from independent samples of the non-equilibrium steady state. We present both an exact inference algorithm and a computationally more efficient, approximate algorithm for weak interactions based on a systematic expansion around mean-field theory. Obtaining expressions for magnetisations, two- and three-point spin correlations, we establish that these observables are sufficient to infer the model parameters. Further, we discuss the symmetries characterising the different orders of the expansion around the mean field and show how different types of dynamics can be distinguished on the basis of samples from the non-equilibrium steady state.

cond-mat.stat-mech

Multiple-line inference of selection on quantitative traits

Trait differences between species may be attributable to natural selection. However, quantifying the strength of evidence for selection acting on a particular trait is a difficult task. Here we develop a population-genetic test for selection acting on a quantitative trait which is based on multiple-line crosses. We show that using multiple lines increases both the power and the scope of selection inference. First, a test based on three or more lines detects selection with strongly increased statistical significance, and we show explicitly how the sensitivity of the test depends on the number of lines. Second, a multiple-line test allows to distinguish different lineage-specific selection scenarios. Our analytical results are complemented by extensive numerical simulations. We then apply the multiple-line test to QTL data on floral character traits in plant species of the Mimulus genus and on photoperiodic traits in different maize strains, where we find a signatures of lineage-specific selection not seen in a two-line test.

q-bio.PE

Pervasive adaptation of gene expression in Drosophila

Gene expression levels are important molecular quantitative traits that link genotypes to molecular functions and fitness. In Drosophila, population-genetic studies in recent years have revealed substantial adaptive evolution at the genomic level. However, the evolutionary modes of gene expression have remained controversial. Here we present evidence that adaptation dominates the evolution of gene expression levels in flies. We show that 63% of the observed expression divergence across seven Drosophila species are adaptive changes driven by directional selection. Our results are derived from the variation of expression within species and the time-resolved divergence across a family of related species, using a new inference method for selection. We identify functional classes of adaptively regulated genes, as well as sex-specific adaptation occurring predominantly in males. Our analysis opens a new avenue to map system-wide selection on molecular quantitative traits independently of their genetic basis.

q-bio.PE

What makes the lac-pathway switch: identifying the fluctuations that trigger phenotype switching in gene regulatory systems

Multistable gene regulatory systems sustain different levels of gene expression under identical external conditions. Such multistability is used to encode phenotypic states in processes including nutrient uptake and persistence in bacteria, fate selection in viral infection, cell cycle control, and development. Stochastic switching between different phenotypes can occur as the result of random fluctuations in molecular copy numbers of mRNA and proteins arising in transcription, translation, transport, and binding. However, which component of a pathway triggers such a transition is generally not known. By linking single-cell experiments on the lactose-uptake pathway in E. coli to molecular simulations, we devise a general method to pinpoint the particular fluctuation driving phenotype switching and apply this method to the transition between the uninduced and induced states of the lac genes. We find that the transition to the induced state is not caused only by the single event of lac-repressor unbinding, but depends crucially on the time period over which the repressor remains unbound from the lac-operon. We confirm this notion in strains with a high expression level of the repressor (leading to shorter periods over which the lac-operon remains unbound), which show a reduced switching rate. Our techniques apply to multi-stable gene regulatory systems in general and allow to identify the molecular mechanisms behind stochastic transitions in gene regulatory circuits.

q-bio.MN

Can we always sweep the details of RNA-processing under the carpet?

RNA molecules follow a succession of enzyme-mediated processing steps from transcription until maturation. The participating enzymes, for example the spliceosome for mRNAs and Drosha and Dicer for microRNAs, are also produced in the cell and their copy-numbers fluctuate over time. Enzyme copy-number changes affect the processing rate of the substrate molecules; high enzyme numbers increase the processing probability, low enzyme numbers decrease it. We study different RNA processing cascades where enzyme copy-numbers are either fixed or fluctuate. We find that for fixed enzyme-copy numbers the substrates at steady-state are Poisson-distributed, and the whole RNA cascade dynamics can be understood as a single birth-death process of the mature RNA product. In this case, solely fluctuations in the timing of RNA processing lead to variation in the number of RNA molecules. However, we show analytically and numerically that when enzyme copy-numbers fluctuate, the strength of RNA fluctuations increases linearly with the RNA transcription rate. This linear effect becomes stronger as the speed of enzyme dynamics decreases relative to the speed of RNA dynamics. Interestingly, we find that under certain conditions, the RNA cascade can reduce the strength of fluctuations in the expression level of the mature RNA product. Finally, by investigating the effects of processing polymorphisms we show that it is possible for the effects of transcriptional polymorphisms to be enhanced, reduced, or even reversed. Our results provide a framework to understand the dynamics of RNA processing.

q-bio.QM

Quantitative analysis of competition in post-transcriptional regulation reveals a novel signature in target expression variation

When small RNAs are loaded onto Argonaute proteins they can form the RNA-induced silencing complexes (RISCs), which mediate RNA interference. RISC-formation is dependent on a shared pool of Argonaute proteins and RISC loading factors, and is thus susceptible to competition among small RNAs for loading. We present a mathematical model that aims to understand how small RNA competition for the PTR resources affects target gene repression. We discuss that small RNA activity is limited by RISC-formation, RISC-degradation and the availability of Argonautes. Together, these observations explain a number of PTR saturation effects encountered experimentally. We show that different competition conditions for RISC-loading result in different signatures of PTR activity determined also by the amount of RISC-recycling taking place. In particular, we find that the small RNAs less efficient at RISC-formation, using fewer resources of the PTR pathway, can perform in the low RISC-recycling range equally well as their more effective counterparts. Additionally, we predict a novel signature of PTR in target expression levels. Under conditions of low RISC-loading efficiency and high RISC-recycling, the variation in target levels increases linearly with the target transcription rate. Furthermore, we show that RISC-recycling determines the effect that Argonaute scarcity conditions have on target expression variation. Our observations taken together offer a framework of predictions which can be used in order to infer from experimental data the particular characteristics of underlying PTR activity.

q-bio.MN

A statistical mechanics approach to the sample deconvolution problem

In a multicellular organism different cell types express a gene in different amounts. Samples from which gene expression levels can be measured typically contain a mixture of different cell types, the resulting measurements thus give only averages over the different cell types present. Based on fluctuations in the mixture proportions from sample to sample it is in principle possible to reconstruct the underlying expression levels of each cell type: to deconvolute the sample. We use a statistical mechanics approach to the problem of deconvoluting such partial concentrations from mixed samples, give analytical results for when and how well samples can be unmixed, and suggest an algorithm for sample deconvolution.

q-bio.QM

Mean-field theory for the inverse Ising problem at low temperatures

The large amounts of data from molecular biology and neuroscience have lead to a renewed interest in the inverse Ising problem: how to reconstruct parameters of the Ising model (couplings between spins and external fields) from a number of spin configurations sampled from the Boltzmann measure. To invert the relationship between model parameters and observables (magnetisations and correlations) mean-field approximations are often used, allowing to determine model parameters from data. However, all known mean-field methods fail at low temperatures with the emergence of multiple thermodynamic states. Here we show how clustering spin configurations can approximate these thermodynamic states, and how mean-field methods applied to thermodynamic states allow an efficient reconstruction of Ising models also at low temperatures.

cond-mat.dis-nn

Bethe-Peierls approximation and the inverse Ising model

We apply the Bethe-Peierls approximation to the problem of the inverse Ising model and show how the linear response relation leads to a simple method to reconstruct couplings and fields of the Ising model. This reconstruction is exact on tree graphs, yet its computational expense is comparable to other mean-field methods. We compare the performance of this method to the independent-pair, naive mean- field, Thouless-Anderson-Palmer approximations, the Sessak-Monasson expansion, and susceptibility propagation in the Cayley tree, SK-model and random graph with fixed connectivity. At low temperatures, Bethe reconstruction outperforms all these methods, while at high temperatures it is comparable to the best method available so far (Sessak-Monasson). The relationship between Bethe reconstruction and other mean- field methods is discussed.

cond-mat.dis-nn

Significance analysis and statistical mechanics: an application to clustering

This paper addresses the statistical significance of structures in random data: Given a set of vectors and a measure of mutual similarity, how likely does a subset of these vectors form a cluster with enhanced similarity among its elements? The computation of this cluster p-value for randomly distributed vectors is mapped onto a well-defined problem of statistical mechanics. We solve this problem analytically, establishing a connection between the physics of quenched disorder and multiple testing statistics in clustering and related problems. In an application to gene expression data, we find a remarkable link between the statistical significance of a cluster and the functional relationships between its genes.

q-bio.MN