SearcharxivSearch

arXiv subjects

Ethan Levien

Publications and source records attributed to Ethan Levien.

8 recordsLinked to original sources

EM-based iterations for multiple instance learning on a query-value model

In multiple instance regression (MIR) data are organized into bags (collections of instances in feature space) and the goal is to learn a mapping that assigns labels to bags. A typical assumption is that there is a so-called concept point in feature space, the proximity to which dictates the bag label. Motivated by modern MIR architectures which are based on attention, we study a softmax model that decouples the concept point and the labeling scheme. The two are respectively determined by a query direction and a value direction value in feature space. This problem isolates a basic challenge of learning both the query and value vectors from bag-level supervision. From this model we derive a parametric family of iterations in the noiseless limit, which generalizes a method known as the EM-DD algorithm. We then derive concentration results for the MLE estimators of the query and value vectors obtained from a random selection of instances. Our result for the value vector shows that a single random initialization of the value vector already points in the correct direction on average, so that a polynomial (in the number of instances per bag and the feature dimension) number of bags is enough for the EM algorithm to converge in $O(1)$ steps with high probability. A key aspect of this analysis is the interplay between concentration of empirical covariance matrices and extremal statistics arising from the selection rule.

math.ST

Coarse-graining and stochastic oscillations in a phenomenological model of cell-size homeostasis

Within a continuous-time, stochastic model of single-cell size homeostasis, we study how the structure of feedback from size to growth rates and cell-cycle progression shapes overall size dynamics, both within and across cell cycles. We focus on a model in which the feedback from cell size to these other processes occurs only through the size deviations, defined as the difference between the absolute size and the progression through the cell cycle. In a linear regime of this model, the dynamics reduce to a stochastically forced simple harmonic oscillator, yielding closed-form expressions for mother-daughter size correlations. We compare these to the higher order regression coefficients that measure the size memory over many generations. Our analysis reveals how the interplay between cell-cycle timing and intrinsic fluctuations shapes the apparent coarse-grained size control strategy, and in-particular, that coarse-grained correlations may not reflect the mechanistic feedback structure. We compare this model to a more commonly used approach where the coarse-grained dynamics are hard-coded into the model; hence, the first order autoregressive model for sizes is a perfect description of the size dynamics and therefore more accurately reflects the feedback structure.

q-bio.CB

Size-structured populations with growth fluctuations: Feynman--Kac formula and decoupling

We study a size-structured population model in which individual cells grow at a rate determined by a fluctuating internal variable (e.g., gene expression levels). Many previous models of phenotypically heterogeneous populations can be viewed as special cases of this model, and it has previously been observed that the internal variable decouples from cell size under certain conditions. In this work, we generalize these results and connect them to the Feynman-Kac formula, which yields relationships between the lineage dynamics and population distribution in branching processes. To this end, we derive conditions for decoupling, both in the lineage and population ensemble. When decoupling occurs in both ensembles, the size dynamics can be transformed, via a random time change, into a growth-homogeneous process, and expectations can be evaluated through an exponential tilting procedure that follows from the Feynman-Kac formula. We further characterize weaker, ensemble-specific forms of decoupling that hold in either the lineage or the population ensemble, but not both. We provide a more general interpretation of tilted expectations in terms of the mass-weighted phenotype distribution

cond-mat.stat-mech

Extremal events dictate population growth rate inference

Recent methods have been developed to map single-cell lineage statistics to population growth. Because population growth selects for exponentially rare phenotypes, these methods inherently depend on sampling large deviations from finite data, which introduces systematic errors. A comprehensive understanding of these errors in the context of finite data remains elusive. To address this gap, we study the error in growth rate estimates across different models. We show that under the usual bias-variance decomposition, the bias can be decomposed into a finite-time bias and nonlinear averaging bias. We demonstrate that finite-time bias, which dominates at short times, can be mitigated by fitting its monotonic behavior. In contrast, at longer times, nonlinear averaging bias becomes the predominant source of error, leading to a phase transition. This transition can be understood through the Random Energy Model, a mean-field model of disordered systems, where a few lineages dominate the estimator. Applying these methods to experimental data demonstrates that correcting for biases in lineage-based approaches yields consistent results for the long-term growth rate across multiple methods and enables the reverse-engineering of dynamic models. This new framework provides a quantitative understanding of growth rate estimators, clarifies the conditions under which they can be effectively applied to finite data, and introduces model-free approaches for studying the connections between physiology and cell growth.

cond-mat.stat-mech

Coalescent processes emerging from large deviations

The classical model for the genealogies of a neutrally evolving population in a fixed environment is due to Kingman. Kingman's coalescent process, which produces a binary tree, universally emerges from many microscopic models in which the variance in the number of offspring is finite. It is understood that power-law offspring distributions with infinite variance can result in a very different type of coalescent structure with merging of more than two lineages. Here we investigate the regime where the variance of the offspring distribution is finite but comparable to the population size. This is achieved by studying a model in which the log offspring sizes have a stretched exponential form. Such offspring distributions are motivated by biology, where they emerge from a toy model of growth in a heterogenous environment, but also mathematics and statistical physics, where limit theorems and phase transitions for sums over random exponentials have received considerable attention due to their appearance in the partition function of Derrida's Random Energy Model (REM). We find that the limit coalescent is a $β$-coalescent -- a previously studied model emerging from evolutionary dynamics models with heavy-tailed offspring distributions. We also discuss the connection to previous results on the REM.

q-bio.PE

Non-genetic variability: survival strategy or nuisance?

The observation that phenotypic variability is ubiquitous in isogenic populations has led to a multitude of experimental and theoretical studies seeking to probe the causes and consequences of this variability. Whether it be in the context of antibiotic treatments or exponential growth in constant environments, non-genetic variability has shown to have significant effects on population dynamics. Here, we review research that elucidates the relationship between cell-to-cell variability and population dynamics. After summarizing the relevant experimental observations, we discuss models of bet-hedging and phenotypic switching. In the context of these models, we discuss how switching between phenotypes at the single-cell level can help populations survive in uncertain environments. Next, we review more fine-grained models of phenotypic variability where the relationship between single-cell growth rates, generation times and cell sizes is explicitly considered. Variability in these traits can have significant effects on the population dynamics, even in a constant environment. We show how these effects can be highly sensitive to the underlying model assumptions. We close by discussing a number of open questions, such as how environmental and intrinsic variability interact and what the role of non-genetic variability in evolutionary dynamics is.

q-bio.PE

A large deviation principle linking lineage statistics to fitness in microbial populations

In exponentially proliferating populations of microbes, the population typically doubles at a rate less than the average doubling time of a single-cell due to variability at the single-cell level. It is known that the distribution of generation times obtained from a single lineage is, in general, insufficient to determine a population's growth rate. Is there an explicit relationship between observables obtained from a single lineage and the population growth rate? We show that a population's growth rate can be represented in terms of averages over isolated lineages. This lineage representation is related to a large deviation principle that is a generic feature of exponentially proliferating populations. Due to the large deviation structure of growing populations, the number of lineages needed to obtain an accurate estimate of the growth rate depends exponentially on the duration of the lineages, leading to a non-monotonic convergence of the estimate, which we verify in both synthetic and experimental data sets.

q-bio.PE

Coupling sample paths to the partial thermodynamic limit in stochastic chemical reaction networks

Many biochemical systems appearing in applications have a multiscale structure so that they converge to piecewise deterministic Markov processes in a thermodynamic limit. The statistics of the piecewise deterministic process can be obtained much more efficiently than those of the exact process. We explore the possibility of coupling sample paths of the exact model to the piecewise deterministic process in order to reduce the variance of their difference. We then apply this coupling to reduce the computational complexity of a Monte Carlo estimator. In addition to rigorous results concerning the asymptotic computational complexity of the Monte Carlo estimator, numerical simulations are performed on some simple biological models confirming that computational gains are made.

physics.comp-ph