SearcharxivSearch

arXiv subjects

Justin B. Kinney

Publications and source records attributed to Justin B. Kinney.

14 recordsLinked to original sources

On learning functions over biological sequence space: relating Gaussian process priors, regularization, and gauge fixing

Mappings from biological sequences (DNA, RNA, protein) to quantitative measures of sequence functionality play an important role in contemporary biology. We are interested in the related tasks of (i) inferring predictive sequence-to-function maps and (ii) decomposing sequence-function maps to elucidate the contributions of individual subsequences. Because each sequence-function map can be written as a weighted sum over subsequences in multiple ways, meaningfully interpreting these weights requires ``gauge-fixing,'' i.e., defining a unique representation for each map. Recent work has established that most existing gauge-fixed representations arise as the unique solutions to $L_2$-regularized regression in an overparameterized ``weight space'' where the choice of regularizer defines the gauge. Here, we establish the relationship between regularized regression in overparameterized weight space and Gaussian process approaches that operate in ``function space,'' i.e.~the space of all real-valued functions on a finite set of sequences. We disentangle how weight space regularizers both impose an implicit prior on the learned function and restrict the optimal weights to a particular gauge. We show how to construct regularizers that correspond to arbitrary explicit Gaussian process priors combined with a wide variety of gauges and characterize the implicit function space priors associated with the most common weight space regularizers. Finally, we derive the posterior distribution of a broad class of sequence-to-function statistics, including gauge-fixed weights and multiple systems for expressing higher-order epistatic coefficients. We show that such distributions can be efficiently computed for product-kernel priors using a kernel trick.

cs.LG

Algebraic and diagrammatic methods for the rule-based modeling of multi-particle complexes

The formation, dissolution, and dynamics of multi-particle complexes is of fundamental interest in the study of stochastic chemical systems. In 1976, Masao Doi introduced a Fock space formalism for modeling classical particles. Doi's formalism, however, does not support the assembly of multiple particles into complexes. Starting in the 2000's, multiple groups developed rule-based methods for computationally simulating biochemical systems involving large macromolecular complexes. However, these methods are based on graph-rewriting rules and/or process algebras that are mathematically disconnected from the statistical physics methods generally used to analyze equilibrium and nonequilibrium systems. Here we bridge these two approaches by introducing an operator algebra for the rule-based modeling of multi-particle complexes. Our formalism is based on a Fock space that supports not only the creation and annihilation of classical particles, but also the assembly of multiple particles into complexes, as well as the disassembly of complexes into their components. Rules are specified by algebraic operators that act on particles through a manifestation of Wick's theorem. We further describe diagrammatic methods that facilitate rule specification and analytic calculations. We demonstrate our formalism on systems in and out of thermal equilibrium, and for nonequilibrium systems we present a stochastic simulation algorithm based on our formalism. The results provide a unified approach to the mathematical and computational study of stochastic chemical systems in which multi-particle complexes play an important role.

physics.bio-ph

Biophysical models of cis-regulation as interpretable neural networks

The adoption of deep learning techniques in genomics has been hindered by the difficulty of mechanistically interpreting the models that these techniques produce. In recent years, a variety of post-hoc attribution methods have been proposed for addressing this neural network interpretability problem in the context of gene regulation. Here we describe a complementary way of approaching this problem. Our strategy is based on the observation that two large classes of biophysical models of cis-regulatory mechanisms can be expressed as deep neural networks in which nodes and weights have explicit physiochemical interpretations. We also demonstrate how such biophysical networks can be rapidly inferred, using modern deep learning frameworks, from the data produced by certain types of massively parallel reporter assays (MPRAs). These results suggest a scalable strategy for using MPRAs to systematically characterize the biophysical basis of gene regulation in a wide range of biological contexts. They also highlight gene regulation as a promising venue for the development of scientifically interpretable approaches to deep learning.

q-bio.MN

Deciphering the regulatory genome of $\textit{Escherichia coli}$, one hundred promoters at a time

Advances in DNA sequencing have revolutionized our ability to read genomes. However, even in the most well-studied of organisms, the bacterium ${\it Escherichia coli}$, for $\approx$ 65$\%$ of the promoters we remain completely ignorant of their regulation. Until we have cracked this regulatory Rosetta Stone, efforts to read and write genomes will remain haphazard. We introduce a new method (Reg-Seq) linking a massively-parallel reporter assay and mass spectrometry to produce a base pair resolution dissection of more than 100 promoters in ${\it E. coli}$ in 12 different growth conditions. First, we show that our method recapitulates regulatory information from known sequences. Then, we examine the regulatory architectures for more than 80 promoters in the ${\it E. coli}$ genome which previously had no known regulation. In many cases, we also identify which transcription factors mediate their regulation. The method introduced here clears a path for fully characterizing the regulatory genome of model organisms, with the potential of moving on to an array of other microbes of ecological and medical relevance.

q-bio.GN

Density estimation on small datasets

How might a smooth probability distribution be estimated, with accurately quantified uncertainty, from a limited amount of sampled data? Here we describe a field-theoretic approach that addresses this problem remarkably well in one dimension, providing an exact nonparametric Bayesian posterior without relying on tunable parameters or large-data approximations. Strong non-Gaussian constraints, which require a non-perturbative treatment, are found to play a major role in reducing distribution uncertainty. A software implementation of this method is provided.

physics.data-an

Physical epistatic landscape of antibody binding affinity

Affinity maturation produces antibodies that bind antigens with high specificity by accumulating mutations in the antibody sequence. Mapping out the antibody-antigen affinity landscape can give us insight into the accessible paths during this rapid evolutionary process. By developing a carefully controlled null model for noninteracting mutations, we characterized epistasis in affinity measurements of a large library of antibody variants obtained by Tite-Seq, a recently introduced Deep Mutational Scan method yielding physical values of the binding constant. We show that representing affinity as the binding free energy minimizes epistasis. Yet, we find that epistatically interacting sites contribute substantially to binding. In addition to negative epistasis, we report a large amount of beneficial epistasis, enlarging the space of high-affinity antibodies as well as their mutational accessibility. These properties suggest that the degeneracy of antibody sequences that can bind a given antigen is enhanced by epistasis - an important property for vaccine design.

q-bio.PE

Measuring the sequence-affinity landscape of antibodies with massively parallel titration curves

Despite the central role that antibodies play in the adaptive immune system and in biotechnology, much remains unknown about the quantitative relationship between an antibody's amino acid sequence and its antigen binding affinity. Here we describe a new experimental approach, called Tite-Seq, that is capable of measuring binding titration curves and corresponding affinities for thousands of variant antibodies in parallel. The measurement of titration curves eliminates the confounding effects of antibody expression and stability that arise in standard deep mutational scanning assays. We demonstrate Tite-Seq on the CDR1H and CDR3H regions of a well-studied scFv antibody. Our data shed light on the structural basis for antigen binding affinity and suggests a role for secondary CDR loops in establishing antibody stability. Tite-Seq fills a large gap in the ability to measure critical aspects of the adaptive immune system, and can be readily used for studying sequence-affinity landscapes in other protein systems.

q-bio.QM

Modeling multi-particle complexes in stochastic chemical systems

Large complexes of classical particles play central roles in biology, in polymer physics, and in other disciplines. However, physics currently lacks mathematical methods for describing such complexes in terms of component particles, interaction energies, and assembly rules. Here we describe a Fock space structure that addresses this need, as well as diagrammatic methods that facilitate the use of this formalism. These methods can dramatically simplify the equations governing both equilibrium and non-equilibrium stochastic chemical systems. A mathematical relationship between the set of all complexes and a list of rules for complex assembly is also identified.

q-bio.QM

Learning quantitative sequence-function relationships from massively parallel experiments

A fundamental aspect of biological information processing is the ubiquity of sequence-function relationships -- functions that map the sequence of DNA, RNA, or protein to a biochemically relevant activity. Most sequence-function relationships in biology are quantitative, but only recently have experimental techniques for effectively measuring these relationships been developed. The advent of such "massively parallel" experiments presents an exciting opportunity for the concepts and methods of statistical physics to inform the study of biological systems. After reviewing these recent experimental advances, we focus on the problem of how to infer parametric models of sequence-function relationships from the data produced by these experiments. Specifically, we retrace and extend recent theoretical work showing that inference based on mutual information, not the standard likelihood-based approach, is often necessary for accurately learning the parameters of these models. Closely connected with this result is the emergence of "diffeomorphic modes" -- directions in parameter space that are far less constrained by data than likelihood-based inference would suggest. Analogous to Goldstone modes in physics, diffeomorphic modes arise from an arbitrarily broken symmetry of the inference problem. An analytically tractable model of a massively parallel experiment is then described, providing an explicit demonstration of these fundamental aspects of statistical inference. This paper concludes with an outlook on the theoretical and computational challenges currently facing studies of quantitative sequence-function relationships.

q-bio.QM

Unification of field theory and maximum entropy methods for learning probability densities

The need to estimate smooth probability distributions (a.k.a. probability densities) from finite sampled data is ubiquitous in science. Many approaches to this problem have been described, but none is yet regarded as providing a definitive solution. Maximum entropy estimation and Bayesian field theory are two such approaches. Both have origins in statistical physics, but the relationship between them has remained unclear. Here I unify these two methods by showing that every maximum entropy density estimate can be recovered in the infinite smoothness limit of an appropriate Bayesian field theory. I also show that Bayesian field theory estimation can be performed without imposing any boundary conditions on candidate densities, and that the infinite smoothness limit of these theories recovers the most common types of maximum entropy estimates. Bayesian field theory is thus seen to provide a natural test of the validity of the maximum entropy null hypothesis. Bayesian field theory also returns a lower entropy density estimate when the maximum entropy hypothesis is falsified. The computations necessary for this approach can be performed rapidly for one-dimensional data, and software for doing this is provided. Based on these results, I argue that Bayesian field theory is poised to provide a definitive solution to the density estimation problem in one dimension.

physics.data-an

Rapid and deterministic estimation of probability densities using scale-free field theories

The question of how best to estimate a continuous probability density from finite data is an intriguing open problem at the interface of statistics and physics. Previous work has argued that this problem can be addressed in a natural way using methods from statistical field theory. Here I describe new results that allow this field-theoretic approach to be rapidly and deterministically computed in low dimensions, making it practical for use in day-to-day data analysis. Importantly, this approach does not impose a privileged length scale for smoothness of the inferred probability density, but rather learns a natural length scale from the data due to the tradeoff between goodness-of-fit and an Occam factor. Open source software implementing this method in one and two dimensions is provided.

physics.data-an

Parametric inference in the large data limit using maximally informative models

Motivated by data-rich experiments in transcriptional regulation and sensory neuroscience, we consider the following general problem in statistical inference. When exposed to a high-dimensional signal S, a system of interest computes a representation R of that signal which is then observed through a noisy measurement M. From a large number of signals and measurements, we wish to infer the "filter" that maps S to R. However, the standard method for solving such problems, likelihood-based inference, requires perfect a priori knowledge of the "noise function" mapping R to M. In practice such noise functions are usually known only approximately, if at all, and using an incorrect noise function will typically bias the inferred filter. Here we show that, in the large data limit, this need for a pre-characterized noise function can be circumvented by searching for filters that instead maximize the mutual information I[M;R] between observed measurements and predicted representations. Moreover, if the correct filter lies within the space of filters being explored, maximizing mutual information becomes equivalent to simultaneously maximizing every dependence measure that satisfies the Data Processing Inequality. It is important to note that maximizing mutual information will typically leave a small number of directions in parameter space unconstrained. We term these directions "diffeomorphic modes" and present an equation that allows these modes to be derived systematically. The presence of diffeomorphic modes reflects a fundamental and nontrivial substructure within parameter space, one that is obscured by standard likelihood-based inference.

q-bio.QM

Equitability, mutual information, and the maximal information coefficient

Reshef et al. recently proposed a new statistical measure, the "maximal information coefficient" (MIC), for quantifying arbitrary dependencies between pairs of stochastic quantities. MIC is based on mutual information, a fundamental quantity in information theory that is widely understood to serve this need. MIC, however, is not an estimate of mutual information. Indeed, it was claimed that MIC possesses a desirable mathematical property called "equitability" that mutual information lacks. This was not proven; instead it was argued solely through the analysis of simulated data. Here we show that this claim, in fact, is incorrect. First we offer mathematical proof that no (non-trivial) dependence measure satisfies the definition of equitability proposed by Reshef et al.. We then propose a self-consistent and more general definition of equitability that follows naturally from the Data Processing Inequality. Mutual information satisfies this new definition of equitability while MIC does not. Finally, we show that the simulation evidence offered by Reshef et al. was artifactual. We conclude that estimating mutual information is not only practical for many real-world applications, but also provides a natural solution to the problem of quantifying associations in large data sets.

q-bio.QM

The r-modes in accreting neutron stars with magneto-viscous boundary layers

We explore the dynamics of the r-modes in accreting neutron stars in two ways. First, we explore how dissipation in the magneto-viscous boundary layer (MVBL) at the crust-core interface governs the damping of r-mode perturbations in the fluid interior. Two models are considered: one assuming an ordinary-fluid interior, the other taking the core to consist of superfluid neutrons, type II superconducting protons, and normal electrons. We show, within our approximations, that no solution to the magnetohydrodynamic equations exists in the superfluid model when both the neutron and proton vortices are pinned. However, if just one species of vortex is pinned, we can find solutions. When the neutron vortices are pinned and the proton vortices are unpinned there is much more dissipation than in the ordinary-fluid model, unless the pinning is weak. When the proton vortices are pinned and the neutron vortices are unpinned the dissipation is comparable or slightly less than that for the ordinary-fluid model, even when the pinning is strong. We also find in the superfluid model that relatively weak radial magnetic fields ~ 10^9 G (10^8 K / T)^2 greatly affect the MVBL, though the effects of mutual friction tend to counteract the magnetic effects. Second, we evolve our two models in time, accounting for accretion, and explore how the magnetic field strength, the r-mode saturation amplitude, and the accretion rate affect the cyclic evolution of these stars. If the r-modes control the spin cycles of accreting neutron stars we find that magnetic fields can affect the clustering of the spin frequencies of low mass x-ray binaries (LMXBs) and the fraction of these that are currently emitting gravitational waves.

gr-qc