SearcharxivSearch

arXiv subjects

Christopher R. Myers

Publications and source records attributed to Christopher R. Myers.

At least 19 recordsLinked to original sources

$Γ$-VAE: Curvature regularized variational autoencoders for uncovering emergent low dimensional geometric structure in high dimensional data

Natural systems with emergent behaviors often organize along low-dimensional subsets of high-dimensional spaces. For example, despite the tens of thousands of genes in the human genome, the principled study of genomics is fruitful because biological processes rely on coordinated organization that results in lower dimensional phenotypes. To uncover this organization, many nonlinear dimensionality reduction techniques have successfully embedded high-dimensional data into low-dimensional spaces by preserving local similarities between data points. However, the nonlinearities in these methods allow for too much curvature to preserve general trends across multiple non-neighboring data clusters, thereby limiting their interpretability and generalizability to out-of-distribution data. Here, we address both of these limitations by regularizing the curvature of manifolds generated by variational autoencoders, a process we coin ``$Γ$-VAE''. We demonstrate its utility using two example data sets: bulk RNA-seq from the The Cancer Genome Atlas (TCGA) and the Genotype Tissue Expression (GTEx); and single cell RNA-seq from a lineage tracing experiment in hematopoietic stem cell differentiation. We find that the resulting regularized manifolds identify mesoscale structure associated with different cancer cell types, and accurately re-embed tissues from completely unseen, out-of distribution cancers as if they were originally trained on them. Finally, we show that preserving long-range relationships to differentiated cells separates undifferentiated cells -- which have not yet specialized -- according to their eventual fate. Broadly, we anticipate that regularizing the curvature of generative models will enable more consistent, predictive, and generalizable models in any high-dimensional system with emergent low-dimensional behavior.

cs.LG

Aristotle Cloud Federation: Container Runtimes Technical Report

A National Science Foundation-sponsored container runtimes investigation was conducted by the Aristotle Cloud Federation to better understand the challenges of selecting and using Docker, Singularity, and X-Containers. The main goal of this investigation was to identify the "pain points" experienced by users when selecting and using containers for scientific research and to share lessons learned. Application performance characteristics are included in this report as well as user experiences with Kubernetes and container orchestration on cloud and HPC platforms. Scientists, research computing practitioners, and educators may find value in this report when considering the use and/or deployment of containers or when preparing students to meet the unique challenges of using containers in scientific research.

cs.DC

Emergent regularities and scaling in armed conflict data

Armed conflict exhibits regularities beyond known power law distributions of fatalities and duration over varying culture and geography. We systematically cluster conflict reports from a database of $10^5$ events from Africa spanning 20 years into conflict avalanches. Conflict profiles collapse over a range of scales. Duration, diameter, extent, fatalities, and report totals satisfy mutually consistent scaling relations captured with a model combining geographic spread and local conflict-site growth. The emergence of such social scaling laws hints at principles guiding conflict evolution.

physics.soc-ph

A scaling theory of armed conflict avalanches

Armed conflict data display scaling and universal dynamics in both social and physical properties like fatalities and geographic extent. We propose a randomly branching, armed-conflict model that relates multiple properties to one another in a way consistent with data. The model incorporates a fractal lattice on which conflict spreads, uniform dynamics driving conflict growth, and regional virulence that modulates local conflict intensity. The quantitative constraints on scaling and universal dynamics we use to develop our minimal model serve more generally as a set of constraints for other models for armed conflict dynamics. We show how this approach akin to thermodynamics imparts mechanistic intuition and unifies multiple conflict properties, giving insight into causation, prediction, and intervention timing.

physics.soc-ph

Fragility of reaction-diffusion models to competing advective processes

We study the coupling of a Fisher-Kolmogorov-Petrovsky-Piskunov (FKPP) equation to a separate, advection-only transport process. We find that the front dynamics can be described by an FKPP-like equation only at sufficiently fast diffusion or large coupling strength. For such parameter regimes, we find a mapping to an effective FKPP equation. We also find that FKPP equation is fragile with respect to the coupling to an advection-only mechanism, discover conditions when the front width diverges, and when front speed is insensitive to the coupling. At zero diffusion in this mean-field description, the downwind front speed goes to a finite value as the coupling goes to zero.

nlin.PS

Overshoot during phenotypic switching of cancer cell populations

The dynamics of tumor cell populations is hotly debated: do populations derive hierarchically from a subpopulation of cancer stem cells (CSCs), or are stochastic transitions that mutate differentiated cancer cells to CSCs important? Here we argue that regulation must also be important. We sort human melanoma cells using three distinct cancer stem cell (CSC) markers - CXCR6, CD271 and ABCG2 - and observe that the fraction of non-CSC-marked cells first overshoots to a higher level and then returns to the level of unsorted cells. This clearly indicates that the CSC population is homeostatically regulated. Combining experimental measurements with theoretical modeling and numerical simulations, we show that the population dynamics of cancer cells is associated with a complex miRNA network regulating the Wnt and PI3K pathways. Hence phenotypic switching is not stochastic, but is tightly regulated by the balance between positive and negative cells in the population. Reducing the fraction of CSCs below a threshold triggers massive phenotypic switching, suggesting that a therapeutic strategy based on CSC eradication is unlikely to succeed.

q-bio.CB

A robust and efficient method for estimating enzyme complex abundance and metabolic flux from expression data

A major theme in constraint-based modeling is unifying experimental data, such as biochemical information about the reactions that can occur in a system or the composition and localization of enzyme complexes, with highthroughput data including expression data, metabolomics, or DNA sequencing. The desired result is to increase predictive capability resulting in improved understanding of metabolism. The approach typically employed when only gene (or protein) intensities are available is the creation of tissue-specific models, which reduces the available reactions in an organism model, and does not provide an objective function for the estimation of fluxes, which is an important limitation in many modeling applications. We develop a method, flux assignment with LAD (least absolute deviation) convex objectives and normalization (FALCON), that employs metabolic network reconstructions along with expression data to estimate fluxes. In order to use such a method, accurate measures of enzyme complex abundance are needed, so we first present a new algorithm that addresses quantification of complex abundance. Our extensions to prior techniques include the capability to work with large models and significantly improved run-time performance even for smaller models, an improved analysis of enzyme complex formation logic, the ability to handle very large enzyme complex rules that may incorporate multiple isoforms, and depending on the model constraints, either maintained or significantly improved correlation with experimentally measured fluxes. FALCON has been implemented in MATLAB and ATS, and can be downloaded from: https://github.com/bbarker/FALCON. ATS is not required to compile the software, as intermediate C source code is available, and binaries are provided for Linux x86-64 systems. FALCON requires use of the COBRA Toolbox, also implemented in MATLAB.

q-bio.MN

Driven synchronization in random networks of oscillators

Synchronization is a universal phenomenon found in many non-equilibrium systems. Much recent interest in this area has overlapped with the study of complex networks, where a major focus is determining how a system's connectivity patterns affect the types of behavior that it can produce. Thus far, modeling efforts have focused on the tendency of networks of oscillators to mutually synchronize themselves, with less emphasis on the effects of external driving. In this work we discuss the interplay between mutual and driven synchronization in networks of phase oscillators of the Kuramoto type, and explore how the structure and emergence of such states depends on the underlying network topology for simple random networks with a given degree distribution. We find a variety of interesting dynamical behaviors, including bifurcations and bistability patterns that are qualitatively different for heterogeneous and homogeneous networks, and which are separated by a Takens-Bogdanov-Cusp singularity in the parameter region where the coupling strength between oscillators is weak. Our analysis is connected to the underlying dynamics of oscillator clusters for important states and transitions.

nlin.AO

You Can Run, You Can Hide: The Epidemiology and Statistical Mechanics of Zombies

We use a popular fictional disease, zombies, in order to introduce techniques used in modern epidemiology modelling, and ideas and techniques used in the numerical study of critical phenomena. We consider variants of zombie models, from fully connected continuous time dynamics to a full scale exact stochastic dynamic simulation of a zombie outbreak on the continental United States. Along the way, we offer a closed form analytical expression for the fully connected differential equation, and demonstrate that the single person per site two dimensional square lattice version of zombies lies in the percolation universality class. We end with a quantitative study of the full scale US outbreak, including the average susceptibility of different geographical regions.

q-bio.PE

Multiscale metabolic modeling of C4 plants: connecting nonlinear genome-scale models to leaf-scale metabolism in developing maize leaves

C4 plants, such as maize, concentrate carbon dioxide in a specialized compartment surrounding the veins of their leaves to improve the efficiency of carbon dioxide assimilation. Nonlinear relationships between carbon dioxide and oxygen levels and reaction rates are key to their physiology but cannot be handled with standard techniques of constraint-based metabolic modeling. We demonstrate that incorporating these relationships as constraints on reaction rates and solving the resulting nonlinear optimization problem yields realistic predictions of the response of C4 systems to environmental and biochemical perturbations. Using a new genome-scale reconstruction of maize metabolism, we build an 18000-reaction, nonlinearly constrained model describing mesophyll and bundle sheath cells in 15 segments of the developing maize leaf, interacting via metabolite exchange, and use RNA-seq and enzyme activity measurements to predict spatial variation in metabolic state by a novel method that optimizes correlation between fluxes and expression data. Though such correlations are known to be weak in general, here the predicted fluxes achieve high correlation with the data, successfully capture the experimentally observed base-to-tip transition between carbon-importing tissue and carbon-exporting tissue, and include a nonzero growth rate, in contrast to prior results from similar methods in other systems. We suggest that developmental gradients may be particularly suited to the inference of metabolic fluxes from expression data.

q-bio.MN

Sloppiness and Emergent Theories in Physics, Biology, and Beyond

Large scale models of physical phenomena demand the development of new statistical and computational tools in order to be effective. Many such models are `sloppy', i.e., exhibit behavior controlled by a relatively small number of parameter combinations. We review an information theoretic framework for analyzing sloppy models. This formalism is based on the Fisher Information Matrix, which we interpret as a Riemannian metric on a parameterized space of models. Distance in this space is a measure of how distinguishable two models are based on their predictions. Sloppy model manifolds are bounded with a hierarchy of widths and extrinsic curvatures. We show how the manifold boundary approximation can extract the simple, hidden theory from complicated sloppy models. We attribute the success of simple effective models in physics as likewise emerging from complicated processes exhibiting a low effective dimensionality. We discuss the ramifications and consequences of sloppy models for biochemistry and science more generally. We suggest that the reason our complex world is understandable is due to the same fundamental reason: simple theories of macroscopic behavior are hidden inside complicated microscopic processes.

cond-mat.stat-mech

Two-strain competition in quasi-neutral stochastic disease dynamics

We develop a new perturbation method for studying quasi-neutral competition in a broad class of stochastic competition models, and apply it to the analysis of fixation of competing strains in two epidemic models. The first model is a two-strain generalization of the stochastic Susceptible-Infected-Susceptible (SIS) model. Here we extend previous results due to Parsons and Quince (2007), Parsons et al (2008) and Lin, Kim and Doering (2012). The second model, a two-strain generalization of the stochastic Susceptible-Infected-Recovered (SIR) model with population turnover, has not been studied previously. In each of the two models, when the basic reproduction numbers of the two strains are identical, a system with an infinite population size approaches a point on the deterministic coexistence line (CL): a straight line of fixed points in the phase space of sub-population sizes. Shot noise drives one of the strain populations to fixation, and the other to extinction, on a time scale proportional to the total population size. Our perturbation method explicitly tracks the dynamics of the probability distribution of the sub-populations in the vicinity of the CL. We argue that, whereas the slow strain has a competitive advantage for mathematically "typical" initial conditions, it is the fast strain that is more likely to win in the important situation when a few infectives of both strains are introduced into a susceptible population.

cond-mat.stat-mech

Outbreak statistics and scaling laws for externally driven epidemics

Power-law scalings are ubiquitous to physical phenomena undergoing a continuous phase transition. The classic Susceptible-Infectious-Recovered (SIR) model of epidemics is one such example where the scaling behavior near a critical point has been studied extensively. In this system the distribution of outbreak sizes scales as $P(n) \sim n^{-3/2}$ at the critical point as the system size $N$ becomes infinite. The finite-size scaling laws for the outbreak size and duration are also well understood and characterized. In this work, we report scaling laws for a model with SIR structure coupled with a constant force of infection per susceptible, akin to a `reservoir forcing'. We find that the statistics of outbreaks in this system are fundamentally different than those in a simple SIR model. Instead of fixed exponents, all scaling laws exhibit tunable exponents parameterized by the dimensionless rate of external forcing. As the external driving rate approaches a critical value, the scale of the average outbreak size converges to that of the maximal size, and above the critical point, the scaling laws bifurcate into two regimes. Whereas a simple SIR process can only exhibit outbreaks of size $\mathcal{O}(N^{1/3})$ and $\mathcal{O}(N)$ depending on whether the system is at or above the epidemic threshold, a driven SIR process can exhibit a richer spectrum of outbreak sizes that scale as $O(N^ξ)$ where $ξ\in (0,1] \backslash \{2/3\}$ and $\mathcal{O}((N/\log N)^{2/3})$ at the multi-critical point.

q-bio.PE

The structure of infectious disease outbreaks across the animal-human interface

Despite the enormous relevance of zoonotic infections to world- wide public health, and despite much effort in modeling individual zoonoses, a fundamental understanding of the disease dynamics and the nature of outbreaks arising in such systems is still lacking. We introduce a simple stochastic model of susceptible-infected- recovered dynamics in a coupled animal-human metapopulation, and solve analytically for several important properties of the cou- pled outbreaks. At early timescales, we solve for the probability and time of spillover, and the disease prevalence in the animal population at spillover as a function of model parameters. At long times, we characterize the distribution of outbreak sizes and the critical threshold for a large human outbreak, both of which show a strong dependence on the basic reproduction number in the animal population. The coupling of animal and human infection dynamics has several crucial implications, most importantly al- lowing for the possibility of large human outbreaks even when human-to-human transmission is subcritical.

q-bio.PE

Epidemic fronts in complex networks with metapopulation structure

Infection dynamics have been studied extensively on complex networks, yielding insight into the effects of heterogeneity in contact patterns on disease spread. Somewhat separately, metapopulations have provided a paradigm for modeling systems with spatially extended and "patchy" organization. In this paper we expand on the use of multitype networks for combining these paradigms, such that simple contagion models can include complexity in the agent interactions and multiscale structure. We first present a generalization of the Volz-Miller mean-field approximation for Susceptible-Infected-Recovered (SIR) dynamics on multitype networks. We then use this technique to study the special case of epidemic fronts propagating on a one-dimensional lattice of interconnected networks - representing a simple chain of coupled population centers - as a necessary first step in understanding how macro-scale disease spread depends on micro-scale topology. Using the formalism of front propagation into unstable states, we derive the effective transport coefficients of the linear spreading: asymptotic speed, characteristic wavelength, and diffusion coefficient for the leading edge of the pulled fronts, and analyze their dependence on the underlying graph structure. We also derive the epidemic threshold for the system and study the front profile for various network configurations. To our knowledge, this is the first such application of front propagation concepts to random network models.

physics.soc-ph

The Context Sensitivity Problem in Biological Sequence Segmentation

In this paper, we describe the context sensitivity problem encountered in partitioning a heterogeneous biological sequence into statistically homogeneous segments. After showing signatures of the problem in the bacterial genomes of Escherichia coli K-12 MG1655 and Pseudomonas syringae DC3000, when these are segmented using two entropic segmentation schemes, we clarify the contextual origins of these signatures through mean-field analyses of the segmentation schemes. Finally, we explain why we believe all sequence segmentation schems are plagued by the context sensitivity problem.

q-bio.GN

Extending the Recursive Jensen-Shannon Segmentation of Biological Sequences

In this paper, we extend a previously developed recursive entropic segmentation scheme for applications to biological sequences. Instead of Bernoulli chains, we model the statistically stationary segments in a biological sequence as Markov chains, and define a generalized Jensen-Shannon divergence for distinguishing between two Markov chains. We then undertake a mean-field analysis, based on which we identify pitfalls associated with the recursive Jensen-Shannon segmentation scheme. Following this, we explain the need for segmentation optimization, and describe two local optimization schemes for improving the positions of domain walls discovered at each recursion stage. We also develop a new termination criterion for recursive Jensen-Shannon segmentation based on the strength of statistical fluctuations up to a minimum statistically reliable segment length, avoiding the need for unrealistic null and alternative segment models of the target sequence. Finally, we compare the extended scheme against the original scheme by recursively segmenting the Escherichia coli K-12 MG1655 genome.

q-bio.GN

Variational method for estimating the rate of convergence of Markov Chain Monte Carlo algorithms

We demonstrate the use of a variational method to determine a quantitative lower bound on the rate of convergence of Markov Chain Monte Carlo (MCMC) algorithms as a function of the target density and proposal density. The bound relies on approximating the second largest eigenvalue in the spectrum of the MCMC operator using a variational principle and the approach is applicable to problems with continuous state spaces. We apply the method to one dimensional examples with Gaussian and quartic target densities, and we contrast the performance of the Random Walk Metropolis-Hastings (RWMH) algorithm with a ``smart'' variant that incorporates gradient information into the trial moves. We find that the variational method agrees quite closely with numerical simulations. We also see that the smart MCMC algorithm often fails to converge geometrically in the tails of the target density except in the simplest case we examine, and even then care must be taken to choose the appropriate scaling of the deterministic and random parts of the proposed moves. Again, this calls into question the utility of smart MCMC in more complex problems. Finally, we apply the same method to approximate the rate of convergence in multidimensional Gaussian problems with and without importance sampling. There we demonstrate the necessity of importance sampling for target densities which depend on variables with a wide range of scales.

physics.data-an