SearcharxivSearch

arXiv subjects

Hong-Li Zeng

Publications and source records attributed to Hong-Li Zeng.

16 recordsLinked to original sources

Fitness Inference in Presence of Migrations between Coupled Evolving Populations

The phase of Quasi-Linkage Equilibrium (QLE) in evolutionary populations is analogous to the thermal equilibrium state in statistical mechanics, a concept pioneered by Kimura in 1965 for two-locus two-allele models. QLE describes a stationary state maintained by the interplay of selection, mutation, recombination and genetic drift. Here we extend QLE theory to populations connected by migration, a fundamental evolutionary force that couples the evolutionary dynamics of interacting subpopulations. Specifically, we examine two populations interacting via symmetric or asymmetric migration while evolving under multi-locus selection. Using whole-genome time-series data generated through FFPopSim, we demonstrate that the QLE phase persists under conditions of sufficiently low migration rates. In this regime, we derive analytical inference relations that allow for the accurate and quantitative estimation of both additive fitness and epistatic interactions.

q-bio.PE

Classification of SARS-CoV-2 Variants through The Epistatical Circos Plots with Convolutional Neural Networks

The COVID-19 pandemic has profoundly affected global health, driven by the remarkable transmissibility and mutational adaptability of the SARS-CoV-2 virus. Although five variants of concern, Alpha, Beta, Gamma, Delta, and Omicron, have been identified, the classification task in this study is formulated using four classes: Alpha, Delta, Omicron, and Else, reflecting the sequence availability and temporal coverage of the dataset. Here, we develop an integrative framework that combines direct coupling analysis (DCA), Circos-based visualization, and convolutional neural networks (CNNs) to characterize lineage-specific epistatic signatures from large-scale SARS-CoV-2 genomic sequences. DCA-inferred pairwise mutational couplings were transformed into Circos images, which were then used as inputs for CNN-based classification models. The proposed framework achieved robust variant classification, with the best-performing model reaching a weighted-average F1-score of $98.68\pm 0.75\%$ and an AUC close to 1.

q-bio.GN

Fitness inference tested by in silico population genetics

We consider populations evolving according to natural selection, mutation, and recombination, and assume that the genomes of all or a representative selection of individuals are known. We pose the problem if it is possible to infer fitness parameters and genotype fitness order from such data. We tested this hypothesis in simulated populations. We delineate parameter ranges where this is possible and other ranges where it is not.Our work provides a framework for determining when fitness inference is feasible from population-wide, whole-genome, time-stratified data and highlights settings where it is not. We give a brief survey of biological model organisms and human pathogens that fit into this framework.

q-bio.PE

Two fitness inference schemes compared using allele frequencies from 1,068,391 sequences sampled in the UK during the COVID-19 pandemic

Throughout the course of the SARS-CoV-2 pandemic, genetic variation has contributed to the spread and persistence of the virus. For example, various mutations have allowed SARS-CoV-2 to escape antibody neutralization or to bind more strongly to the receptors that it uses to enter human cells. Here, we compared two methods that estimate the fitness effects of viral mutations using the abundant sequence data gathered over the course of the pandemic. Both approaches are grounded in population genetics theory but with different assumptions. One approach, tQLE, features an epistatic fitness landscape and assumes that alleles are nearly in linkage equilibrium. Another approach, MPL, assumes a simple, additive fitness landscape, but allows for any level of correlation between alleles. We characterized differences in the distributions of fitness values inferred by each approach and in the ranks of fitness values that they assign to sequences across time. We find that in a large fraction of weeks the two methods are in good agreement as to their top-ranked sequences, i.e., as to which sequences observed that week are most fit. We also find that agreement between ranking of sequences varies with genetic unimodality in the population in a given week.

q-bio.PE

Statistical Genetics in and out of Quasi-Linkage Equilibrium (Extended)

This review is about statistical genetics, an interdisciplinary topic between statistical physics and population biology. The focus is on the phase of quasi-linkage equilibrium (QLE). Our goals here are to clarify under which conditions the QLE phase can be expected to hold in population biology and how the stability of the QLE phase is lost. The QLE state, which has many similarities to a thermal equilibrium state in statistical mechanics, was discovered by M Kimura for a two-locus two-allele model, and was extended and generalized to the global genome scale by (Neher and Shraiman, 2011). What we will refer to as the Kimura-Neher-Shraiman (KNS) theory describes a population evolving due to the mutations, recombination, natural selection and possibly genetic drift. A QLE phase exists at sufficiently high recombination rate and/or mutation rates with respect to selection strength. We show how in QLE it is possible to infer the epistatic parameters of the fitness function from the knowledge of the (dynamical) distribution of genotypes in a population. We further consider the breakdown of the QLE regime for high enough selection strength. We review recent results for the selection-mutation and selection-recombination dynamics. Finally, we identify and characterize a new phase which we call the non-random coexistence (NRC) where variability persists in the population without either fixating or disappearing.

q-bio.PE

Temporal epistasis inference from more than 3,500,000 SARS-CoV-2 Genomic Sequences

We use Direct Coupling Analysis (DCA) to determine epistatic interactions between loci of variability of the SARS-CoV-2 virus, segmenting genomes by month of sampling. We use full-length, high-quality genomes from the GISAID repository up to October 2021, in total over 3,500,000 genomes. We find that DCA terms are more stable over time than correlations, but nevertheless change over time as mutations disappear from the global population or reach fixation. Correlations are enriched for phylogenetic effects, and in particularly statistical dependencies at short genomic distances, while DCA brings out links at longer genomic distance. We discuss the validity of a DCA analysis under these conditions in terms of a transient Quasi-Linkage Equilibrium state. We identify putative epistatic interaction mutations involving loci in Spike.

q-bio.GN

Mutation frequency time series reveal complex mixtures of clones in the world-wide SARS-CoV-2 viral population

We compute the allele frequencies of the alpha (B.1.1.7), beta (B.1.351) and delta (B.167.2) variants of SARS-CoV-2 from almost two million genome sequences on the GISAID repository. We find that the frequencies of a majority of the defining mutations in alpha rose towards the end of 2020 but drifted apart during spring 2021, a similar pattern being followed by delta during summer of 2021. For beta we find a more complex scenario with frequencies of some mutations rising and some remaining close to zero. Our results point to that what is generally reported as single variants is in fact a collection of variants with different genetic characteristics. For all three variants we further find some alleles with a clearly deviating time series.

q-bio.PE

Inferring epistasis from genomic data with comparable mutation and outcrossing rate

We consider a population evolving due to mutation, selection and recombination, where selection includes single-locus terms (additive fitness) and two-loci terms (pairwise epistatic fitness). We further consider the problem of inferring fitness in the evolutionary dynamics from one or several snap-shots of the distribution of genotypes in the population. In the recent literature this has been done by applying the Quasi-Linkage Equilibrium (QLE) regime first obtained by Kimura in the limit of high recombination. Here we show that the approach also works in the interesting regime where the effects of mutations are comparable to or larger than recombination. This leads to a modified main epistatic fitness inference formula where the rates of mutation and recombination occur together. We also derive this formula using by a previously developed Gaussian closure that formally remains valid when recombination is absent. The findings are validated through numerical simulations.

q-bio.PE

Global analysis of more than 50,000 SARS-Cov-2 genomes reveals epistasis between 8 viral genes

Genome-wide epistasis analysis is a powerful tool to infer gene interactions, which can guide drug and vaccine development and lead to a deeper understanding of microbial pathogenesis. We have considered all complete SARS-CoV-2 genomes deposited in the GISAID repository until \textbf{four} different cut-off dates, and used Direct Coupling Analysis together with an assumption of Quasi-Linkage Equilibrium to infer epistatic contributions to fitness from polymorphic loci. We find \textbf{eight} interactions, of which three between pairs where one locus lies in gene ORF3a, both loci holding non-synonymous mutations. We also find interactions between two loci in gene nsp13, both holding non-synonymous mutations, and four interactions involving one locus holding a synonymous mutation. Altogether we infer interactions between loci in viral genes ORF3a and nsp2, nsp12 and nsp6, between ORF8 and nsp4, and between loci in genes nsp2, nsp13 and nsp14. The paper opens the prospect to use prominent epistatically linked pairs as a starting point to search for combinatorial weaknesses of recombinant viral pathogens.

q-bio.QM

Longitudinal Support Vector Machines for High Dimensional Time Series

We consider the problem of learning a classifier from observed functional data. Here, each data-point takes the form of a single time-series and contains numerous features. Assuming that each such series comes with a binary label, the problem of learning to predict the label of a new coming time-series is considered. Hereto, the notion of {\em margin} underlying the classical support vector machine is extended to the continuous version for such data. The longitudinal support vector machine is also a convex optimization problem and its dual form is derived as well. Empirical results for specified cases with significance tests indicate the efficacy of this innovative algorithm for analyzing such long-term multivariate data.

cs.LG

Network reconstruction from asynchronously updated evolutionary game

The interactions between players of prisoner's dilemma (PD) game are reconstructed with evolutionary game data. All participants play the game with their counterparts and gain corresponding rewards during each round of the game. However, their strategies are updated asynchronously during the evolutionary PD game. Two inference methods of the interactions between players are derived with naive mean-field (nMF) approximation and maximum log-likelihood estimation (MLE) respectively. The two methods are tested numerically also for fully connected asymmetric Sherrington-Kirkpatrick (SK) models, varying the data length, asymmetric degree, payoff and system noise (coupling strength). We find that the reconstruction mean square error (MSE) of MLE method is proportional to the inverse of data length and typically half (benefit from the extra information of update times) of that by nMF. Both methods are robust to the asymmetric degree but works better for large payoff. Compared with MLE, nMF is more sensitive to the couplings strength which prefers weak couplings.

physics.soc-ph

Inverse Ising techniques to infer underlying mechanisms from data

As a problem in data science the inverse Ising (or Potts) problem is to infer the parameters of a Gibbs-Boltzmann distributions of an Ising (or Potts) model from samples drawn from that distribution. The algorithmic and computational interest stems from the fact that this inference task cannot be done efficiently by the maximum likelihood criterion, since the normalizing constant of the distribution (the partition function) can not be calculated exactly and efficiently. The practical interest on the other hand flows from several outstanding applications, of which the most well known has been predicting spatial contacts in protein structures from tables of homologous protein sequences. Most applications to date have been to data that has been produced by a dynamical process which, as far as it is known, cannot be expected to satisfy detailed balance. There is therefore no a priori reason to expect the distribution to be of the Gibbs-Boltzmann type, and no a priori reason to expect that inverse Ising (or Potts) techniques should yield useful information. In this review we discuss two types of problems where progress nevertheless can be made. We find that depending on model parameters there are phases where, in fact, the distribution is close to Gibbs-Boltzmann distribution, a non-equilibrium nature of the under-lying dynamics notwithstanding. We also discuss the relation between inferred Ising model parameters and parameters of the underlying dynamics.

stat.OT

Inferring genetic fitness from genomic data

The genetic composition of a naturally developing population is considered as due to mutation, selection, genetic drift and recombination. Selection is modeled as single-locus terms (additive fitness) and two-loci terms (pairwise epistatic fitness). The problem is posed to infer epistatic fitness from population-wide whole-genome data from a time series of a developing population. We generate such data in silico, and show that in the Quasi-Linkage Equilibrium (QLE) phase of Kimura, Neher and Shraiman, that pertains at high enough recombination rates and low enough mutation rates, epistatic fitness can be quantitatively correctly inferred using inverse Ising/Potts methods.

q-bio.PE

Maximum likelihood reconstruction for Ising models with asynchronous updates

We describe how the couplings in an asynchronous kinetic Ising model can be inferred. We consider two cases, one in which we know both the spin history and the update times and one in which we only know the spin history. For the first case, we show that one can average over all possible choices of update times to obtain a learning rule that depends only on spin correlations and can also be derived from the equations of motion for the correlations. For the second case, the same rule can be derived within a further decoupling approximation. We study all methods numerically for fully asymmetric Sherrington-Kirkpatrick models, varying the data length, system size, temperature, and external field. Good convergence is observed in accordance with the theoretical expectations.

physics.data-an

L$_1$ Regularization for Reconstruction of a non-equilibrium Ising Model

The couplings in a sparse asymmetric, asynchronous Ising network are reconstructed using an exact learning algorithm. L$_1$ regularization is used to remove the spurious weak connections that would otherwise be found by simply minimizing the minus likelihood of a finite data set. In order to see how L$_1$ regularization works in detail, we perform the calculation in several ways including (1) by iterative minimization of a cost function equal to minus the log likelihood of the data plus an L$_1$ penalty term, and (2) an approximate scheme based on a quadratic expansion of the cost function around its minimum. In these schemes, we track how connections are pruned as the strength of the L$_1$ penalty is increased from zero to large values. The performance of the methods for various coupling strengths is quantified using ROC curves.

stat.ME

Network inference using asynchronously updated kinetic Ising Model

Network structures are reconstructed from dynamical data by respectively naive mean field (nMF) and Thouless-Anderson-Palmer (TAP) approximations. For TAP approximation, we use two methods to reconstruct the network: a) iteration method; b) casting the inference formula to a set of cubic equations and solving it directly. We investigate inference of the asymmetric Sherrington- Kirkpatrick (S-K) model using asynchronous update. The solutions of the sets cubic equation depend of temperature T in the S-K model, and a critical temperature Tc is found around 2.1. For T < Tc, the solutions of the cubic equation sets are composed of 1 real root and two conjugate complex roots while for T > Tc there are three real roots. The iteration method is convergent only if the cubic equations have three real solutions. The two methods give same results when the iteration method is convergent. Compared to nMF, TAP is somewhat better at low temperatures, but approaches the same performance as temperature increase. Both methods behave better for longer data length, but for improvement arises, TAP is well pronounced.

stat.CO