SearcharxivSearch

arXiv subjects

Xavier Didelot

Publications and source records attributed to Xavier Didelot.

7 recordsLinked to original sources

Bayesian copula-based modelling for multi-type spatio-temporal epidemic data

The study of infectious disease epidemiology for multi-type disease pathogens requires modelling techniques that account for the complex interactions existing between strains across geography and time. In this paper, we propose a novel multi-type spatio-temporal infectious disease model to better support the understanding of these pathogens. We formulate a joint state-space for all epidemics arising for a given multi-type pathogen as well as biologically informed representations of how these epidemic states may interact. We introduce the use of several copula models to uncover the dependence structure of epidemics between strains. We develop a computationally efficient Markov chain Monte Carlo (MCMC) sampling scheme for all proposed models. We also provide robust model comparison techniques using bridge sampling and importance sampling to evaluate model evidence in high-dimensional space. We demonstrate the performance of our proposed models using simulated datasets, where simulated epidemics were successfully identified and associated parameters correctly inferred. The proposed models were also fitted to monthly multi-type incidence data on invasive meningococcal disease from 26 European countries. The accompanying software is freely available as a R package at https://github.com/Matthewadeoye/MultiOutbreaks.

stat.ME

Bayesian spatio-temporal modelling for infectious disease outbreak detection

The Bayesian analysis of infectious disease surveillance data from multiple locations typically involves building and fitting a spatio-temporal model of how the disease spreads in the structured population. Here we present new generally applicable methodology to perform this task. We introduce a parsimonious representation of seasonality and a biologically informed specification of the outbreak component to avoid parameter identifiability issues. We develop a computationally efficient Bayesian inference methodology for the proposed models, including techniques to detect outbreaks by computing marginal posterior probabilities at each spatial location and time point. We show that it is possible to efficiently integrate out the discrete parameters associated with outbreak states, enabling the use of dynamic Hamiltonian Monte Carlo (HMC) as a complementary alternative to a hybrid Markov chain Monte Carlo (MCMC) algorithm. Furthermore, we introduce a robust Bayesian model comparison framework based on importance sampling to approximate model evidence in high-dimensional space. The performance of our methodology is validated through systematic simulation studies, where simulated outbreaks were successfully detected, and our model comparison strategy demonstrates strong reliability. We also apply our new methodology to monthly incidence data on invasive meningococcal disease from 28 European countries. The results highlight outbreaks across multiple countries and months, with model comparison analysis showing that the new specification outperforms previous approaches. The accompanying software is freely available as a R package at https://github.com/Matthewadeoye/DetectOutbreaks.

stat.ME

Ancestral process for infectious disease outbreaks with superspreading

When an infectious disease outbreak is of a relatively small size, describing the ancestry of a sample of infected individuals is difficult because most ancestral models assume large population sizes. Given a set of infected individuals, we show that it is possible to express exactly the probability that they have the same infector, either inclusively (so that other individuals may have the same infector too) or exclusively (so that they may not). To compute these probabilities requires knowledge of the offspring distribution, which determines how many infections each infected individual causes. We consider transmission both without and with superspreading, in the form of a Poisson and a Negative-Binomial offspring distribution, respectively. We show how our results can be incorporated into a new lambda-coalescent model which allows multiple lineages to coalesce together. We call this new model the omega-coalescent, we compare it with previously proposed alternatives, and advocate its use in future studies of infectious disease outbreaks.

q-bio.PE

A Bayesian Modelling Framework with Model Comparison for Epidemics with Super-Spreading

The transmission dynamics of an epidemic are rarely homogeneous. Super-spreading events and super-spreading individuals are two types of heterogeneous transmissibility. Inference of super-spreading is commonly carried out on secondary case data, the expected distribution of which is known as the offspring distribution. However, this data is seldom available. Here we introduce a multi-model framework fit to incidence time-series, data that is much more readily available. The framework consists of five discrete-time, stochastic, branching-process models of epidemics spread through a susceptible population. The framework includes a baseline model of homogeneous transmission, a unimodal and a bimodal model for super-spreading events, as well as a unimodal and a bimodal model for super-spreading individuals. Bayesian statistics is used to infer model parameters using Markov Chain Monte-Carlo. Model comparison is conducted by computing Bayes factors, with importance sampling used to estimate the marginal likelihood of each model. This estimator is selected for its consistency and lower variance compared to alternatives. Application to simulated data from each model identifies the correct model for the majority of simulations and accurately infers the true parameters, such as the basic reproduction number. We also apply our methods to incidence data from the 2003 SARS outbreak and the Covid-19 pandemic. Model selection consistently identifies the same model and mechanism for a given disease, even when using different time series. Our estimates are consistent with previous studies based on secondary case data. Quantifying the contribution of super-spreading to disease transmission has important implications for infectious disease management and control. Our modelling framework is disease-agnostic and implemented as an R package, with potential to be a valuable tool for public health.

q-bio.QM

Bayesian Inference of Reproduction Number from Epidemiological and Genetic Data Using Particle MCMC

Inference of the reproduction number through time is of vital importance during an epidemic outbreak. Typically, epidemiologists tackle this using observed prevalence or incidence data. However, prevalence and incidence data alone is often noisy or partial. Models can also have identifiability issues with determining whether a large amount of a small epidemic or a small amount of a large epidemic has been observed. Sequencing data however is becoming more abundant, so approaches which can incorporate genetic data are an active area of research. We propose using particle MCMC methods to infer the time-varying reproduction number from a combination of prevalence data reported at a set of discrete times and a dated phylogeny reconstructed from sequences. We validate our approach on simulated epidemics with a variety of scenarios. We then apply the method to real data sets of HIV-1 in North Carolina, USA and tuberculosis in Buenos Aires, Argentina. The models and algorithms are implemented in an open source R package called EpiSky which is available at https://github.com/alicia-gill/EpiSky.

stat.ME

Epidemic clones, oceanic gene pools and eco-LD in the free living marine pathogen Vibrio parahaemolyticus

We investigated global patterns of variation in 157 whole genome sequences of Vibrio parahaemolyticus, a free-living and seafood associated marine bacterium. Pandemic clones, responsible for recent outbreaks of gastroenteritis in humans have spread globally. However, there are oceanic gene pools, one located in the oceans surrounding Asia and another in the Mexican Gulf. Frequent recombination means that most isolates have acquired the genetic profile of their current location. We investigated the genetic structure in the Asian gene pool by calculating the effective population size in two different ways. Under standard neutral models, the two estimates should give similar answers but we found a thirty fold difference. We propose that this discrepancy is caused by the subdivision of the species into a hundred or more ecotypes which are maintained stably in the population. To investigate the genetic factors involved, we used 51 unrelated isolates to conduct a genome-wide scan for epistatically interacting loci. We found a single example of strong epistasis between distant genome regions. A majority of strains had a type VI secretion system associated with bacterial killing. The remaining strains had genes associated with biofilm formation and regulated by c-di-GMP signaling. All strains had one or other of the two systems and none of isolate had complete complements of both systems, although several strains had remnants. Further top-down analysis of patterns of linkage disequilibrium within frequently recombining species will allow a detailed understanding of how selection acts to structure the pattern of variation within natural bacterial populations.

q-bio.PE

A Bayesian approach to inferring the phylogenetic structure of communities from metagenomic data

Metagenomics provides a powerful new tool set for investigating evolutionary interactions with the environment. However, an absence of model-based statistical methods means that researchers are often not able to make full use of this complex information. We present a Bayesian method for inferring the phylogenetic relationship among related organisms found within metagenomic samples. Our approach exploits variation in the frequency of taxa among samples to simultaneously infer each lineage haplotype, the phylogenetic tree connecting them, and their frequency within each sample. Applications of the algorithm to simulated data show that our method can recover a substantial fraction of the phylogenetic structure even in the presence of strong mixing among samples. We provide examples of the method applied to data from green sulfur bacteria recovered from an Antarctic lake, plastids from mixed Plasmodium falciparum infections, and virulent Neisseria meningitidis samples.

q-bio.QM