SearcharxivSearch

arXiv subjects

Peter Neal

Publications and source records attributed to Peter Neal.

15 recordsLinked to original sources

Household epidemic models revisited

We analyse a generalized stochastic household epidemic model defined by a bivariate random variable $(X_G, X_L)$, representing the number of global and local infectious contacts that an infectious individual makes during their infectious period. Each global contact is selected uniformly among all individuals and each local contact is selected uniformly among all other household members. The main focus is when all households have the same size $h \geq 2$, and the number of households is large. Large population properties of the model are derived including a central limit theorem for the final size of a major epidemic, the proof of which utilises an enhanced embedding argument. A modification of the epidemic model is considered where local contacts are replaced by global contacts independently with probability $p$. We then prove monotonicity results for the probability of the major outbreak and the limiting final fraction infected $z$ (conditioned on a major outbreak). a) The probability of a major outbreak is shown to be increasing in both $h$ and $p$ for any distribution of $X_L$. b) The final size $z$ increases monotonically with both $h$ and $p$ if the probability generating function (pgf) of $X_L$ is log-convex, which is satisfied by traditional household epidemic models where $X_L$ has a mixed-Poisson distribution. Additionally, we provide counter examples to b) when the pgf of $X_L$ is not log-convex.

math.PR

An epidemic model on a network having two group structures with tunable overlap

A network epidemic model is studied. The underlying social network has two different types of group structures, households and workplaces, such that each individual belongs to exactly one household and one workplace. The random network is constructed such that a parameter $\theta$ controls the degree of overlap between the two group structures: $\theta=0$ corresponding to all household members belonging to the same workplace and $\theta=1$ to all household members belonging to distinct workplaces. On the network a stochastic SIR epidemic is defined, having an arbitrary but specified infectious period distribution, with global (community), household and workplace infectious contacts. The stochastic epidemic model is analysed as the population size $n\to\infty$ with the (asymptotic) probability, and size, of a major outbreak obtained. These results are proved in greater generality than existing results in the literature by allowing for any fixed $0 \leq \theta \leq 1$, a non-constant infectious period distribution, the presence or absence of global infection and potentially (asymptotically) infinite local outbreaks.

math.PR

Efficient Bayesian model selection for coupled hidden Markov models with application to infectious diseases

Performing model selection for coupled hidden Markov models (CHMMs) is highly challenging, owing to the large dimension of the hidden state process. Whilst in principle the hidden state process can be marginalized out via forward filtering, in practice the computational cost of doing so increases exponentially with the number of coupled Markov chains, making this approach infeasible in most applications. Monte Carlo methods can be utilized, but despite many remarkable developments in model selection methodology, generic approaches continue to be ill-suited for such high-dimensional problems. Here we develop specialized solutions for CHMMs with weak inter-chain dependencies. Specifically we construct effective proposal distributions for the hidden state process that remain computationally viable as the number of chains increases, and that require little user input or tuning. This methodology is particularly applicable to individual-level infectious disease models characterized as CHMMs, in which each chain represents an individual, and the coupling represents contact between individuals. Since the only significant contacts are between susceptible and infectious individuals, and since multiple infection pathways are often possible, the resulting CHMMs naturally have low inter-chain dependencies. We demonstrate the utility of our methodology with an application to a study of highly pathogenic avian influenza in chickens.

stat.ME

Real time analysis of epidemic data

Infectious diseases have severe health and economic consequences for society. It is important in controlling the spread of an emerging infectious disease to be able to both estimate the parameters of the underlying model and identify those individuals most at risk of infection in a timely manner. This requires having a mechanism to update inference on the model parameters and the progression of the disease as new data becomes available. However, Markov chain Monte Carlo (MCMC), the gold standard for statistical inference for infectious disease models, is not equipped to deal with this important problem. Motivated by the need to develop effective statistical tools for emerging diseases and using the 2001 UK Foot-and-Mouth disease outbreak as an exemplar, we introduce a Sequential Monte Carlo (SMC) algorithm to enable real-time analysis of epidemic outbreaks. Naive application of SMC methods leads to significant particle degeneracy which are successfully overcome by particle perturbation and incorporating MCMC-within-SMC updates.

stat.CO

The basic reproduction number, $R_0$, in structured populations

In this paper, we provide a straightforward approach to defining and deriving the key epidemiological quantity, the basic reproduction number, $R_0$, for Markovian epidemics in structured populations. The methodology derived is applicable to, and demonstrated on, both $SIR$ and $SIS$ epidemics and allows for population as well as epidemic dynamics. The approach taken is to consider the epidemic process as a multitype process by identifying and classifying the different types of infectious units along with the infections from, and the transitions between, infectious units. For the household model, we show that our expression for $R_0$ agrees with earlier work despite the alternative nature of the construction of the mean reproductive matrix, and hence, the basic reproduction number.

q-bio.PE

Model comparison with missing data using MCMC and importance sampling

Selecting between competing statistical models is a challenging problem especially when the competing models are non-nested. In this paper we offer a simple solution by devising an algorithm which combines MCMC and importance sampling to obtain computationally efficient estimates of the marginal likelihood which can then be used to compare the models. The algorithm is successfully applied to longitudinal epidemic and time series data sets and shown to outperform existing methods for computing the marginal likelihood.

stat.CO

Optimal scaling of the independence sampler: Theory and Practice

The independence sampler is one of the most commonly used MCMC algorithms usually as a component of a Metropolis-within-Gibbs algorithm. The common focus for the independence sampler is on the choice of proposal distribution to obtain an as high as possible acceptance rate. In this paper we have a somewhat different focus concentrating on the use of the independence sampler for updating augmented data in a Bayesian framework where a natural proposal distribution for the independence sampler exists. Thus we concentrate on the proportion of the augmented data to update to optimise the independence sampler. Generic guidelines for optimising the independence sampler are obtained for independent and identically distributed product densities mirroring findings for the random walk Metropolis algorithm. The generic guidelines are shown to be informative beyond the narrow confines of idealised product densities in two epidemic examples.

stat.CO

On expected durations of birth-death processes, with applications to branching processes and SIS epidemics

We study continuous-time birth-death type processes, where individuals have independent and identically distributed lifetimes, according to a random variable Q, with E[Q]=1, and where the birth rate if the population is currently in state (has size) n is \alpha(n). We focus on two important examples, namely \alpha(n)=\lambda n being a branching process, and \alpha(n)=\lambda n(N-n)/N which corresponds to an SIS epidemic model in a homogeneously mixing community of fixed size N. The processes are assumed to start with a single individual, i.e. in state 1. Let T, A_n, C and S denote the (random) time to extinction, the total time spent in state $n$, the total number of individuals ever alive and the sum of the lifetimes of all individuals in the birth-death process, respectively. The main results of the paper give expressions for the expectation of all these quantities, and shows that these expectations are insensitive to the distribution of Q. We also derive an asymptotic expression for the expected time to extinction of the SIS epidemic, but now starting at the endemic state, which is not independent of the distribution of Q. The results are also applied to the household SIS epidemic, showing that its threshold parameter R_* is insensitive to the distribution of Q, contrary to the household SIR epidemic, for which R_* does depend on Q.

math.PR

Simulation based sequential Monte Carlo methods for discretely observed Markov processes

Parameter estimation for discretely observed Markov processes is a challenging problem. However, simulation of Markov processes is straightforward using the Gillespie algorithm. We exploit this ease of simulation to develop an effective sequential Monte Carlo (SMC) algorithm for obtaining samples from the posterior distribution of the parameters. In particular, we introduce two key innovations, coupled simulations, which allow us to study multiple parameter values on the basis of a single simulation, and a simple, yet effective, importance sampling scheme for steering simulations towards the observed data. These innovations substantially improve the efficiency of the SMC algorithm with minimal effect on the speed of the simulation process. The SMC algorithm is successfully applied to two examples, a Lotka-Volterra model and a Repressilator model.

stat.CO

On the expected time a branching process has K individuals alive

Consider a homogeneous time-continuous branching process where individuals have constant birth rate $\delta$, and life length distribution $Q$ having mean $E(Q)=1$. Let $X(u)$ denote the number of individuals alive at time $u$, and assume that $X(0)=1$. Let $K$ be a positive integer and define $A_K:=\int_0^\infty 1_{\{X(u)=K\}}du$, the accumulated time that the branching process has exactly $K$ individuals alive. In this paper we prove that $E(A_K)=\delta^{K-1}/\left(k(1\vee\delta)^K\right)$, irrespective of the life length distribution $Q$, subject to the normalizing condition $E(Q)=1$.

math.PR

Robust estimation of microbial diversity in theory and in practice

Quantifying diversity is of central importance for the study of structure, function and evolution of microbial communities. The estimation of microbial diversity has received renewed attention with the advent of large-scale metagenomic studies. Here, we consider what the diversity observed in a sample tells us about the diversity of the community being sampled. First, we argue that one cannot reliably estimate the absolute and relative number of microbial species present in a community without making unsupported assumptions about species abundance distributions. The reason for this is that sample data do not contain information about the number of rare species in the tail of species abundance distributions. We illustrate the difficulty in comparing species richness estimates by applying Chao's estimator of species richness to a set of in silico communities: they are ranked incorrectly in the presence of large numbers of rare species. Next, we extend our analysis to a general family of diversity metrics ("Hill diversities"), and construct lower and upper estimates of diversity values consistent with the sample data. The theory generalizes Chao's estimator, which we retrieve as the lower estimate of species richness. We show that Shannon and Simpson diversity can be robustly estimated for the in silico communities. We analyze nine metagenomic data sets from a wide range of environments, and show that our findings are relevant for empirically-sampled communities. Hence, we recommend the use of Shannon and Simpson diversity rather than species richness in efforts to quantify and compare microbial diversity.

q-bio.PE

Optimal scaling of random walk Metropolis algorithms with discontinuous target densities

We consider the optimal scaling problem for high-dimensional random walk Metropolis (RWM) algorithms where the target distribution has a discontinuous probability density function. Almost all previous analysis has focused upon continuous target densities. The main result is a weak convergence result as the dimensionality d of the target densities converges to infinity. In particular, when the proposal variance is scaled by $d^{-2}$, the sequence of stochastic processes formed by the first component of each Markov chain converges to an appropriate Langevin diffusion process. Therefore optimizing the efficiency of the RWM algorithm is equivalent to maximizing the speed of the limiting diffusion. This leads to an asymptotic optimal acceptance rate of $e^{-2}$ (=0.1353) under quite general conditions. The results have major practical implications for the implementation of RWM algorithms by highlighting the detrimental effect of choosing RWM algorithms over Metropolis-within-Gibbs algorithms.

math.PR

The time to extinction for an SIS-household-epidemic model

We analyse a stochastic SIS epidemic amongst a finite population partitioned into households. Since the population is finite, the epidemic will eventually go extinct, i.e., have no more infectives in the population. We study the effects of population size and within household transmission upon the time to extinction. This is done through two approximations. The first approximation is suitable for all levels of within household transmission and is based upon an Ornstein-Uhlenbeck process approximation for the diseases fluctuations about an endemic level relying on a large population. The second approximation is suitable for high levels of within household transmission and approximates the number of infectious households by a simple homogeneously mixing SIS model with the households replaced by individuals. The analysis, supported by a simulation study, shows that the mean time to extinction is minimized by moderate levels of within household transmission.

q-bio.PE

Multitype randomized Reed--Frost epidemics and epidemics upon random graphs

We consider a multitype epidemic model which is a natural extension of the randomized Reed--Frost epidemic model. The main result is the derivation of an asymptotic Gaussian limit theorem for the final size of the epidemic. The method of proof is simpler, and more direct, than is used for similar results elsewhere in the epidemics literature. In particular, the results are specialized to epidemics upon extensions of the Bernoulli random graph.

math.PR

Optimal scaling for partially updating MCMC algorithms

In this paper we shall consider optimal scaling problems for high-dimensional Metropolis--Hastings algorithms where updates can be chosen to be lower dimensional than the target density itself. We find that the optimal scaling rule for the Metropolis algorithm, which tunes the overall algorithm acceptance rate to be 0.234, holds for the so-called Metropolis-within-Gibbs algorithm as well. Furthermore, the optimal efficiency obtainable is independent of the dimensionality of the update rule. This has important implications for the MCMC practitioner since high-dimensional updates are generally computationally more demanding, so that lower-dimensional updates are therefore to be preferred. Similar results with rather different conclusions are given for so-called Langevin updates. In this case, it is found that high-dimensional updates are frequently most efficient, even taking into account computing costs.

math.PR