SearcharxivSearch

arXiv subjects

Daniela De Angelis

Publications and source records attributed to Daniela De Angelis.

At least 19 recordsLinked to original sources

GENIE: Generative Neural Inference for Epidemics

The SARS-CoV-2 pandemic highlighted the ongoing risk infectious diseases pose to society and the value of reliable information on the likely future burden. When forecasting an epidemic at fine spatial resolution, traditionally used mechanistic compartmental model struggle to capture highly complex granular transmission dynamics, resulting in inaccurate and overconfident forecasts. However, detailed Agent-Based Models (ABMs), are challenging to calibrate and are too computationally expensive to use in real-time. Amortized simulation-based inference promises to overcome this difficulty by exploiting the power of machine learning (ML) to perform approximate forecasting at near-real-time using arbitrarily complex models of epidemics. In this work we introduce Generative Neural Inference for Epidemics (GENIE), a spatio-temporal ML-based framework for high-resolution forecasting of the burden of respiratory pathogens. GENIE is designed to reflect two key characteristics of outbreaks: (i) shared biological mechanisms across locations and (ii) location-specific characteristics affecting transmission dynamics. This results in the model architecture having two modules: (i) a Local Infection Encoder - which learns to represent disease dynamics shared across all locations and (ii) a Local Profile Encoder - which learns location-specific representations. Using simulations from a high-resolution spatio-temporal ABM, GENIE is trained to generate samples from an approximate posterior predictive distribution of future epidemic trajectories. Benchmarked against established statistical and ML models, GENIE demonstrates superior performance across a range of measures including the timing and magnitude of peak hospitalisations.

stat.ME

A Bayesian factor analysis model for non-randomised staggered designs

The employment of peer supporter workers starting in 2018 was one of the interventions deployed by National Health Service England as part of its Hepatitis C virus (HCV) elimination plan. Peers are individuals with relevant lived experience who educate their communities about the virus and promote testing and treatment. In this paper, we assess the causal effect of the peers intervention on HCV patient case-finding, using data on 22 administrative regions from January 2016 to May 2021. To do this, we develop a Bayesian causal factor analysis model for count outcomes and ordinal interventions. Our method provides uncertainty quantification for all causal estimands of interest, gains efficiency by jointly modelling the intervention assignment process, pre- and post-intervention outcomes, and provides estimates of both conditional average and individual treatment effects (ITEs). For ITEs, we propose a copula-based approach that allows practitioners to perform sensitivity analysis to assumptions made regarding the joint distribution of potential outcomes, that are necessary to estimate these quantities. Our analysis suggests that the introduction of peers led to an increase in HCV patient case-finding. Further, we found that the effect of the intervention increased with intervention intensity, and was stronger during the national COVID-19 lockdown.

stat.AP

A computationally efficient framework for realistic epidemic modelling through Gaussian Markov random fields

We tackle limitations of ordinary differential equation-driven Susceptible-Infections-Removed (SIR) models and their extensions that have recently be employed for epidemic nowcasting and forecasting. In particular, we deal with challenges related to the extension of SIR-type models to account for the so-called \textit{environmental stochasticity}, i.e., external factors, such as seasonal forcing, social cycles and vaccinations that can dramatically affect outbreaks of infectious diseases. Typically, in SIR-type models environmental stochasticity is modelled through stochastic processes. However, this stochastic extension of epidemic models leads to models with large dimension that increases over time. Here we propose a Bayesian approach to build an efficient modelling and inferential framework for epidemic nowcasting and forecasting by using Gaussian Markov random fields to model the evolution of these stochastic processes over time and across population strata. Importantly, we also develop a bespoke and computationally efficient Markov chain Monte Carlo algorithm to estimate the large number of parameters and latent states of the proposed model. We test our approach on simulated data and we apply it to real data from the Covid-19 pandemic in the United Kingdom.

stat.CO

Estimating the duration of RT-PCR positivity for SARS-CoV-2 from doubly interval censored data with undetected infections

Monitoring the incidence of new infections during a pandemic is critical for an effective public health response. General population prevalence surveys for SARS-CoV-2 can provide high-quality data to estimate incidence. However, estimation relies on understanding the distribution of the duration that infections remain detectable. This study addresses this need using data from the Coronavirus Infection Survey (CIS), a long-term, longitudinal, general population survey conducted in the UK. Analyzing these data presents unique challenges, such as doubly interval censoring, undetected infections, and false negatives. We propose a Bayesian nonparametric survival analysis approach, estimating a discrete-time distribution of durations and integrating prior information derived from a complementary study. Our methodology is validated through a simulation study, including its resilience to model misspecification, and then applied to the CIS dataset. This results in the first estimate of the full duration distribution in a general population, as well as methodology that could be transferred to new contexts.

stat.ME

The Lifebelt Particle Filter for robust estimation from low-valued count data

Particle filtering methods can be applied to estimation problems in discrete spaces on bounded domains, to sample from and marginalise over unknown hidden states. As in continuous settings, problems such as particle degradation can arise: proposed particles can be incompatible with the data, lying in low probability regions or outside the boundary constraints, and the discrete system could result in all particles having weights of zero. In this paper we introduce the Lifebelt Particle Filter (LBPF), a novel method for robust likelihood estimation in low-valued count problems. The LBPF combines a standard particle filter with one (or more) lifebelt particles which, by construction, lie within the boundaries of the discrete random variables, and therefore are compatible with the data. A mixture of resampled and non-resampled particles allows for the preservation of the lifebelt particle, which, together with the remaining particle swarm, provides samples from the filtering distribution, and can be used to generate unbiased estimates of the likelihood. The main benefit of the LBPF is that only one or few, wisely chosen, particles are sufficient to prevent particle collapse. Differently from other methods, there is no need to increase the number of particles, and therefore the computational effort, in regions of the parameter space that generate less likely hidden states. The LBPF can be used within a pseudo-marginal scheme to draw inferences on static parameters, $ \boldsymbolθ $, governing the system. We address here the estimation of a parameter governing probabilities of deaths and recoveries of hospitalised patients during an epidemic.

stat.CO

Inferring Epidemics from Multiple Dependent Data via Pseudo-Marginal Methods

Health-policy planning requires evidence on the burden that epidemics place on healthcare systems. Multiple, often dependent, datasets provide a noisy and fragmented signal from the unobserved epidemic process including transmission and severity dynamics. This paper explores important challenges to the use of state-space models for epidemic inference when multiple dependent datasets are analysed. We propose a new semi-stochastic model that exploits deterministic approximations for large-scale transmission dynamics while retaining stochasticity in the occurrence and reporting of relatively rare severe events. This model is suitable for many real-time situations including large seasonal epidemics and pandemics. Within this context, we develop algorithms to provide exact parameter inference and test them via simulation. Finally, we apply our joint model and the proposed algorithm to several surveillance data on the 2017-18 influenza epidemic in England to reconstruct transmission dynamics and estimate the daily new influenza infections as well as severity indicators such as the case-hospitalisation risk and the hospital-intensive care risk.

stat.AP

Real-time modelling of the SARS-CoV-2 pandemic in England 2020-2023: a challenging data integration

A central pillar of the UK's response to the SARS-CoV-2 pandemic was the provision of up-to-the moment nowcasts and short term projections to monitor current trends in transmission and associated healthcare burden. Here we present a detailed deconstruction of one of the 'real-time' models that was key contributor to this response, focussing on the model adaptations required over three pandemic years characterised by the imposition of lockdowns, mass vaccination campaigns and the emergence of new pandemic strains. The Bayesian model integrates an array of surveillance and other data sources including a novel approach to incorporating prevalence estimates from an unprecedented large-scale household survey. We present a full range of estimates of the epidemic history and the changing severity of the infection, quantify the impact of the vaccination programme and deconstruct contributing factors to the reproduction number. We further investigate the sensitivity of model-derived insights to the availability and timeliness of prevalence data, identifying its importance to the production of robust estimates.

stat.AP

Sample-efficient neural likelihood-free Bayesian inference of implicit HMMs

Likelihood-free inference methods based on neural conditional density estimation were shown to drastically reduce the simulation burden in comparison to classical methods such as ABC. When applied in the context of any latent variable model, such as a Hidden Markov model (HMM), these methods are designed to only estimate the parameters, rather than the joint distribution of the parameters and the hidden states. Naive application of these methods to a HMM, ignoring the inference of this joint posterior distribution, will thus produce an inaccurate estimate of the posterior predictive distribution, in turn hampering the assessment of goodness-of-fit. To rectify this problem, we propose a novel, sample-efficient likelihood-free method for estimating the high-dimensional hidden states of an implicit HMM. Our approach relies on learning directly the intractable posterior distribution of the hidden states, using an autoregressive-flow, by exploiting the Markov property. Upon evaluating our approach on some implicit HMMs, we found that the quality of the estimates retrieved using our method is comparable to what can be achieved using a much more computationally expensive SMC algorithm.

stat.ML

The NOSTRA model: coherent estimation of infection sources in the case of possible nosocomial transmission

Nosocomial infections have important consequences for patients and hospital staff: they worsen patient outcomes and their management stresses already overburdened health systems. Accurate judgements of whether an infection is nosocomial helps staff make appropriate choices to protect other patients within the hospital. Nosocomiality cannot be properly assessed without considering whether the infected patient came into contact with high risk potential infectors within the hospital. We developed a Bayesian model that integrates epidemiological, contact and pathogen genetic data to determine how likely an infection is to be nosocomial and the probability of given infection candidates being the source of the infection.

stat.AP

An approximate diffusion process for environmental stochasticity in infectious disease transmission modelling

Modelling the transmission dynamics of an infectious disease is a complex task. Not only it is difficult to accurately model the inherent non-stationarity and heterogeneity of transmission, but it is nearly impossible to describe, mechanistically, changes in extrinsic environmental factors including public behaviour and seasonal fluctuations. An elegant approach to capturing environmental stochasticity is to model the force of infection as a stochastic process. However, inference in this context requires solving a computationally expensive ``missing data" problem, using data-augmentation techniques. We propose to model the time-varying transmission-potential as an approximate diffusion process using a path-wise series expansion of Brownian motion. This approximation replaces the ``missing data" imputation step with the inference of the expansion coefficients: a simpler and computationally cheaper task. We illustrate the merit of this approach through two examples: modelling influenza using a canonical SIR model, and the modelling of COVID-19 pandemic using a multi-type SEIR model.

stat.CO

Trends in COVID-19 hospital outcomes in England before and after vaccine introduction, a cohort study

Widespread vaccination campaigns have changed the landscape for COVID-19, vastly altering symptoms and reducing morbidity and mortality. We estimate trends in mortality by month of admission and vaccination status among those hospitalised with COVID-19 in England between March 2020 to September 2021, controlling for demographic factors and hospital load. Among 259,727 hospitalised COVID-19 cases, 51,948 (20.0%) experienced mortality in hospital. Hospitalised fatality risk ranged from 40.3% (95% confidence interval 39.4-41.3%) in March 2020 to 8.1% (7.2-9.0%) in June 2021. Older individuals and those with multiple co-morbidities were more likely to die or else experienced longer stays prior to discharge. Compared to unvaccinated people, the hazard of hospitalised mortality was 0.71 (0.67-0.77) with a first vaccine dose, and 0.56 (0.52-0.61) with a second vaccine dose. Compared to hospital load at 0-20% of the busiest week, the hazard of hospitalised mortality during periods of peak load (90-100%), was 1.23 (1.12-1.34). The prognosis for people hospitalised with COVID-19 in England has varied substantially throughout the pandemic and according to case-mix, vaccination, and hospital load. Our estimates provide an indication for demands on hospital resources, and the relationship between hospital burden and outcomes.

stat.AP

A comparison of two frameworks for multi-state modelling, applied to outcomes after hospital admissions with COVID-19

We compare two multi-state modelling frameworks that can be used to represent dates of events following hospital admission for people infected during an epidemic. The methods are applied to data from people admitted to hospital with COVID-19, to estimate the probability of admission to ICU, the probability of death in hospital for patients before and after ICU admission, the lengths of stay in hospital, and how all these vary with age and gender. One modelling framework is based on defining transition-specific hazard functions for competing risks. A less commonly used framework defines partially-latent subpopulations who will experience each subsequent event, and uses a mixture model to estimate the probability that an individual will experience each event, and the distribution of the time to the event given that it occurs. We compare the advantages and disadvantages of these two frameworks, in the context of the COVID-19 example. The issues include the interpretation of the model parameters, the computational efficiency of estimating the quantities of interest, implementation in software and assessing goodness of fit. In the example, we find that some groups appear to be at very low risk of some events, in particular ICU admission, and these are best represented by using "cure-rate" models to define transition-specific hazards. We provide general-purpose software to implement all the models we describe in the "flexsurv" R package, which allows arbitrarily-flexible distributions to be used to represent the cause-specific hazards or times to events.

stat.ME

Evaluating the impact of local tracing partnerships on the performance of contact tracing for COVID-19 in England

Assessing the impact of an intervention using time-series observational data on multiple units and outcomes is a frequent problem in many fields of scientific research. In this paper, we present a novel method to estimate intervention effects in such a setting by generalising existing approaches based on the factor analysis model and developing a Bayesian algorithm for inference. Our method is one of the few that can simultaneously: deal with outcomes of mixed type (continuous, binomial, count); increase efficiency in the estimates of the causal effects by jointly modelling multiple outcomes affected by the intervention; easily provide uncertainty quantification for all causal estimands of interest. We use the proposed approach to evaluate the impact that local tracing partnerships (LTP) had on the effectiveness of England's Test and Trace (TT) programme for COVID-19. Our analyses suggest that, overall, LTPs had a small positive impact on TT. However, there is considerable heterogeneity in the estimates of the causal effects over units and time.

stat.AP

Hospitalisation risk for COVID-19 patients infected with SARS-CoV-2 variant B.1.1.7: cohort analysis

Objective: To evaluate the relationship between coronavirus disease 2019 (COVID-19) diagnosis with SARS-CoV-2 variant B.1.1.7 (also known as Variant of Concern 202012/01) and the risk of hospitalisation compared to diagnosis with wildtype SARS-CoV-2 variants. Design: Retrospective cohort, analysed using stratified Cox regression. Setting: Community-based SARS-CoV-2 testing in England, individually linked with hospitalisation data. Participants: 839,278 laboratory-confirmed COVID-19 patients, of whom 36,233 had been hospitalised within 14 days, tested between 23rd November 2020 and 31st January 2021 and analysed at a laboratory with an available TaqPath assay that enables assessment of S-gene target failure (SGTF). SGTF is a proxy test for the B.1.1.7 variant. Patient data were stratified by age, sex, ethnicity, deprivation, region of residence, and date of positive test. Main outcome measures: Hospitalisation between 1 and 14 days after the first positive SARS-CoV-2 test. Results: 27,710 of 592,409 SGTF patients (4.7%) and 8,523 of 246,869 non-SGTF patients (3.5%) had been hospitalised within 1-14 days. The stratum-adjusted hazard ratio (HR) of hospitalisation was 1.52 (95% confidence interval [CI] 1.47 to 1.57) for COVID-19 patients infected with SGTF variants, compared to those infected with non-SGTF variants. The effect was modified by age (P<0.001), with HRs of 0.93-1.21 for SGTF compared to non-SGTF patients below age 20 years, 1.29 in those aged 20-29, and 1.45-1.65 in age groups 30 years or older. Conclusions: The results suggest that the risk of hospitalisation is higher for individuals infected with the B.1.1.7 variant compared to wildtype SARS-CoV-2, likely reflecting a more severe disease. The higher severity may be specific to adults above the age of 30.

stat.AP

Quantifying efficiency gains of innovative designs of two-arm vaccine trials for COVID-19 using an epidemic simulation model

Clinical trials of a vaccine during an epidemic face particular challenges, such as the pressure to identify an effective vaccine quickly to control the epidemic, and the effect that time-space-varying infection incidence has on the power of a trial. We illustrate how the operating characteristics of different trial design elements may be evaluated using a network epidemic and trial simulation model, based on COVID-19 and individually randomised two-arm trials with a binary outcome. We show that "ring" recruitment strategies, prioritising participants at high risk of infection, can result in substantial improvement in terms of power, if sufficiently many contacts of observed cases are at high risk. In addition, we introduce a novel method to make more efficient use of the data from the earliest cases of infection observed in the trial, whose infection may have been too early to be vaccine-preventable. Finally, we compare several methods of response-adaptive randomisation, discussing their advantages and disadvantages in this two-arm context and identifying particular adaptation strategies that preserve power and estimation properties, while slightly reducing the number of infections, given an effective vaccine.

q-bio.PE

Trends in risks of severe events and lengths of stay for COVID-19 hospitalisations in England over the pre-vaccination era: results from the Public Health England SARI-Watch surveillance scheme

Background: Trends in hospitalised case-fatality risk (HFR), risk of intensive care unit (ICU) admission and lengths of stay for patients hospitalised for COVID-19 in England over the pre-vaccination era are unknown. Methods: Data on hospital and ICU admissions with COVID-19 at 31 NHS trusts in England were collected by Public Health England's Severe Acute Respiratory Infections surveillance system and linked to death information. We applied parametric multi-state mixture models, accounting for censored outcomes and regressing risks and times between events on month of admission, geography, and baseline characteristics. Findings: 20,785 adults were admitted with COVID-19 in 2020. Between March and June/July/August estimated HFR reduced from 31.9% (95% confidence interval 30.3-33.5%) to 10.9% (9.4-12.7%), then rose steadily from 21.6% (18.4-25.5%) in September to 25.7% (23.0-29.2%) in December, with steeper increases among older patients, those with multi-morbidity and outside London/South of England. ICU admission risk reduced from 13.9% (12.8-15.2%) in March to 6.2% (5.3-7.1%) in May, rising to a high of 14.2% (11.1-17.2%) in September. Median length of stay in non-critical care increased during 2020, from 6.6 to 12.3 days for those dying, and from 6.1 to 9.3 days for those discharged. Interpretation: Initial improvements in patient outcomes, corresponding to developments in clinical practice, were not sustained throughout 2020, with HFR in December approaching the levels seen at the start of the pandemic, whilst median hospital stays have lengthened. The role of increased transmission, new variants, case-mix and hospital pressures in increasing COVID-19 severity requires urgent further investigation.

stat.AP

HIV transmission in men who have sex with men in England: on track for elimination by 2030?

Background: After a decade of a treatment as prevention (TasP) strategy based on progressive HIV testing scale-up and earlier treatment, a reduction in the estimated number of new infections in men-who-have-sex-with-men (MSM) in England had yet to be identified by 2010. To achieve internationally agreed targets for HIV control and elimination, test-and-treat prevention efforts have been dramatically intensified over the period 2010-2015, and, from 2016, further strengthened by pre-exposure prophylaxis (PrEP). Methods: Application of a novel age-stratified back-calculation approach to data on new HIV diagnoses and CD4 count-at-diagnosis, enabled age-specific estimation of HIV incidence, undiagnosed infections and mean time-to-diagnosis across both the 2010-2015 and 2016-2018 periods. Estimated incidence trends were then extrapolated, to quantify the likelihood of achieving HIV elimination by 2030. Findings: A fall in HIV incidence in MSM is estimated to have started in 2012/3, eighteen months before the observed fall in new diagnoses. A steep decrease from 2,770 annual infections (95% credible interval 2.490-3,040) in 2013 to 1,740 (1,500-2,010) in 2015 is estimated, followed by steady decline from 2016, reaching 854 (441-1,540) infections in 2018. A decline is consistently estimated in all age groups, with a fall particularly marked in the 24-35 age group, and slowest in the 45+ group. Comparable declines are estimated in the number of undiagnosed infections. Interpretation: The peak and subsequent sharp decline in HIV incidence occurred prior to the phase-in of PrEP. Definining elimination as a public health threat to be < 50 new infections (1.1 infections per 10,000 at risk), 40% of incidence projections hit this threshold by 2030. In practice, targeted policies will be required, particularly among the 45+y where STIs are increasing most rapidly.

q-bio.QM

Assessing the causal effect of binary interventions from observational panel data with few treated units

Researchers are often challenged with assessing the impact of an intervention on an outcome of interest in situations where the intervention is non-randomised, the intervention is only applied to one or few units, the intervention is binary, and outcome measurements are available at multiple time points. In this paper, we review existing methods for causal inference in these situations. We detail the assumptions underlying each method, emphasize connections between the different approaches and provide guidelines regarding their practical implementation. Several open problems are identified thus highlighting the need for future research.

stat.AP