SearcharxivSearch

arXiv subjects

Ben Lambert

Publications and source records attributed to Ben Lambert.

At least 19 recordsLinked to original sources

Multi-scale measures of time-varying epidemic spread on human mobility networks

Human movement drives the spatial spread and persistence of many infectious diseases, yet existing theory and real-time operational tools for inferring the instantaneous reproduction number R(t) often assume static and/or homogeneously mixing populations and cannot describe how individuals generate and acquire infections heterogeneously based on their movement patterns within a day. Renewal equations underpin many such popular estimators of R(t), and here, we develop a network-based modelling framework from which we derive new mechanism-led renewal equations and control indicators for outbreaks of infectious diseases. These equations directly integrate within-day human movement to rigorously define a family of instantaneous reproduction numbers; inward, outward, and type R(t) for individual locations, R(t) between locations, R(t) at meeting locations, and R(t) for the entire mobility network. These quantities correct for the unsuitability of existing location-specific R(t) estimators that operate in closed, static populations. Applying our framework to epidemics on diverse types of networks alongside mobile phone data, we demonstrate how our new framework's outputs provide new, multi-scale control indicators at the network, location, and transmission corridor scales, and can be used to design targeted disease control interventions including the strength, type, and length of intervention required across space and time. We capture the biasing effects of different existing ways to measure location-specific and network-level transmission potential without capturing within-day human movements. This generalisable framework redefines reproduction numbers in real-world outbreaks that are shaped by individuals moving across connected locations, enabling more spatially and temporally precise interventions.

q-bio.QM

A note on conditional densities, Bayes' rule, and recent criticisms of Bayesian inference

When performing Bayesian inference, we frequently need to work with conditional probability densities. For example, the posterior function is the conditional density of the parameters given the data. Some might worry that conditional densities are ill-defined, considering that for a continuous random variable $Y$, the event $\{Y=y\}$ has probability zero, meaning the formula $\mathbb{P}(A|B)=\mathbb{P}(A\cap B)/\mathbb{P}(B)$ is inapplicable. In reality, when we work with conditional densities, we never condition directly on the zero-probability event $\{Y=y\}$; rather, we first condition on the random variable $Y$, and then we may plug in an observed value $y$. The first purpose of our article is to provide an exposition on conditional densities that elaborates on this point. While we have aimed to make this explanation accessible, we follow it with a roadmap of the measure theory needed to make it rigorous. A recent preprint (arXiv:2411.13570) has expressed the concern that probability densities are ill-defined and that as a result Bayes' theorem cannot be used, and they provide examples that allegedly demonstrate inconsistencies in the Bayesian framework. The second purpose of our article is to investigate their claims. We contend that the examples given in their work do not demonstrate any inconsistencies; we find that there are mathematical errors and that they deviate significantly from the Bayesian framework.

stat.ME

Foliation of null cones by surfaces of constant spacetime mean curvature near MOTS

Marginally Outer Trapped Surfaces (MOTS) in spacetimes are well-known to indicate the existence of black holes. Using flow techniques, we prove that a neighbourhood of a stable MOTS in a null cone may be foliated by hypersurfaces of constant spacetime mean curvature. We also provide methods to construct prescribed spacetime mean curvature surfaces within null cones.

math.DG

Misspecification of the generation time distribution and its impact on Rt estimates in structured populations

Due to its ability to summarise 'real-time' epidemic behaviour, the time-dependent reproduction number, Rt, is a useful metric for tracking pathogen transmission and quantifying the effects of interventions during infectious disease outbreaks. The predominant models underlying inferred Rt trajectories are renewal equations, their success owing in part to the relatively few assumptions they require. One necessary assumption is the generation time distribution, which summarises the time periods between infections in infector-infectee transmission pairs. This distribution is typically assumed to be the same across all members of a population. In reality, however, it may vary systematically between population groups. In this study, we consider two Rt inference frameworks based on renewal equation models: one for a single, homogeneous group and another accounting for a structured population. We compare the estimates of Rt generated by the two models and investigate, both analytically and through simulations, under which conditions the conclusions drawn from these modelling paradigms differ. We also demonstrate a methodology for selecting the generation time for the one-group model that correctly encapsulates variations between different population groups; this allows us to use a renewal framework for a one-group model to infer Rt when, in fact, the population is structured. Finally, we use real epidemic data to demonstrate that practical Rt estimates can differ depending on whether the underlying model is the one-group model or the multi-group model. Our results motivate the need for rigorous collection of detailed epidemic data and consideration of differences between population groups to improve the accuracy of Rt estimates that are used to guide public health policy responses.

q-bio.PE

Assessing the performance of compartmental and renewal models for learning $R_{t}$ using spatially heterogeneous epidemic simulations on real geographies

The time-varying reproduction number ($R_t$) gives an indication of the trajectory of an infectious disease outbreak. Commonly used frameworks for inferring $R_t$ from epidemiological time series include those based on compartmental models (such as the SEIR model) and renewal equation models. These inference methods are usually validated using synthetic data generated from a simple model, often from the same class of model as the inference framework. However, in a real outbreak the transmission processes, and thus the infection data collected, are much more complex. The performance of common $R_t$ inference methods on data with similar complexity to real world scenarios has been subject to less comprehensive validation. We therefore propose evaluating these inference methods on outbreak data generated from a sophisticated, geographically accurate agent-based model. We illustrate this proposed method by generating synthetic data for two outbreaks in Northern Ireland: one with minimal spatial heterogeneity, and one with additional heterogeneity. We find that the simple SEIR model struggles with the greater heterogeneity, while the renewal equation model demonstrates greater robustness to spatial heterogeneity, though is sensitive to the accuracy of the generation time distribution used in inference. Our approach represents a principled way to benchmark epidemiological inference tools and is built upon an open-source software platform for reproducible epidemic simulation and inference.

q-bio.PE

The time-dependent reproduction number for epidemics in heterogeneous populations

The time-dependent reproduction number Rt can be used to track pathogen transmission and to assess the efficacy of interventions. This quantity can be estimated by fitting renewal equation models to time series of infectious disease case counts. These models almost invariably assume a homogeneous population. Individuals are assumed not to differ systematically in the rates at which they come into contact with others. It is also assumed that the typical time that elapses between one case and those it causes (known as the generation time distribution) does not differ across groups. But contact patterns are known to widely differ by age and according to other demographic groupings, and infection risk and transmission rates have been shown to vary across groups for a range of directly transmitted diseases. Here, we derive from first principles a renewal equation framework which accounts for these differences in transmission across groups. We use a generalisation of the classic McKendrick-von Foerster equation to handle populations structured into interacting groups. This system of partial differential equations allows us to derive a simple analytical expression for Rt which involves only group-level contact patterns and infection risks. We show that the same expression emerges from both deterministic and stochastic discrete-time versions of the model and demonstrate via simulations that our Rt expression governs the long-run fate of epidemics. Our renewal equation model provides a basis from which to account for more realistic, diverse populations in epidemiological models and opens the door to inferential approaches which use known group characteristics to estimate Rt.

q-bio.PE

Renewal equations for vector-borne diseases

During infectious disease outbreaks, estimates of time-varying pathogen transmissibility, such as the instantaneous reproduction number R(t) or epidemic growth rate r(t), are used to inform decision-making by public health authorities. For directly transmitted infectious diseases, the renewal equation framework is a widely used method for measuring time-varying transmissibility. The framework uses information on the typical time elapsing between an infection and the offspring infections (quantified by the generation time distribution), and R(t), to describe the rate at which currently infected individuals generate new infections. For diseases with transmission cycles involving hosts and vectors, however, renewal equation models have been far less used. This is likely due to difficulties in mechanistically defining generation times that can capture the complexity of multi-stage, human-vector relationships. Here, using dengue as an example, we provide general renewal equations that are derived from first principles using age-structured systems of coupled partial differential equations across human and vector sub-populations. Our framework tracks the multi-stage transmission cycle over calendar time and across stage-specific ages, resulting in governing renewal equations that quantify how the rate at which new infections are generated from existing infections depends on stage-specific processes. The framework provides a foundation on which to base inferential frameworks for estimating R(t) and r(t) for infectious diseases with multiple stages in the transmission cycle

q-bio.PE

Ten simple rules for training scientists to make better software

Computational methods and associated software implementations are central to every field of scientific investigation. Modern biological research, particularly within systems biology, has relied heavily on the development of software tools to process and organize increasingly large datasets, simulate complex mechanistic models, provide tools for the analysis and management of data, and visualize and organize outputs. However, developing high-quality research software requires scientists to develop a host of software development skills, and teaching these skills to students is challenging. There has been a growing importance placed on ensuring reproducibility and good development practices in computational research. However, less attention has been devoted to informing the specific teaching strategies which are effective at nurturing in researchers the complex skillset required to produce high-quality software that, increasingly, is required to underpin both academic and industrial biomedical research. Recent articles in the Ten Simple Rules collection have discussed the teaching of foundational computer science and coding techniques to biology students. We advance this discussion by describing the specific steps for effectively teaching the necessary skills scientists need to develop sustainable software packages which are fit for (re-)use in academic research or more widely. Although our advice is likely to be applicable to all students and researchers hoping to improve their software development skills, our guidelines are directed towards an audience of students that have some programming literacy but little formal training in software development or engineering, typical of early doctoral students. These practices are also applicable outside of doctoral training environments, and we believe they should form a key part of postgraduate training schemes more generally in the life sciences.

cs.CY

EpiGeoPop: A Tool for Developing Spatially Accurate Country-level Epidemiological Models

Mathematical models play a crucial role in understanding the spread of infectious disease outbreaks and influencing policy decisions. These models aid pandemic preparedness by predicting outcomes under hypothetical scenarios and identifying weaknesses in existing frameworks. However, their accuracy, utility, and comparability are being scrutinized. Agent-based models (ABMs) have emerged as a valuable tool, capturing population heterogeneity and spatial effects, particularly when assessing intervention strategies. Here we present EpiGeoPop, a user-friendly tool for rapidly preparing spatially accurate population configurations of entire countries. EpiGeoPop helps to address the problem of complex and time-consuming model set up in ABMs, specifically improving the integration of spatial detail. We subsequently demonstrate the importance of accurate spatial detail in ABM simulations of disease outbreaks using Epiabm, an ABM based on Imperial College London's CovidSim with improved modularity, documentation and testing. Our investigation involves the interplay between population density, the implementation of spatial transmission, and realistic interventions implemented in Epiabm.

q-bio.PE

Understanding the impact of numerical solvers on inference for differential equation models

Most ordinary differential equation (ODE) models used to describe biological or physical systems must be solved approximately using numerical methods. Perniciously, even those solvers which seem sufficiently accurate for the forward problem, i.e., for obtaining an accurate simulation, may not be sufficiently accurate for the inverse problem, i.e., for inferring the model parameters from data. We show that for both fixed step and adaptive step ODE solvers, solving the forward problem with insufficient accuracy can distort likelihood surfaces, which may become jagged, causing inference algorithms to get stuck in local "phantom" optima. We demonstrate that biases in inference arising from numerical approximation of ODEs are potentially most severe in systems involving low noise and rapid nonlinear dynamics. We reanalyze an ODE changepoint model previously fit to the COVID-19 outbreak in Germany and show the effect of the step size on simulation and inference results. We then fit a more complicated rainfall-runoff model to hydrological data and illustrate the importance of tuning solver tolerances to avoid distorted likelihood surfaces. Our results indicate that when performing inference for ODE model parameters, adaptive step size solver tolerances must be set cautiously and likelihood surfaces should be inspected for characteristic signs of numerical issues.

math.ST

A capillary problem for spacelike mean curvature flow in a cone of Minkowski space

Consider a convex cone in three-dimensional Minkowski space which either contains the lightcone or is contained in it. This work considers mean curvature flow of a proper spacelike strictly mean convex disc in the cone which is graphical with respect to its rays. Its boundary is required to have constant intersection angle with the boundary of the cone. We prove that the corresponding parabolic boundary value problem for the graph admits a solution for all time which rescales to a self-similarly expanding solution.

math.DG

Epidemiological Agent-Based Modelling Software (Epiabm)

Epiabm is a fully tested, open-source software package for epidemiological agent-based modelling, re-implementing the well-known CovidSim model from the MRC Centre for Global Infectious Disease Analysis at Imperial College London. It has been developed as part of the first-year training programme in the EPSRC SABS:R3 Centre for Doctoral Training at the University of Oxford. The model builds an age-stratified, spatially heterogeneous population and offers a modular approach to configure and run epidemic scenarios, allowing for a broad scope of investigative and comparative studies. Two simulation backends are provided: a pedagogical Python backend (with full functionality) and a high performance C++ backend for use with larger population simulations. Both are highly modular, with comprehensive testing and documentation for ease of understanding and extensibility. Epiabm is publicly available through GitHub at https://github.com/SABS-R3-Epidemiology/epiabm.

q-bio.PE

Autocorrelated measurement processes and inference for ordinary differential equation models of biological systems

Ordinary differential equation models are used to describe dynamic processes across biology. To perform likelihood-based parameter inference on these models, it is necessary to specify a statistical process representing the contribution of factors not explicitly included in the mathematical model. For this, independent Gaussian noise is commonly chosen, with its use so widespread that researchers typically provide no explicit justification for this choice. This noise model assumes `random' latent factors affect the system in ephemeral fashion resulting in unsystematic deviation of observables from their modelled counterparts. However, like the deterministically modelled parts of a system, these latent factors can have persistent effects on observables. Here, we use experimental data from dynamical systems drawn from cardiac physiology and electrochemistry to demonstrate that highly persistent differences between observations and modelled quantities can occur. Considering the case when persistent noise arises due only to measurement imperfections, we use the Fisher information matrix to quantify how uncertainty in parameter estimates is artificially reduced when erroneously assuming independent noise. We present a workflow to diagnose persistent noise from model fits and describe how to remodel accounting for correlated errors.

stat.ME

Nonlocal estimates for the Volume Preserving Mean Curvature Flow and applications

We obtain estimates on nonlocal quantities appearing in the Volume Preserving Mean Curvature Flow (VPMCF) in the closed, Euclidean setting. As a result we demonstrate that blowups of finite time singularities of VPMCF are ancient solutions to Mean Curvature Flow (MCF), prove that monotonicity methods may always be applied at finite times and obtain information on the asymptotics of the flow.

math.DG

A note on Alexandrov immersed mean curvature flow

We demonstrate that the property of being Alexandrov immersed is preserved along mean curvature flow. Furthermore, we demonstrate that mean curvature flow techniques for mean convex embedded flows such as noncollapsing and gradient estimates also hold in this setting. We also indicate the necessary modifications to the work of Brendle--Huisken to allow for mean curvature flow with surgery for the Alexandrov immersed, $2$-dimensional setting.

math.DG

Using flexible noise models to avoid noise model misspecification in inference of differential equation time series models

When modelling time series, it is common to decompose observed variation into a "signal" process, the process of interest, and "noise", representing nuisance factors that obfuscate the signal. To separate signal from noise, assumptions must be made about both parts of the system. If the signal process is incorrectly specified, our predictions using this model may generalise poorly; similarly, if the noise process is incorrectly specified, we can attribute too much or too little observed variation to the signal. With little justification, independent Gaussian noise is typically chosen, which defines a statistical model that is simple to implement but often misstates system uncertainty and may underestimate error autocorrelation. There are a range of alternative noise processes available but, in practice, none of these may be entirely appropriate, as actual noise may be better characterised as a time-varying mixture of these various types. Here, we consider systems where the signal is modelled with ordinary differential equations and present classes of flexible noise processes that adapt to a system's characteristics. Our noise models include a multivariate normal kernel where Gaussian processes allow for non-stationary persistence and variance, and nonparametric Bayesian models that partition time series into distinct blocks of separate noise structures. Across the scenarios we consider, these noise processes faithfully reproduce true system uncertainty: that is, parameter estimate uncertainty when doing inference using the correct noise model. The models themselves and the methods for fitting them are scalable to large datasets and could help to ensure more appropriate quantification of uncertainty in a host of time series models.

stat.ME

Model Evidence with Fast Tree Based Quadrature

High dimensional integration is essential to many areas of science, ranging from particle physics to Bayesian inference. Approximating these integrals is hard, due in part to the difficulty of locating and sampling from regions of the integration domain that make significant contributions to the overall integral. Here, we present a new algorithm called Tree Quadrature (TQ) that separates this sampling problem from the problem of using those samples to produce an approximation of the integral. TQ places no qualifications on how the samples provided to it are obtained, allowing it to use state-of-the-art sampling algorithms that are largely ignored by existing integration algorithms. Given a set of samples, TQ constructs a surrogate model of the integrand in the form of a regression tree, with a structure optimised to maximise integral precision. The tree divides the integration domain into smaller containers, which are individually integrated and aggregated to estimate the overall integral. Any method can be used to integrate each individual container, so existing integration methods, like Bayesian Monte Carlo, can be combined with TQ to boost their performance. On a set of benchmark problems, we show that TQ provides accurate approximations to integrals in up to 15 dimensions; and in dimensions 4 and above, it outperforms simple Monte Carlo and the popular Vegas method.

stat.ML