Searcharxiv⌕ Search

arXiv subjects

Jennifer A. Flegg

Publications and source records attributed to Jennifer A. Flegg.

At least 19 recordsLinked to original sources

Multiscale Modelling of Birth-Death Processes

Many biological systems exhibit multiscale dynamics, where some species occur in high copy numbers while others remain rare. This heterogeneity necessitates hybrid modelling approaches: deterministic models are computationally efficient but inaccurate for low-count species, while fully stochastic simulations are accurate but prohibitively expensive. Threshold-based hybrid methods, such as the Jump--Switch--Flow (JSF) algorithm, address this by simulating low-count species stochastically and high-count species deterministically, switching at a user-chosen threshold $Ω$. In such methods, the choice of $Ω$ controls the trade-off between computational cost and accuracy, but is typically made by trial-and-error: there is no principled way to choose $Ω$ a priori for a given observable of interest. We close this gap for extinction probability. Our contribution is a computable, method-agnostic error bound that quantifies the discrepancy introduced by the threshold and yields an explicit rule for selecting $Ω$ to meet a user-specified error tolerance. We formalise JSF as a piecewise-deterministic Markov process and derive backward equations for extinction under exact and hybrid dynamics. Near extinction boundaries, the complex nonlinear dynamics reduce to tractable time-inhomogeneous linear birth--death processes; this structure yields a rigorous error decomposition into early and late excursions, whose dominant term becomes a fast, actionable heuristic requiring only the solution of a scalar Riccati equation. Monte Carlo studies on a stochastic Lotka--Volterra model confirm that the heuristic reliably upper-bounds the empirical error in extinction probability across a wide parameter range. The framework depends only on the birth and death rates near extinction, not on the specific simulation method, and therefore applies beyond JSF to any threshold-based hybrid scheme.

q-bio.PE↗

A Bayesian spatio-temporal nearest neighbor Gaussian process model for pooled genetic data

Large scale genetic datasets often aggregate the total allele counts of distinct genetic markers. Inferring haplotype frequencies (i.e.\ the frequency of multimarker alleles) from these pooled data is a challenge. Previous spatio-temporal modelling in this context has been limited to 3 markers due to the computational cost. In this work, we propose a nearest neighbor Gaussian process (NNGP) model to improve scaling with the number of markers and observations. To infer the parameters of our model, we develop a novel sequential Monte Carlo squared algorithm, which uses particle Gibbs with ancestor sampling to mutate the NNGP function values. The latter has a linear cost in the number of observations and the number of NNGPs, and can be applied to a broad range of NNGP models. As a case study, we analyse genetic data relating to antimalarial drug resistance in Africa, and show our scaling results empirically on a 3 and 6 genetic marker dataset.

stat.ME↗

Reliable model selection in the presence of parameter non-identifiability

Mathematical models are invaluable for understanding and predicting how biological systems behave, although their construction requires specifying mechanisms and relationships that are often not perfectly known. In the presence of multiple competing models, model uncertainty should be accounted for when performing inference based on available data. Bayesian model selection is a framework for testing mechanistic hypotheses and generating predictions under model uncertainty, which generally requires computation of the model evidence. In this work, we investigate the reliability of evidence computation methods when parameter non-identifiability -- the inability to distinguish between parameter values given available data -- is present, and find that deterministic evidence approximations can produce misleading model selection results because their underlying assumptions are violated. We propose a novel implementation of adaptive multiple importance sampling for evidence estimation, and demonstrate its robustness against non-identifiability. We use ecological case studies to demonstrate how simple model selection methods fail to produce accurate results, whereas our method yields model selection results that are comparable to those obtained by Markov chain Monte Carlo methods at substantially lower computational cost. Given the pervasiveness of parameter non-identifiability in mathematical biology, this work provides a practical approach to reliable model selection in the presence of poorly identified parameters.

stat.ME↗

Reduced-Precision Stochastic Simulation for Mathematical Biology

The stochastic simulation algorithm (SSA) is widely used to perform exact forward simulation of discrete stochastic processes in biology. However, the computational cost, driven by sequential event-by-event sampling across large ensembles, remains a computational barrier. We investigate whether reduced-precision floating-point arithmetic can accelerate SSA without degrading statistical fidelity, drawing on the success of reduced-precision methods in weather and climate modelling. We evaluate two strategies across five canonical models (birth--death, Schlögl, Telegraph, dimerisation, repressilator): (i) mixed precision, computing propensities in 16-bit while maintaining accumulators in 32-bit; and (ii) uniform precision, performing all arithmetic in 16-bit. Mixed-precision SSA produces ensemble statistics that closely match the 64-bit reference for all models, as measured by Kolmogorov--Smirnov tests and Wasserstein distances. Under uniform precision, deterministic rounding introduces systematic biases across several models, with catastrophic failures in some cases. Stochastic rounding (SR) and propensity normalisation eliminate these biases, restoring distributional fidelity across all models tested (KS $p > 0.05$). Our results establish mixed-precision SSA with SR as a viable acceleration strategy for mathematical biology: 16-bit formats shrink per-variable data size by $2$--$4\times$ relative to \texttt{fp32}/\texttt{fp64}, yielding comparable reductions in memory footprint and up to $\sim 1.5\times$ wall-clock speedup on CPU hardware that lacks native 16-bit arithmetic. As a hardware-level acceleration, mixed-precision SSA complements algorithmic methods such as tau-leaping and maps naturally onto modern GPU and TPU architectures with native 16-bit arithmetic.

q-bio.QM↗

Quantifying structural uncertainty in chemical reaction network inference

Dynamical systems in biology are complex, and one often does not have comprehensive knowledge about the interactions involved. Chemical reaction network (CRN) inference aims to identify, from observing species concentrations over time, the unknown reactions between the species. Existing approaches such as sparse regularisation largely focus on identifying a single, most likely CRN, without addressing uncertainty about the network structure. However, it is important to quantify structural uncertainty to have confidence in our inference and predictions. In this work, we explore how effective sparse regularisation methods are for quantifying structural uncertainty. Locally optimal solutions to sparse regularisation are mapped to CRN structures; however, it is unclear whether this approach encompasses all plausible CRNs. We find that inducing sparsity with nonconvex penalty functions results in better coverage of the plausible CRNs compared to the popular lasso regularisation. To validate our approach, we apply our methods to real-world data examples, and are able to simultaneously recover reactions proposed across multiple literature sources for a reaction system. Our emphasis on network-level probabilities enables a novel, hierarchical representation of structural ambiguities in the space of CRNs. This representation translates into alternative reaction pathways suggested by the available data, thus guiding the efforts of future experimental design.

stat.ME↗

Balancing training load, rest and musculoskeletal injury risk: a mathematical modelling study in Thoroughbred racehorses

Musculoskeletal injuries (MSI) in Thoroughbred racehorses are a leading cause of death and premature retirement in racehorses and are heavily influenced by training practices. Greater distances of high-speed galloping accumulated during racing campaigns are associated with MSI. Bone injury is the most common MSI, and understanding how training practices influence bone damage accumulation is critical for improving both horse welfare and racing outcomes. This study builds on an existing mathematical model of bone adaptation and damage to investigate the impact of different training programs on bone injury risk. Several training programs (three progressive, four race-fit, six rest programs and two with rest replaced by low-intensity training) were constructed to reflect representative practices undertaken by professional trainers in Victoria, Australia. Training programs varied in training volume, rest frequency and program duration. Lower volume training programs that included high-speed training, achieved sufficient bone adaptation with less accumulation of bone damage, and subsequently lower risk of bone failure. In addition, incorporating more frequent rests (at least 2 per year) and/or longer rest periods (at least 6 weeks) reduced bone damage due to the extended opportunity to remove and repair bone damage. These results provide an in-silico mathematical model of the bone response to training, demonstrating the effects of training programs on bone adaptation, damage formation and repair. The findings can guide the design of training programs that balance both bone adaptation and bone health throughout horses racing career.

q-bio.PE↗

A nonparametric approach to practical identifiability of nonlinear mixed effects models

Mathematical modelling is a widely used approach to understand and interpret clinical trial data. This modelling typically involves fitting mechanistic mathematical models to data from individual trial participants. Despite the widespread adoption of this individual-based fitting, it is becoming increasingly common to take a hierarchical approach to parameter estimation, where modellers characterize the population parameter distributions, rather than considering each individual independently. This hierarchical parameter estimation is standard in pharmacometric modelling. However, many of the existing techniques for parameter identifiability do not immediately translate from the individual-based fitting to the hierarchical setting. Here, we propose a nonparametric approach to study practical identifiability within a hierarchical parameter estimation framework. We focus on the commonly used nonlinear mixed effects framework and investigate two well-studied examples from the pharmacometrics and viral dynamics literature to illustrate the potential utility of our approach.

stat.ME↗

A hybrid framework for compartmental models enabling simulation-based inference

Multi-scale systems often exhibit a combination of stochastic and deterministic dynamics. In compartmental models, low occupancy compartments tend to exhibit stochastic dynamics while high occupancy compartments tend to follow deterministic dynamics. Representing both dynamics with existing methods is challenging. Failing to account for stochasticity in small populations can produce ``atto-foxes'', for example in the Lotka-Volterra ordinary differential equation (ODE) model. This limitation becomes problematic when studying the extinction of species or the clearance of infection, but it can be overcome by using discrete stochastic models, such as continuous time Markov chains (CTMCs). Unfortunately, simulating CTMCs is impractical for many realistic models, where discrete events have very high frequencies. In this work, we develop a novel mathematical framework to couple continuous ODEs and discrete CTMCs: ``Jump-Switch-Flow'' (JSF). In this framework, compartments can reach extinct states (``absorbing states''), thereby resolving atto-fox-type problems. JSF has the desired behaviours of exact CTMC simulation, but is substantially computationally faster than existing alternatives, by at least one order of magnitude, and can even obtain constant scaling, irrespective of compartment occupancy. We demonstrate JSF's utility for simulation-based inference, particularly multi-scale problems, with several case-studies. In a simulation study, we demonstrate how JSF can enable a more nuanced analysis of the efficacy of public health interventions. We also carry out a novel analysis of longitudinal within-host data from SARS-CoV-2 infections to quantify the timing of viral clearance. In this work, we show how JSF offers a novel approach to compartmental model simulation.

q-bio.PE↗

Evaluating interventions for Plasmodium vivax forest malaria using a three-scale mathematical model

The rising proportion of Plasmodium vivax cases concentrated in forest-fringe areas across the Greater Mekong Subregion highlights the importance of pharmaceutical and mosquito control techniques specifically targeted towards forest-going populations. To mathematically assess best-possible antimalarial interventions in the context of hypnozoite reactivation and seasonal forest migration, we extend a previously developed three-scale integro-differential equations model of P. vivax transmission. In particular, we fit the model to data gathered over a four-year period in Vietnam to gain insight into local P. vivax dynamics and validate the model's ability to capture epidemiological trends. The calibrated model is then used to generate optimal schedules for mass-drug administration (MDA) in forest-goers and gauge the efficacy of vector control techniques (such as long-lasting insecticide nets and indoor residual spraying) in forest-adjacent areas. Our results highlight the dependence of optimal MDA timing on the demographics of the human population, the importance of interventions targeting the mosquito bite rate, and the need for efficacy in hypnozoite-targeting antimalarial drugs.

q-bio.QM↗

Spatio-temporal agent-based modelling of malaria

Plasmodium falciparum is responsible for the majority of malaria morbidity and mortality each year. Malaria transmission rates vary by location and time of year due to climate and environmental conditions. We show the impact of these factors by developing a stochastic spatiotemporal agent-based malaria model that captures the impact of spatially distributed interventions on malaria transmission. Our model uses spatiotemporal estimates of mosquito climatic suitability and household location data to model the interaction between human and mosquito agents. We apply our model to investigate how strategies for distributing interventions to households in Vietnam impact the disease burden. Our study shows that providing some level of protection to a wide range of households reduces malaria prevalence more compared to providing a strong level of protection to a limited number of households.

q-bio.PE↗

Sunlight-heated refugia protect frogs from chytridiomycosis: a mathematical modelling study

The fungal disease Chytridiomycosis poses a threat to frog populations worldwide. It has driven over 90 amphibian species to extinction and severely affected hundreds more. Difficulties in disease management have shown a need for novel conservation approaches. We present a novel mathematical model for chytridiomycosis transmission in frogs that includes the natural history of infection, to test the hypothesis that sunlight-heated refugia reduce transmission. The model is fit using approximate Bayesian computation to experimental data where a cohort of frogs, a fixed subset of which had cleared a prior infection, were provided access to either sunlight-heated or shaded refugia. Using our model, we can estimate the extent to which prior chytridiomycosis infection protects against subsequent infection, and quantify the effect of sunlight-heating of refugia. Results estimate a 40% reduction in chytridiomycosis transmission when frogs have access to sunlight-heated refugia, compared to shaded refugia. This strongly supports the hypothesis that the sunlight-heated refugia reduce disease transmission. Frogs that were infected and recovered were estimated to have a reduction in susceptibility of approximately 97% compared to frogs with no prior infection. This research provides quantitative evidence supporting sunlight-heated refugia as an effective disease management tool for chytridiomycosis in frog populations. By estimating both the impact of refugia and the protective effects of prior infection, the model provides an evidence base for implementing sunlight-heated refugia as part of amphibian conservation strategies. This work represents an important first step in using mathematical modelling to inform policy on the design and implementation of habitat-based interventions to support amphibian population recovery and long-term sustainability.

q-bio.PE↗

Accurate stochastic simulation algorithm for multiscale models of infectious diseases

In the infectious disease literature, significant effort has been devoted to studying dynamics at a single scale. For example, compartmental models describing population-level dynamics are often formulated using differential equations. In cases where small numbers or noise play a crucial role, these differential equations are replaced with memoryless Markovian models, where discrete individuals can be members of a compartment and transition stochastically. Classic stochastic simulation algorithms, such as the next reaction method, can be employed to solve these Markovian models exactly. The intricate coupling between models at different scales underscores the importance of multiscale modelling in infectious diseases. However, several computational challenges arise when the multiscale model becomes non-Markovian. In this paper, we address these challenges by developing a novel exact stochastic simulation algorithm. We apply it to a showcase multiscale system where all individuals share the same deterministic within-host model while the population-level dynamics are governed by a stochastic formulation. We demonstrate that as long as the within-host information is harvested at a reasonable resolution, the novel algorithm will always be accurate. Furthermore, our implementation is still efficient even at finer resolutions. Beyond infectious disease modelling, the algorithm is widely applicable to other multiscale systems, providing a versatile, accurate, and computationally efficient framework.

q-bio.PE↗

Investigation of P. Vivax Elimination via Mass Drug Administration

Plasmodium vivax is the most geographically widespread malaria parasite due to its ability to remain dormant (as a hypnozoite) in the human liver and subsequently reactivate. Given the majority of P. vivax infections are due to hypnozoite reactivation, targeting the hypnozoite reservoir with a radical cure is crucial for achieving P. vivax elimination. Stochastic effects can strongly influence dynamics when disease prevalence is low or when the population size is small. Hence, it is important to account for this when modelling malaria elimination.cWe use a stochastic multiscale model of P. vivax transmission to study the impacts of multiple rounds of mass drug administration (MDA) with a radical cure, accounting for superinfection and hypnozoite dynamics. Our results indicate multiple rounds of MDA with a high-efficacy drug are needed to achieve a substantial probability of elimination. This work has the potential to help guide P. vivax elimination strategies by quantifying elimination probabilities for an MDA approach.

q-bio.PE↗

A spatial multiscale mathematical model of Plasmodium vivax transmission

The epidemiological behavior of Plasmodium vivax malaria occurs across spatial scales including within-host, population, and metapopulation levels. On the within-host scale, P. vivax sporozoites inoculated in a host may form latent hypnozoites, the activation of which drives secondary infections and accounts for a large proportion of P. vivax illness; on the metapopulation level, the coupled human-vector dynamics characteristic of the population level are further complicated by the migration of human populations across patches with different malaria forces of (re-)infection. To explore the interplay of all three scales in a single two-patch model of Plasmodium vivax dynamics, we construct and study a system of eight integro-differential equations with periodic forcing (arising from the single-frequency sinusoidal movement of a human sub-population). Under the numerically-informed ansatz that the limiting solutions to the system are closely bounded by sinusoidal ones for certain regions of parameter space, we derive a single nonlinear equation from which all approximate limiting solutions may be drawn, and devise necessary and sufficient conditions for the equation to have only a disease-free solution. Our results illustrate the impact of movement on P. vivax transmission and suggest a need to focus vector control efforts on forest mosquito populations. The three-scale model introduced here provides a more comprehensive framework for studying the clinical, behavioral, and geographical factors underlying P. vivax malaria endemicity.

q-bio.PE↗

Mathematical models of Plasmodium vivax transmission: a scoping review

Plasmodium vivax is one of the most geographically widespread malaria parasites in the world due to its ability to remain dormant in the human liver as hypnozoites and subsequently reactivate after the initial infection (i.e. relapse infections). More than 80% of P. vivax infections are due to hypnozoite reactivation. Mathematical modelling approaches have been widely applied to understand P. vivax dynamics and predict the impact of intervention outcomes. In this article, we provide a scoping review of mathematical models that capture P. vivax transmission dynamics published between January 1988 and May 2023 to provide a comprehensive summary of the mathematical models and techniques used to model P. vivax dynamics. We aim to assist researchers working on P. vivax transmission and other aspects of P. vivax malaria by highlighting best practices in currently published models and highlighting where future model development is required. We provide an overview of the different strategies used to incorporate the parasite's biology, use of multiple scales (within-host and population-level), superinfection, immunity, and treatment interventions. In most of the published literature, the rationale for different modelling approaches was driven by the research question at hand. Some models focus on the parasites' complicated biology, while others incorporate simplified assumptions to avoid model complexity. Overall, the existing literature on mathematical models for P. vivax encompasses various aspects of the parasite's dynamics. We recommend that future research should focus on refining how key aspects of P. vivax dynamics are modelled, including spatial heterogeneity in exposure risk, the accumulation of hypnozoite variation, the interaction between P. falciparum and P. vivax, acquisition of immunity, and recovery under superinfection.

q-bio.PE↗

Haplotype frequency inference from pooled genetic data with a latent multinomial model

In genetic studies, haplotype data provide more refined information than data about separate genetic markers. However, large-scale studies that genotype hundreds to thousands of individuals may only provide results of pooled data, where only the total allele counts of each marker in each pool are reported. Methods for inferring haplotype frequencies from pooled genetic data that scale well with pool size rely on a normal approximation, which we observe to produce unreliable inference when applied to real data. We illustrate cases where the approximation breaks down, due to the normal covariance matrix being near-singular. As an alternative to approximate methods, in this paper we propose exact methods to infer haplotype frequencies from pooled genetic data based on a latent multinomial model, where the observed allele counts are considered integer combinations of latent, unobserved haplotype counts. One of our methods, latent count sampling via Markov bases, achieves approximately linear runtime with respect to pool size. Our exact methods produce more accurate inference over existing approximate methods for synthetic data and for data based on haplotype information from the 1000 Genomes Project. We also demonstrate how our methods can be applied to time-series of pooled genetic data, as a proof of concept of how our methods are relevant to more complex hierarchical settings, such as spatiotemporal models.

stat.ME↗

Comparison of new computational methods for geostatistical modelling of malaria

Geostatistical analysis of health data is increasingly used to model spatial variation in malaria prevalence, burden, and other metrics. Traditional inference methods for geostatistical modelling are notoriously computationally intensive, motivating the development of newer, approximate methods. The appeal of faster methods is particularly great as the size of the region and number of spatial locations being modelled increases. Methods We present an applied comparison of four proposed `fast' geostatistical modelling methods and the software provided to implement them -- Integrated Nested Laplace Approximation (INLA), tree boosting with Gaussian processes and mixed effect models (GPBoost), Fixed Rank Kriging (FRK) and Spatial Random Forests (SpRF). We illustrate the four methods by estimating malaria prevalence on two different spatial scales -- country and continent. We compare the performance of the four methods on these data in terms of accuracy, computation time, and ease of implementation. Results Two of these methods -- SpRF and GPBoost -- do not scale well as the data size increases, and so are likely to be infeasible for larger-scale analysis problems. The two remaining methods -- INLA and FRK -- do scale well computationally, however the resulting model fits are very sensitive to the user's modelling assumptions and parameter choices. Conclusions INLA and FRK both enable scalable geostatistical modelling of malaria prevalence data. However care must be taken when using both methods to assess the fit of the model to data and plausibility of predictions, in order to select appropriate model assumptions and approximation parameters.

stat.AP↗

Optimal interruption of P. vivax malaria transmission using mass drug administration

\textit{Plasmodium vivax} is the most geographically widespread malaria-causing parasite resulting in significant associated global morbidity and mortality. One of the factors driving this widespread phenomenon is the ability of the parasites to remain dormant in the liver. Known as hypnozoites, they reside in the liver following an initial exposure, before activating later to cause further infections, referred to as relapses. As around 79-96$\%$ of infections are attributed to relapses, we expect it will be highly impactful to apply treatment to target the hypnozoite reservoir to eliminate \textit{P. vivax}. Treatment with a radical cure to target the hypnozoite reservoir is a potential tool to control or eliminate \textit{P. vivax}. We have developed a multiscale mathematical model as a system of integro-differential equations that captures the complex dynamics of \textit{P. vivax} hypnozoites and the effect of hypnozoite relapse on disease transmission. Here, we use our model to study the anticipated effect of radical cure treatment administered via a mass drug administration (MDA) program. We implement multiple rounds of MDA with a fixed interval between rounds, starting from different steady-state disease prevalences. We then construct an optimisation model to obtain the optimal MDA interval. We also incorporate mosquito seasonality in our model to study its effect on the optimal treatment regime. We find that the effect of MDA interventions is temporary and depends on the pre-intervention disease prevalence (and choice of model parameters) as well as the number of MDA rounds under consideration. We find radical cure alone may not be enough to lead to \textit{P. vivax} elimination under our mathematical model (and choice of model parameters) since the prevalence of infection eventually returns to pre-MDA levels.

q-bio.PE↗