SearcharxivSearch

arXiv subjects

Richard Reeve

Publications and source records attributed to Richard Reeve.

7 recordsLinked to original sources

Visualization for Epidemiological Modelling: Challenges, Solutions, Reflections & Recommendations

We report on an ongoing collaboration between epidemiological modellers and visualization researchers by documenting and reflecting upon knowledge constructs -- a series of ideas, approaches and methods taken from existing visualization research and practice -- deployed and developed to support modelling of the COVID-19 pandemic. Structured independent commentary on these efforts is synthesized through iterative reflection to develop: evidence of the effectiveness and value of visualization in this context; open problems upon which the research communities may focus; guidance for future activity of this type; and recommendations to safeguard the achievements and promote, advance, secure and prepare for future collaborations of this kind. In describing and comparing a series of related projects that were undertaken in unprecedented conditions, our hope is that this unique report, and its rich interactive supplementary materials, will guide the scientific community in embracing visualization in its observation, analysis and modelling of data as well as in disseminating findings. Equally we hope to encourage the visualization community to engage with impactful science in addressing its emerging data challenges. If we are successful, this showcase of activity may stimulate mutually beneficial engagement between communities with complementary expertise to address problems of significance in epidemiology and beyond. https://ramp-vis.github.io/RAMPVIS-PhilTransA-Supplement/

cs.HC

FAIR Data Pipeline: provenance-driven data management for traceable scientific workflows

Modern epidemiological analyses to understand and combat the spread of disease depend critically on access to, and use of, data. Rapidly evolving data, such as data streams changing during a disease outbreak, are particularly challenging. Data management is further complicated by data being imprecisely identified when used. Public trust in policy decisions resulting from such analyses is easily damaged and is often low, with cynicism arising where claims of "following the science" are made without accompanying evidence. Tracing the provenance of such decisions back through open software to primary data would clarify this evidence, enhancing the transparency of the decision-making process. Here, we demonstrate a Findable, Accessible, Interoperable and Reusable (FAIR) data pipeline developed during the COVID-19 pandemic that allows easy annotation of data as they are consumed by analyses, while tracing the provenance of scientific outputs back through the analytical source code to data sources. Such a tool provides a mechanism for the public, and fellow scientists, to better assess the trust that should be placed in scientific evidence, while allowing scientists to support policy-makers in openly justifying their decisions. We believe that tools such as this should be promoted for use across all areas of policy-facing research.

q-bio.QM

Dynamic virtual ecosystems as a tool for detecting large-scale responses of biodiversity to environmental and land-use change

Ecosystems are governed by dynamic processes such as competition for resources, reproduction and dispersal. These shape their biodiversity and how the system responds to change. Current approaches to modelling ecosystems, especially plants, focus on either describing fine-scale processes for individual species or broad-scale patterns for limited groups of plant functional types. Digitisation of herbarium and other plant records has unlocked a wealth of information that can be used to drive models of plant communities and make predictions for their future under different scenarios of climate change. The advent of increased computational capacity and fast, high level programming languages allows for simulation of such landscapes at unprecedented scales. Here, we demonstrate a tool for Ecosystem Simulation through Integrated Species Trait-Environment Modelling (EcoSISTEM), which models plant species across multiple ecosystem sizes, from patches and small islands to regions and entire continents. These simulated ecosystems support the ability to generate many different types of habitat, as well as reproducing different disturbance scenarios such as climate change, habitat loss and invasion. EcoSISTEM also reproduces examples of real-world species distributions by integrating plant occurrence records and global climate reconstructions to simulate plant species throughout the continent of Africa for the past century. EcoSISTEM allows us to flexibly explore the dynamics of tens of thousands of species interacting across a continent. The code parallelises efficiently across multiple nodes on high performance computing platforms, and has been scaled up to run on over 1000 cores. It allows us to study the impact of changes to climate, resources and habitat and investigate real-life mechanisms surrounding climate change and biodiversity loss.

q-bio.QM

Estimation of temporal covariances in pathogen dynamics using Bayesian multivariate autoregressive models

It is well recognised that animal and plant pathogens form complex ecological communities of interacting organisms within their hosts. Although community ecology approaches have been applied to determine pathogen interactions at the within-host scale, methodologies enabling robust inference of the epidemiological impact of pathogen interactions are lacking. Here we developed a novel statistical framework to identify statistical covariances from the infection time-series of multiple pathogens simultaneously. Our framework extends Bayesian multivariate disease mapping models to analyse multivariate time series data by accounting for within- and between-year dependencies in infection risk and incorporating a between-pathogen covariance matrix which we estimate. Importantly, our approach accounts for possible confounding drivers of temporal patterns in pathogen infection frequencies, enabling robust inference of pathogen-pathogen interactions. We illustrate the validity of our statistical framework using simulated data and applied it to diagnostic data available for five respiratory viruses co-circulating in a major urban population between 2005 and 2013: adenovirus, human coronavirus, human metapneumovirus, influenza B virus and respiratory syncytial virus. We found positive and negative covariances indicative of epidemiological interactions among specific virus pairs. This statistical framework enables a community ecology perspective to be applied to infectious disease epidemiology with important utility for public health planning and preparedness.

stat.ME

Improving the identification of antigenic sites in the H1N1 Influenza virus through accounting for the experimental structure in a sparse hierarchical Bayesian model

Understanding how genetic changes allow emerging virus strains to escape the protection afforded by vaccination is vital for the maintenance of effective vaccines. In the current work, we use structural and phylogenetic differences between pairs of virus strains to identify important antigenic sites on the surface of the influenza A(H1N1) virus through the prediction of haemagglutination inhibition (HI) assay, pairwise measures of the antigenic similarity of virus strains. We propose a sparse hierarchical Bayesian model that can deal with the pairwise structure and inherent experimental variability in the H1N1 data through the introduction of latent variables. The latent variables represent the underlying HI assay measurement of any given pair of virus strains and help account for the fact that for any HI assay measurement between the same pair of virus strains, the difference in the viral sequence remains the same. Through accurately representing the structure of the H1N1 data, the model is able to select virus sites which are antigenic, while its latent structure achieves the computational efficiency required to deal with large virus sequence data, as typically available for the influenza virus. In addition to the latent variable model, we also propose a new method, block integrated Widely Applicable Information Criterion (biWAIC), for selecting between competing models. We show how this allows us to effectively select the random effects when used with the proposed model and apply both methods to an A(H1N1) dataset.

stat.AP

How to partition diversity

Diversity measurement underpins the study of biological systems, but measures used vary across disciplines. Despite their common use and broad utility, no unified framework has emerged for measuring, comparing and partitioning diversity. The introduction of information theory into diversity measurement has laid the foundations, but the framework is incomplete without the ability to partition diversity, which is central to fundamental questions across the life sciences: How do we prioritise communities for conservation? How do we identify reservoirs and sources of pathogenic organisms? How do we measure ecological disturbance arising from climate change? The lack of a common framework means that diversity measures from different fields have conflicting fundamental properties, allowing conclusions reached to depend on the measure chosen. This conflict is unnecessary and unhelpful. A mathematically consistent framework would transform disparate fields by delivering scientific insights in a common language. It would also allow the transfer of theoretical and practical developments between fields. We meet this need, providing a versatile unified framework for partitioning biological diversity. It encompasses any kind of similarity between individuals, from functional to genetic, allowing comparisons between qualitatively different kinds of diversity. Where existing partitioning measures aggregate information across the whole population, our approach permits the direct comparison of subcommunities, allowing us to pinpoint distinct, diverse or representative subcommunities and investigate population substructure. The framework is provided as a ready-to-use R package to easily test our approach.

q-bio.QM

Identifying the genetic basis of antigenic change in influenza A(H1N1)

Determining phenotype from genetic data is a fundamental challenge. Influenza A viruses undergo rapid antigenic drift and identification of emerging antigenic variants is critical to the vaccine selection process. Using former seasonal influenza A(H1N1) viruses, hemagglutinin sequence and corresponding antigenic data were analyzed in combination with 3-D structural information. We attributed variation in hemagglutination inhibition to individual amino acid substitutions and quantified their antigenic impact, validating a subset experimentally using reverse genetics. Substitutions identified as low-impact were shown to be a critical component of influenza antigenic evolution and by including these, as well as the high-impact substitutions often focused on, the accuracy of predicting antigenic phenotypes of emerging viruses from genotype was doubled. The ability to quantify the phenotypic impact of specific amino acid substitutions should help refine techniques that predict the fitness and evolutionary success of variant viruses, leading to stronger theoretical foundations for selection of candidate vaccine viruses.

q-bio.PE