SearcharxivSearch

arXiv subjects

Marta Blangiardo

Publications and source records attributed to Marta Blangiardo.

At least 19 recordsLinked to original sources

Comparing statistical learning models in wastewater-based epidemiology: An application to norovirus

Wastewater-based epidemiology (WBE) is an increasingly important tool for infectious disease surveillance, but there has been limited direct comparison of modelling approaches for predicting pathogen concentrations across space and time. We compare the predictive performance of six modelling approaches using norovirus in England as a case study, considering 3,232 wastewater samples from 152 sewage treatment works collected between May 2021 and March 2022 as part of the UK Environmental Monitoring for Health Protection programme. The benchmark was a Bayesian spatio-temporal model using Integrated Nested Laplace Approximation (INLA) with the Stochastic Partial Differential Equation approach (SPDE), compared with Lasso regression, Generalised Additive Models (GAM), Bayesian GAM, Extreme Gradient Boosting (XGBoost), and Random Forest. Models were evaluated using 10-fold spatial-block cross-validation with metrics including mean squared error, bias, correlation, empirical coverage of 95% prediction intervals, interval score, and computational cost. Random Forest was the best overall performing model, achieving the best interval score and nearest to nominal empirical coverage (95.3%), while maintaining point prediction accuracy comparable to XGBoost. The INLA-SPDE model also performed well, with near-nominal empirical coverage, the third best interval score, and consistently low bias across all evaluated metrics. Comparing predicted spatio-temporal trends, both models yielded similar spatial patterns but notable differences in uncertainty estimation. Our findings highlight a trade-off between predictive accuracy, uncertainty quantification, and computational efficiency. Ensemble machine learning methods are well suited to rapid prediction, whereas Bayesian geostatistical models remain valuable when probabilistic decision support is a priority for public health surveillance

stat.AP

A spatio-temporal block aggregation model for latent log Gaussian outcomes: application on modelling wastewater virus concentration in Wales

Wastewater-based epidemiology has emerged as a valuable tool for monitoring community-level infectious disease dynamics, providing population-wide signals that complement clinical surveillance. However, wastewater measurements are often observed as aggregated values over irregular spatial units. This work develops an approach to link an underlying spatially continuous processes and an aggregated outcome. We propose a spatio-temporal model for latent log-Gaussian outcomes that provides a coherent framework for inference and prediction, allowing the process to be integrated over arbitrary spatial configurations. This framework can also be used for subsequent analyses, such as linking wastewater signal to health outcomes at administrative areas. We use a Bayesian framework for inference via the linearised integrated nested Laplace approximation (INLA) approach. We apply the proposed methodology to model SARS-CoV-2 N1 gene copies in wastewater across 47 catchment areas in Wales from the beginning of August 2022 to the end of July 2023. The results demonstrate that the model captures spatial and temporal patterns and has good predictive performance. Results also show that estimated viral gene copies are strongly linked to positivity rates from COVID-19 PCR tests at the local authority level. Our findings highlight the importance of explicitly modelling block aggregation when analysing wastewater surveillance data. The proposed framework provides a flexible and principled approach for integrating environmental surveillance data into public health monitoring systems.

stat.AP

Joint Bayesian models for validating spatial health-event databases against a gold standard: separating global and local discrepancies

The reuse of medico-administrative and synthetic spatial data may overcome some limitations of population-based registries, provided rigorous validation is performed. However, no tool exists to spatially validate a candidate-for-reuse database (CFRD) against a gold standard (GS). We propose a Bayesian framework for two-dimensional (global and local) map-to-map validation of spatial health-event databases. We consider an error-model family (random [REM] and structured [SEM]) in which the CFRD is modelled as a departure from the GS. Both are compared with a shared component model (SCM). Global disagreement is assessed using the database-specific intercept difference ($RR_{\mathrm{global}}$), while local disagreement is measured by the exceedance probability of the database-specific error term. Disturbance scenarios included null, uniform, clustered, and random perturbations in the CFRD. Sensitivity, specificity, false detection rate, and Matthews Correlation Coefficient assessed detection performance. $RR_{\mathrm{global}}$ accurately recovered map-wide shifts across all models and scenarios. REM and SEM behaved were both sensitive and specific to local discrepancies. SCM was more conservative. Applied to Crohn's disease data from the EPIMAD registry and a CFRD, all models reached the same conclusion: the CFRD reproduced global and local spatial structures with an overall signal about 7\% lower. Extensions to other outcome distributions, spatio-temporal models and calibration constitute natural next steps. \textit{Keywords:} data reuse; spatial database validation; Bayesian hierarchical models; disease mapping; shared component model.

stat.ME

Spatially continuous modelling of aggregated outcome data

This work develops a block aggregation approach to spatial estimation and prediction when the response is observed at a coarse spatial scale, for example as counts of events in administrative areas, or blocks, while covariates are available at a finer spatial resolution, typically as raster images. Our approach specifies a linear predictor at the finer resolution as a combination of covariate effects and a latent, spatially continuous Gaussian process. This linear predictor then determines the distribution of the response through an inverse link function and spatial integration. We use a simulation study to evaluate the performance of the proposed approach in comparison to two industry standard approaches: a traditional geostatistical model that associates each response with the centroid of its block; and a Markov random field (MRF) approach that aggregates covariate data to block-level. As expected, the differences in performance among the three approaches are small with respect to block-level prediction. The rationale for, and advantage of, the block aggregation approach lies in its delivery of reliable inferences at whatever spatial resolution is required in a particular application. We describe two applications: a linear Gaussian sampling model of wastewater virus concentrations in England, using population density as covariate; and log-linear Poisson model of cardiovascular hospitalisations in England using socio-demographic variables at fine-scale administrative units as covariates.

stat.ME

A two-stage approach to heat-mortality risk assessment comparing multiple exposure-to-temperature models: the case study in Lazio, Italy

This study investigates how different spatiotemporal temperature models affect the estimation of heat-related mortality in Lazio, Italy (2008--2022). First, we compare three methods to reconstruct daily maximum temperature at the municipality level: 1. a Bayesian quantile regression model with spatial interpolation, 2. a Bayesian Gaussian regression model, 3. the gridded reanalysis data from ERA5-Land. Both Bayesian models are station-based and exhibit higher and more spatially variable temperatures compared to ERA5-Land. Then, using individual mortality data for cardiovascular and respiratory causes, we estimate temperature-mortality associations through Bayesian conditional Poisson models in a case-crossover design. Exposure is defined as the mean maximum temperature over the previous three days. Additional models include heatwave definitions combining different thresholds and durations. All models exhibit a marked increase in relative risk at high temperatures; however, the temperature of minimum risk varies significantly across methods. Stratified analyses reveal higher relative risk increases in females and the elderly (80+). Heatwave effects depend on the definitions used, but all methods capture an increased mortality risk associated with prolonged heat exposure. Results confirm the importance of temperature model choice in epidemiology and provide insights for early warning systems and climate-health adaptation strategies.

stat.AP

Modelling the Spatially Varying Non-Linear Effects of Heat Exposure

Exposure to high ambient temperatures is a significant driver of preventable mortality, with non-linear health effects and elevated risks in specific regions. To capture this complexity and account for spatial dependencies across small areas, we propose a Bayesian framework that integrates non-linear functions with the Besag, York, and Mollie (BYM2) model. Applying this framework to all-cause mortality data in Switzerland, we quantified spatial inequalities in heat-related mortality. We retrieved daily all-cause mortality at small areas (2,145 municipalities) for people older than 65 years from the Swiss Federal Office of Public Health and daily mean temperature at 1km$\times$1km grid from the Swiss Federal Office of Meteorology. By fully propagating uncertainties, we derived key epidemiological metrics, including heat-related excess mortality and minimum mortality temperature (MMT). Heat-related excess mortality rates were higher in northern Switzerland, while lower MMTs were observed in mountainous regions. Further, we explored the role of the proportion of individuals older than 85 years, green space, average temperature, deprivation, urbanicity, air pollution, and language regions in explaining these discrepancies. We found that spatial disparities in heat-related excess mortality were primarily driven by population age distribution, green space, and vulnerabilities associated with elevated temperature exposure.

stat.AP

A Bayesian Multisource Fusion Model for Spatiotemporal PM2.5 in an Urban Setting

Airborne particulate matter (PM2.5) is a major public health concern in urban environments, where population density and emission sources exacerbate exposure risks. We present a novel Bayesian spatiotemporal fusion model to estimate monthly PM2.5 concentrations over Greater London (2014-2019) at 1km resolution. The model integrates multiple PM2.5 data sources, including outputs from two atmospheric air quality dispersion models and predictive variables, such as vegetation and satellite aerosol optical depth, while explicitly modelling a latent spatiotemporal field. Spatial misalignment of the data is addressed through an upscaling approach to predict across the entire area. Building on stochastic partial differential equations (SPDE) within the integrated nested Laplace approximations (INLA) framework, our method introduces spatially- and temporally-varying coefficients to flexibly calibrate datasets and capture fine-scale variability. Model performance and complexity are balanced using predictive metrics such as the predictive model choice criterion and thorough cross-validation. The best performing model shows excellent fit and solid predictive performance, enabling reliable high-resolution spatiotemporal mapping of PM2.5 concentrations with the associated uncertainty. Furthermore, the model outputs, including full posterior predictive distributions, can be used to map exceedance probabilities of regulatory thresholds, supporting air quality management and targeted interventions in vulnerable urban areas, as well as providing refined exposure estimates of PM2.5 for epidemiological applications.

stat.ME

Bayesian Interrupted Time Series for evaluating policy change on mental well-being: an application to England's welfare reform

Factors contributing to social inequalities are also associated with negative mental health outcomes leading to disparities in mental well-being. We propose a Bayesian hierarchical model which can evaluate the impact of policies on population well-being, accounting for spatial/temporal dependencies. Building on an interrupted time series framework, our approach can evaluate how different profiles of individuals are affected in different ways, whilst accounting for their uncertainty. We apply the framework to assess the impact of the United Kingdoms welfare reform, which took place throughout the 2010s, on mental well-being using data from the UK Household Longitudinal Study. The additional depth of knowledge is essential for effective evaluation of current policy and implementation of future policy.

stat.ME

A framework for estimating and visualising excess mortality during the COVID-19 pandemic

COVID-19 related deaths underestimate the pandemic burden on mortality because they suffer from completeness and accuracy issues. Excess mortality is a popular alternative, as it compares observed with expected deaths based on the assumption that the pandemic did not occur. Expected deaths had the pandemic not occurred depend on population trends, temperature, and spatio-temporal patterns. In addition to this, high geographical resolution is required to examine within country trends and the effectiveness of the different public health policies. In this tutorial, we propose a framework using R to estimate and visualise excess mortality at high geographical resolution. We show a case study estimating excess deaths during 2020 in Italy. The proposed framework is fast to implement and allows combining different models and presenting the results in any age, sex, spatial and temporal aggregation desired. This makes it particularly powerful and appealing for online monitoring of the pandemic burden and timely policy making.

stat.AP

A Bayesian hierarchical small-area population model accounting for data source specific methodologies from American Community Survey, Population Estimates Program, and Decennial Census data

Small area estimates of population are necessary for many epidemiological studies, yet their quality and accuracy are often not assessed. In the United States, small area estimates of population counts are published by the United States Census Bureau (USCB) in the form of the Decennial census counts, Intercensal population projections (PEP), and American Community Survey (ACS) estimates. Although there are significant relationships between these data sources, there are important contrasts in data collection and processing methodologies, such that each set of estimates may be subject to different sources and magnitudes of error. Additionally, these data sources do not report identical small area population counts due to post-survey adjustments specific to each data source. Resulting small area disease/mortality rates may differ depending on which data source is used for population counts (denominator data). To accurately capture annual small area population counts, and associated uncertainties, we present a Bayesian population model (B-Pop), which fuses information from all three USCB sources, accounting for data source specific methodologies and associated errors. The main features of our framework are: 1) a single model integrating multiple data sources, 2) accounting for data source specific data generating mechanisms, and specifically accounting for data source specific errors, and 3) prediction of estimates for years without USCB reported data. We focus our study on the 159 counties of Georgia, and produce estimates for years 2005-2021.

stat.ME

Interoperability of statistical models in pandemic preparedness: principles and reality

We present "interoperability" as a guiding framework for statistical modelling to assist policy makers asking multiple questions using diverse datasets in the face of an evolving pandemic response. Interoperability provides an important set of principles for future pandemic preparedness, through the joint design and deployment of adaptable systems of statistical models for disease surveillance using probabilistic reasoning. We illustrate this through case studies for inferring spatial-temporal coronavirus disease 2019 (COVID-19) prevalence and reproduction numbers in England.

stat.ME

School neighbourhood and compliance with WHO-recommended annual NO2 guideline: a case study of Greater London

Despite several national and local policies towards cleaner air in England, many schools in London breach the WHO-recommended concentrations of air pollutants such as NO2 and PM2.5. This is while, previous studies highlight significant adverse health effects of air pollutants on children's health. In this paper we adopted a Bayesian spatial hierarchical model to investigate factors that affect the odds of schools exceeding the WHO-recommended concentration of NO2 (i.e., 40 ug/m3 annual mean) in Greater London (UK). We considered a host of variables including schools' characteristics as well as their neighbourhoods' attributes from household, socioeconomic, transport-related, land use, built and natural environment characteristics perspectives. The results indicated that transport-related factors including the number of traffic lights and bus stops in the immediate vicinity of schools, and borough-level bus fuel consumption are determinant factors that increase the likelihood of non-compliance with the WHO guideline. In contrast, distance from roads, river transport, and underground stations, vehicle speed (an indicator of traffic congestion), the proportion of borough-level green space, and the area of green space at schools reduce the likelihood of exceeding the WHO recommended concentration of NO2. As a sensitivity analysis, we repeated our analysis under a hypothetical scenario in which the recommended concentration of NO2 is 35 ug/m3, instead of 40 ug/m3. Our results underscore the importance of adopting clean fuel technologies on buses, installing green barriers, and reducing motorised traffic around schools in reducing exposure to NO2 concentrations in proximity to schools. This study would be useful for local authority decision making with the aim of improving air quality for school-aged children in urban settings.

stat.AP

A joint bayesian space-time model to integrate spatially misaligned air pollution data in R-INLA

In air pollution studies, dispersion models provide estimates of concentration at grid level covering the entire spatial domain, and are then calibrated against measurements from monitoring stations. However, these different data sources are misaligned in space and time. If misalignment is not considered, it can bias the predictions. We aim at demonstrating how the combination of multiple data sources, such as dispersion model outputs, ground observations and covariates, leads to more accurate predictions of air pollution at grid level. We consider nitrogen dioxide (NO2) concentration in Greater London and surroundings for the years 2007-2011, and combine two different dispersion models. Different sets of spatial and temporal effects are included in order to obtain the best predictive capability. Our proposed model is framed in between calibration and Bayesian melding techniques for data fusion red. Unlike other examples, we jointly model the response (concentration level at monitoring stations) and the dispersion model outputs on different scales, accounting for the different sources of uncertainty. Our spatio-temporal model allows us to reconstruct the latent fields of each model component, and to predict daily pollution concentrations. We compare the predictive capability of our proposed model with other established methods to account for misalignment (e.g. bilinear interpolation), showing that in our case study the joint model is a better alternative.

stat.AP

A spatio-temporal model to understand forest fires causality in Europe

Forest fires are the outcome of a complex interaction between environmental factors, topography and socioeconomic factors (Bedia et al, 2014). Therefore, understand causality and early prediction are crucial elements for controlling such phenomenon and saving lives.The aim of this study is to build spatio-temporal model to understand causality of forest fires in Europe, at NUTS2 level between 2012 and 2016, using environmental and socioeconomic variables.We have considered a disease mapping approach, commonly used in small area studies to assess thespatial pattern and to identify areas characterised by unusually high or low relative risk.

stat.AP

Missing data analysis and imputation via latent Gaussian Markov random fields

In this paper we recast the problem of missing values in the covariates of a regression model as a latent Gaussian Markov random field (GMRF) model in a fully Bayesian framework. Our proposed approach is based on the definition of the covariate imputation sub-model as a latent effect with a GMRF structure. We show how this formulation works for continuous covariates and provide some insight on how this could be extended to categorical covariates. The resulting Bayesian hierarchical model naturally fits within the integrated nested Laplace approximation (INLA) framework, which we use for model fitting. Hence, our work fills an important gap in the INLA methodology as it allows to treat models with missing values in the covariates. As in any other fully Bayesian framework, by relying on INLA for model fitting it is possible to formulate a joint model for the data, the imputed covariates and their missingness mechanism. In this way, we are able to tackle the more general problem of assessing the missingness mechanism by conducting a sensitivity analysis on the different alternatives to model the non-observed covariates. Finally, we illustrate the proposed approach with two examples on modeling health risk factors and disease mapping. Here, we rely on two different imputation mechanisms based on a typical multiple linear regression and a spatial model, respectively. Given the speed of model fitting with INLA we are able to fit joint models in a short time, and to easily conduct sensitivity analyses.

stat.CO

A hierarchical modelling approach to assess multi pollutant effects in time-series studies

When assessing the short term effect of air pollution on health outcomes, it is common practice to consider one pollutant at a time, due to their high correlation. Multi pollutant methods have been recently proposed, mainly consisting of collapsing the different pollutants into air quality indexes or clustering the pollutants and then evaluating the effect of each cluster on the health outcome. A major drawback of such approaches is that it is not possible to evaluate the health impact of each pollutant. In this paper we propose the use of the Bayesian hierarchical framework to deal with multi pollutant concentration in a two-component model: a pollutant model is specified to estimate the `true' concentration values for if your each pollutant and then such concentration is linked to the health outcomes in a time series perspective. Through a simulation study we evaluate the model performance and we apply the modelling framework to investigate the effect of six pollutants on cardiovascular mortality in Greater London in 2011-2012.

stat.AP

Using Ecological Propensity Score to Adjust for Missing Confounders in Small Area Studies

Small area ecological studies are commonly used in epidemiology to assess the impact of area level risk factors on health outcomes when data are only available in an aggregated form. However the resulting estimates are often biased due to unmeasured confounders, which typically are not available from the standard administrative registries used for these studies. Extra information on confounders can be provided through external datasets such as surveys or cohorts, where the data are available at the individual level rather than at the area level; however such data typically lack the geographical coverage of administrative registries. We develop a framework of analysis which combines ecological and individual level data from different sources to provide an adjusted estimate of area level risk factors which is less biased. Our method (i) summarises all available individual level confounders into an area level scalar variable, which we call ecological propensity score (EPS), (ii) implements a hierarchical structured approach to predict the values of EPS whenever they are missing, (iii) includes the estimated and predicted EPS into the ecological regression linking the risk factors to the health outcome. Through a simulation study we show that integrating individual level data into small area analyses via EPS is a promising method to reduce the bias intrinsic in ecological studies due to unmeasured confounders; we also apply the method to a real case study to evaluate the effect of air pollution on coronary heart disease hospital admissions in Greater London.

stat.AP

A Bayesian Approach to Modelling Fine-Scale Spatial Dynamics of Non-State Terrorism: World Study, 2002-2013

To this day, terrorism persists as a worldwide threat, as exemplified by the ongoing lethal attacks perpetrated by ISIS in Iraq, Syria, Al Qaeda in Yemen, and Boko Haram in Nigeria. In response, states deploy various counterterrorism policies, the costs of which could be reduced through efficient preventive measures. Statistical models able to account for complex spatio-temporal dependencies have not yet been applied, despite their potential for providing guidance to explain and prevent terrorism. In an effort to address this shortcoming, we employ hierarchical models in a Bayesian context, where the spatial random field is represented by a stochastic partial differential equation. Our results confirm the contagious nature of the lethality of terrorism and the number of lethal terrorist attacks in both space and time. Moreover, the frequency of lethal attacks tends to be higher in richer areas, close to large cities, and within democratic countries. In contrast, attacks are more likely to be lethal far away from large cities, at higher altitudes, in poorer areas, and in locations with higher ethnic diversity. We argue that, on a local scale, the lethality of terrorism and the frequency of lethal attacks are driven by antagonistic mechanisms.

stat.AP