SearcharxivSearch

arXiv subjects

Sara Martino

Publications and source records attributed to Sara Martino.

At least 19 recordsLinked to original sources

Cross Validation for the log Gaussian Cox Process

The log Gaussian Cox Process (LGCP) is one of the most widely used models for the analysis of spatial point patterns. Although Bayesian methods and software for fitting LGCPs are now well established, practical tools for model criticism, predictive assessment, and model comparison remain comparatively underdeveloped. This paper develops a practical Bayesian cross-validation framework for LGCPs while placing it within a broader framework for Bayesian model assessment. Our approach is motivated by the conceptual and computational challenges that arise when defining predictive validation for point processes, including the choice of holdout data, prediction task, scoring rule, and computational strategy. We propose a leave-region-out cross-validation framework in which the fundamental unit of prediction is a bounded spatial region rather than an individual event. Predictive performance is assessed through the logarithmic scoring rule for joint probabilistic forecasts, while computational efficiency is achieved by combining a new likelihood approximation with a Bayesian unlearning strategy and Laplace approximations to obtain leave-region-out predictive distributions without repeated model refitting. We validate the proposed approximations against Monte Carlo estimates obtained via brute-force refitting, demonstrate the methodology on simulated and real spatial point pattern datasets defined on Euclidean domains and networks, and provide implementations in the R-INLA, inlabru, and MetricGraph packages.

stat.ME

Flexible covariance structures on metric graphs

Whittle-Matérn (WM) Gaussian random fields (GRFs) are defined as solutions of stochastic partial differential equations (SPDEs) and provide a natural analog of Matérn GRFs on non-Euclidean geometry where the Matérn covariance function is not valid. In particular, WM GRFs on metric graphs have been an active area of research motivated by road and river networks where spatial dependence is more naturally described by intrinsic distances in the network than by Euclidean distances. This family of GRFs is controlled by three parameters relating to marginal variance, spatial range, and smoothness, but can be extended to so-called generalized WM GRFs through spatially varying coefficients in the SPDE. Recent work has considered the use of spatially varying covariates, but the full possibilities of flexibility have not been considered. In this work, we introduce latent GRFs that describe the spatially varying coefficients of the SPDE. This flexible model is compared to less flexible models in a simulation study evaluating both the ability to estimate the covariance structure and predictive ability. An important focus is the number of observations and replications necessary to reliably recover the covariance structure. We find that the flexible model improves over less flexible models in the presence of sufficient data. We also demonstrate practical applicability on traffic counts in a part of Madrid, and observe major differences between in-sample and out-of-sample predictive abilities of the models compared.

stat.ME

Joint Modelling of Line and Point Data on Metric Graphs

Metric graphs are useful tools for describing spatial domains like road and river networks, where spatial dependence act along the network. We take advantage of recent developments for such Gaussian Random Fields (GRFs), and consider joint spatial modelling of observations with different spatial supports. Motivated by an application to traffic state modelling in Trondheim, Norway, we consider line-referenced data, which can be described by an integral of the GRF along a line segment on the metric graph, and point-referenced data. Through a simulation study inspired by the application, we investigate the number of replicates that are needed to estimate parameters and to predict unobserved locations. The former is assessed using bias and variability, and the latter is assessed through root mean square error (RMSE), continuous rank probability scores (CRPSs), and coverage. Joint modelling is contrasted with a simplified approach that treat line-referenced observations as point-referenced observations. The results suggest joint modelling leads to strong improvements. The application to Trondheim, Norway, combines point-referenced induction loop data and line-referenced public transportation data. To ensure positive speeds, we use a non-linear link function, which requires integrals of non-linear combinations of the linear predictor. This is made computationally feasible by a combination of the R packages inlabru and MetricGraph, and new code for processing geographical line data to work with existing graph representations and fmesher methods for dealing with line support in inlabru on objects from MetricGraph. We fit the model to two datasets where we expect different spatial dependency and compare the results.

stat.ME

Geographical inequalities in mortality by age and gender in Italy, 2002-2019: insights from a spatial extension of the Lee-Carter model

Italy reports some of the lowest levels of mortality in the developed world. Recent evidence, however, suggests that even in low mortality countries improvements may be slowing and regional inequalities widening. This study contributes new empirical evidence to the debate by analysing mortality data by single year of age for males and females across 107 provinces in Italy from 2002 to 2019. We extend the widely used Lee Carter model to include spatially varying age specific effects, and further specify it to capture space age time interactions. The model is estimated in a Bayesian framework using the inlabru package, which builds on INLA (Integrated Nested Laplace Approximation) for non linear models and facilitates the use of smoothing priors. This approach borrows strength across provinces and years, mitigating random fluctuations in small area death counts. Results demonstrate the value of such a granular approach, highlighting the existence of an uneven geography of mortality despite overall national improvements. Mortality disadvantage is concentrated in parts of the Centre South and North West, while the Centre North and North East fare relatively better. These geographical differences have widened since 2010, with clear age and gender specific patterns, being more pronounced at younger adult ages for men and at older adult ages for women. Future work may involve refining the analysis to mortality by cause of death or socioeconomic status, informing more targeted public health policies to address mortality disparities across Italy's provinces.

stat.AP

Linking climate and dengue in the Philippines using a two-stage Bayesian spatio-temporal model

Dengue is an infectious disease which poses significant socioeconomic and disease burden in many tropical and subtropical regions of the world. This work aims to provide additional insight into the association between dengue and climate in the Philippines. We employ a two-stage modelling framework: the first stage fits climate models, while the second stage fits a health model that uses the climate predictions from the first stage as inputs. We postulate a Bayesian spatio-temporal model and use the integrated nested Laplace approximation (INLA) approach for inference. To account for the uncertainty in the climate models, we perform posterior sampling and then perform Bayesian model averaging to compute the final posterior estimates of second-stage model parameters. The results indicate that temperature is positively associated with dengue, although extremely hot conditions tend to have a negative effect. Moreover, the relationship between rainfall and dengue varies in space. In areas with uniform amounts of rainfall all year round, rainfall is negatively associated with dengue. In contrast, in regions with pronounced dry and wet season, rainfall shows a positive association with dengue. Finally, there remains unexplained structured variation in space and time after accounting for the impact of climate variables and other covariates.

stat.AP

Validating uncertainty propagation approaches for two-stage Bayesian spatial models using simulation-based calibration

This work tackles the problem of uncertainty propagation in two-stage Bayesian models, with a focus on spatial applications. A two-stage modeling framework has the advantage of being more computationally efficient than a fully Bayesian approach when the first-stage model is already complex in itself, and avoids the potential problem of unwanted feedback effects. Two ways of doing two-stage modeling are the crude plug-in method and the posterior sampling method. The former ignores the uncertainty in the first-stage model, while the latter can be computationally expensive. This paper validates the two aforementioned approaches and proposes a new approach to do uncertainty propagation, which we call the $\mathbf{Q}$ uncertainty method, implemented using the Integrated Nested Laplace Approximation (INLA). We validate the different approaches using the simulation-based calibration method, which tests the self-consistency property of Bayesian models. Results show that the crude plug-in method underestimates the true posterior uncertainty in the second-stage model parameters, while the resampling approach and the proposed method are correct. We illustrate the approaches in a real life data application which aims to link relative humidity and Dengue cases in the Philippines for August 2018.

stat.ME

A Data Fusion Model for Meteorological Data using the INLA-SPDE method

This work aims to combine two primary meteorological data sources in the Philippines: data from a sparse network of weather stations and outcomes of a numerical weather prediction model. To this end, we propose a data fusion model which is primarily motivated by the problem of sparsity in the observational data and the use of a numerical prediction model as an additional data source in order to obtain better predictions for the variables of interest. The proposed data fusion model assumes that the different data sources are error-prone realizations of a common latent process. The outcomes from the weather stations follow the classical error model while the outcomes of the numerical weather prediction model involves a constant multiplicative bias parameter and an additive bias which is spatially-structured and time-varying. We use a Bayesian model averaging approach with the integrated nested Laplace approximation (INLA) for doing inference. The proposed data fusion model outperforms the stations-only model and the regression calibration approach, when assessed using leave-group-out cross-validation (LGOCV). We assess the benefits of data fusion and evaluate the accuracy of predictions and parameter estimation through a simulation study. The results show that the proposed data fusion model generally gives better predictions compared to the stations-only approach especially with sparse observational data.

stat.AP

Fast spatial simulation of extreme high-resolution radar precipitation data using INLA

Aiming to deliver improved precipitation simulations for hydrological impact assessment studies, we develop a methodology for modelling and simulating high-dimensional spatial precipitation extremes, focusing on both their marginal distributions and tail dependence structures. Tail dependence is crucial for assessing the consequences of extreme precipitation events, yet most stochastic weather generators do not attempt to capture this property. The spatial distribution of precipitation occurrences is modelled with four competing models, while the spatial distribution of nonzero extreme precipitation intensities are modelled with a latent Gaussian version of the spatial conditional extremes model. Nonzero precipitation marginal distributions are modelled using latent Gaussian models with gamma and generalised Pareto likelihoods. Fast inference is achieved using integrated nested Laplace approximations (INLA). We model and simulate spatial precipitation extremes in Central Norway, using 13 years of hourly radar data with a spatial resolution of $1 \times 1$~km$^2$, over an area of size $6461$~km$^2$, to describe the behaviour of extreme precipitation over a small drainage area. Inference on this high-dimensional data set is achieved within hours, and the simulations capture the main trends of the observed precipitation well.

stat.AP

Spatio-temporal Occupancy Models with INLA

Modern methods for quantifying and predicting species distribution play a crucial part in biodiversity conservation. Occupancy models are a popular choice for analyzing species occurrence data as they allow to separate the observational error induced by imperfect detection, and the sources of bias affecting the occupancy process. However, the spatial and temporal variation in occupancy not accounted for by environmental covariates is often ignored or modelled through simple spatial structures as the computational costs of fitting explicit spatio-temporal models is too high. In this work, we demonstrate how INLA may be used to fit complex occupancy models and how the R-INLA package can provide a user-friendly interface to make such complex models available to users. We show how occupancy models, provided some simplification on the detection process, can be framed as latent Gaussian models and benefit from the powerful INLA machinery. A large selection of complex modelling features, and random effect modelshave already been implemented in R-INLA. These become available for occupancy models, providing the user with an efficient and flexible toolbox. We illustrate how INLA provides a computationally efficient framework for developing and fitting complex occupancy models using two case studies. Through these, we show how different spatio-temporal models that include spatial-varying trends, smooth terms, and spatio-temporal random effects can be fitted. At the cost of limiting the complexity of the detection model, INLA can incorporate a range of complex structures in the process. INLA-based occupancy models provide an alternative framework to fit complex spatiotemporal occupancy models. The need for new and more flexible computationally approaches to fit such models makes INLA an attractive option for addressing complex ecological problems, and a promising area of research.

stat.ME

A joint Bayesian framework for missing data and measurement error using integrated nested Laplace approximations

Measurement error (ME) and missing values in covariates are often unavoidable in disciplines that deal with data, and both problems have separately received considerable attention during the past decades. However, while most researchers are familiar with methods for treating missing data, accounting for ME in covariates of regression models is less common. In addition, ME and missing data are typically treated as two separate problems, despite practical and theoretical similarities. Here, we exploit the fact that missing data in a continuous covariate is an extreme case of classical ME, allowing us to use existing methodology that accounts for ME via a Bayesian framework that employs integrated nested Laplace approximations (INLA), and thus to simultaneously account for both ME and missing data in the same covariate. As a useful by-product, we present an approach to handle missing data in INLA, since this corresponds to the special case when no ME is present. In addition, we show how to account for Berkson ME in the same framework. In its broadest generality, the proposed joint Bayesian framework can thus account for Berkson ME, classical ME, and missing data, or for any combination of these in the same or different continuous covariates of the family of regression models that are feasible with INLA. The approach is exemplified using both simulated and real data. We provide extensive and fully reproducible Supplementary Material with thoroughly documented examples using {R-INLA} and {inlabru}.

stat.ME

An Efficient Workflow for Modelling High-Dimensional Spatial Extremes

A successful model for high-dimensional spatial extremes should, in principle, be able to describe both weakening extremal dependence at increasing levels and changes in the type of extremal dependence class as a function of the distance between locations. Furthermore, the model should allow for computationally tractable inference using inference methods that efficiently extract information from data and that are robust to model misspecification. In this paper, we demonstrate how to fulfil all these requirements by developing a comprehensive methodological workflow for efficient Bayesian modelling of high-dimensional spatial extremes using the spatial conditional extremes model while performing fast inference with R-INLA. We then propose a post hoc adjustment method that results in more robust inference by properly accounting for possible model misspecification. The developed methodology is applied for modelling extreme hourly precipitation from high-resolution radar data in Norway. Inference is computationally efficient, and the resulting model fit successfully captures the main trends in the extremal dependence structure of the data. Robustifying the model fit by adjusting for possible misspecification further improves model performance.

stat.ME

Modelling sub-daily precipitation extremes with the blended generalised extreme value distribution

A new method is proposed for modelling the yearly maxima of sub-daily precipitation, with the aim of producing spatial maps of return level estimates. Yearly precipitation maxima are modelled using a Bayesian hierarchical model with a latent Gaussian field, with the blended generalised extreme value (bGEV) distribution used as a substitute for the more standard generalised extreme value (GEV) distribution. Inference is made less wasteful with a novel two-step procedure that performs separate modelling of the scale parameter of the bGEV distribution using peaks over threshold data. Fast inference is performed using integrated nested Laplace approximations (INLA) together with the stochastic partial differential equation (SPDE) approach, both implemented in R-INLA. Heuristics for improving the numerical stability of R-INLA with the GEV and bGEV distributions are also presented. The model is fitted to yearly maxima of sub-daily precipitation from the south of Norway, and is able to quickly produce high-resolution return level maps with uncertainty. The proposed two-step procedure provides an improved model fit over standard inference techniques when modelling the yearly maxima of sub-daily precipitation with the bGEV distribution.

stat.AP

A spatio-temporal analysis of NO$_2$ concentrations during the Italian 2020 COVID-19 lockdown

When a new environmental policy or a specific intervention is taken in order to improve air quality, it is paramount to assess and quantify - in space and time - the effectiveness of the adopted strategy. The lockdown measures taken worldwide in 2020 to reduce the spread of the SARS-CoV- 2 virus can be envisioned as a policy intervention with an indirect effect on air quality. In this paper we propose a statistical spatio-temporal model as a tool for intervention analysis, able to take into account the effect of weather and other confounding factors, as well as the spatial and temporal correlation existing in the data. In particular, we focus here on the 2019/2020 relative change in nitrogen dioxide (NO$_2$) concentrations in the north of Italy, for the period of March and April during which the lockdown measure was in force. As an output, we provide a collection of weekly continuous maps, describing the spatial pattern of the NO$_2$ 2019/2020 relative changes. We found that during March and April 2020 most of the studied area is characterized by negative relative changes (median values around -25%), with the exception of the first week of March and the fourth week of April (median values around 5%). As these changes cannot be attributed to a weather effect, it is likely that they are a byproduct of the lockdown measures.

stat.AP

Integration of presence-only data from several sources. A case study on dolphins' spatial distribution

Presence-only data are a typical occurrence in species distribution modeling. They include the presence locations and no information on the absence. Their modeling usually does not account for detection biases. In this work, we aim to merge three different sources of information to model the presence of marine mammals. The approach is fully general and it is applied to two species of dolphins in the Central Tyrrhenian Sea (Italy) as a case study. Data come from the Italian Environmental Protection Agency (ISPRA) and Sapienza University of Rome research campaigns, and from a careful selection of social media (SM) images and videos. We build a Log Gaussian Cox process where different detection functions describe each data source. For the SM data, we analyze several choices that allow accounting for detection biases. Our findings allow for a correct understanding of Stenella coeruleoalba and Tursiops truncatus distribution in the study area. The results prove that the proposed approach is broadly applicable, it can be widely used, and it is easily implemented in the R software using INLA and inlabru. We provide examples' code with simulated data in the supplementary materials.

q-bio.QM

Importance Sampling with the Integrated Nested Laplace Approximation

The Integrated Nested Laplace Approximation (INLA) is a deterministic approach to Bayesian inference on latent Gaussian models (LGMs) and focuses on fast and accurate approximation of posterior marginals for the parameters in the models. Recently, methods have been developed to extend this class of models to those that can be expressed as conditional LGMs by fixing some of the parameters in the models to descriptive values. These methods differ in the manner descriptive values are chosen. This paper proposes to combine importance sampling with INLA (IS-INLA), and extends this approach with the more robust adaptive multiple importance sampling algorithm combined with INLA (AMIS-INLA). This paper gives a comparison between these approaches and existing methods on a series of applications with simulated and observed datasets and evaluates their performance based on accuracy, efficiency, and robustness. The approaches are validated by exact posteriors in a simple bivariate linear model; then, they are applied to a Bayesian lasso model, a Bayesian imputation of missing covariate values, and lastly, in parametric Bayesian quantile regression. The applications show that the AMIS-INLA approach, in general, outperforms the other methods, but the IS-INLA algorithm could be considered for faster inference when good proposals are available.

stat.CO

Spatio-temporal modelling of $\text{PM}_{10}$ daily concentrations in Italy using the SPDE approach

This paper illustrates the main results of a spatio-temporal interpolation process of $\text{PM}_{10}$ concentrations at daily resolution using a set of 410 monitoring sites, distributed throughout the Italian territory, for the year 2015. The interpolation process is based on a Bayesian hierarchical model where the spatial-component is represented through the Stochastic Partial Differential Equation (SPDE) approach with a lag-1 temporal autoregressive component (AR1). Inference is performed through the Integrated Nested Laplace Approximation (INLA). Our model includes 11 spatial and spatio-temporal predictors, including meteorological variables and Aerosol Optical Depth. As the predictors' impact varies across months, the regression is based on 12 monthly models with the same set of covariates. The predictive model performance has been analyzed using a cross-validation study. Our results show that the predicted and the observed values are well in accordance (correlation range: 0.79 - 0.91; bias: 0.22 - 1.07 $μ\text{g}/\text{m}^3$; RMSE: 4.9 - 13.9 $μ\text{g}/\text{m}^3$). The model final output is a set of 365 gridded (1km $\times$ 1km) daily $\text{PM}_{10}$ maps over Italy equipped with an uncertainty measure. The spatial prediction performance shows that the interpolation procedure is able to reproduce the large scale data features without unrealistic artifacts in the generated $\text{PM}_{10}$ surfaces. The paper presents also two illustrative examples of practical applications of our model, exceedance probability and population exposure maps.

stat.AP

Integrated Nested Laplace Approximations (INLA)

This is a short description and basic introduction to the Integrated nested Laplace approximations (INLA) approach. INLA is a deterministic paradigm for Bayesian inference in latent Gaussian models (LGMs) introduced in Rue et al. (2009). INLA relies on a combination of analytical approximations and efficient numerical integration schemes to achieve highly accurate deterministic approximations to posterior quantities of interest. The main benefit of using INLA instead of Markov chain Monte Carlo (MCMC) techniques for LGMs is computational; INLA is fast even for large, complex models. Moreover, being a deterministic algorithm, INLA does not suffer from slow convergence and poor mixing. INLA is implemented in the R package R-INLA, which represents a user-friendly and versatile tool for doing Bayesian inference. R-INLA returns posterior marginals for all model parameters and the corresponding posterior summary information. Model choice criteria as well as predictive diagnostics are directly available. Here, we outline the theory behind INLA, present the R-INLA package and describe new developments of combining INLA with MCMC for models that are not possible to fit with R-INLA.

stat.CO

Spatial Modelling of Temperature and Humidity using Systems of Stochastic Partial Differential Equations

This work is motivated by constructing a weather simulator for precipitation. Temperature and humidity are two of the most important driving forces of precipitation, and the strategy is to have a stochastic model for temperature and humidity, and use a deterministic model to go from these variables to precipitation. Temperature and humidity are empirically positively correlated. Generally speaking, if variables are empirically dependent, then multivariate models should be considered. In this work we model humidity and temperature in southern Norway. We want to construct bivariate Gaussian random fields (GRFs) based on this dataset. The aim of our work is to use the bivariate GRFs to capture both the dependence structure between humidity and temperature as well as their spatial dependencies. One important feature for the dataset is that the humidity and temperature are not necessarily observed at the same locations. Both univariate and bivariate spatial models are fitted and compared. For modeling and inference the SPDE approach for univariate models and the systems of SPDEs approach for multivariate models have been used. To evaluate the performance of the difference between the univariate and bivariate models, we compare predictive performance using some commonly used scoring rules: mean absolute error, mean-square error and continuous ranked probability score. The results illustrate that we can capture strong positive correlation between the temperature and the humidity. Furthermore, the results also agree with the physical or empirical knowledge. At the end, we conclude that using the bivariate GRFs to model this dataset is superior to the approach with independent univariate GRFs both when evaluating point predictions and for quantifying prediction uncertainty.

stat.AP