SearcharxivSearch

arXiv subjects

Emanuele Giorgi

Publications and source records attributed to Emanuele Giorgi.

At least 19 recordsLinked to original sources

Comment on: "The Two Cultures of Prevalence Mapping: Small Area Estimation and Model-Based Geostatistics"

Small Area Estimation (SAE) and Model-Based Geostatistics (MBG) provide complementary approaches to prevalence mapping, with their relative advantages depending on the inferential goals and characteristics of the available data. We argue that a fuller comparison should consider model interpretability, the role of epidemiologically motivated covariates, inferential objectives beyond area-level prediction, integration of data from different spatial partitions and surveys, and task-specific model validation. In particular, we question whether survey design variables should routinely be included in MBG models when their effects may instead be mediated by measurable environmental and socio-economic risk factors. We further argue that simulation-based validation, tailored to the operational objectives of prevalence mapping, can provide a more informative assessment of model performance than conventional cross-validation alone.

stat.ME

Bivariate geostatistical latent variable models for the analysis of antibody density data

The increasing availability of serosurveys that measure antibody responses to multiple antigens requires the development of methods that can exploit the full information content of such data, both biological and spatial. However, the non-Gaussian and potentially multimodal distributional behaviour of antibody responses makes the development of such methods inherently complex, especially in a multivariate setting. Here, we extend the latent variable framework of Giorgi and Wallin (2026), in which continuous antibody concentrations are modelled through an individual-level latent seroreactivity process that represents the level of immune activation to a given antigen. We focus primarily on the bivariate setting and set out a series of guiding principles that justify the resulting joint modelling structure. The proposed model captures distinct sources of correlation between antibody responses, arising both from shared exposure to the same environment and from biological processes occurring within the same host. Spatial dependence is introduced through a novel bivariate Matérn random field, which we use to construct a parsimonious class of cross-covariance functions between antigen-specific spatial processes. We illustrate the application of the framework to analyse data on bivariate antibody measurements from a malaria serosurvey in the Kenyan highlands. Results from the application and a simulation study show that ignoring this correlation substantially degrades inference on joint properties of the antibody distributions and on individual-level seroreactivity, but matters less when interest lies exclusively in each antibody's marginal distribution. Finally, we discuss how the framework could be extended to settings with more than two antigens, and highlight the modelling challenges that arise as the number of antigens grows.

stat.ME

A decay-adjusted spatio-temporal model to account for the impact of mass drug administration on neglected tropical disease prevalence

Prevalence surveys are routinely used to monitor the effectiveness of mass drug administration (MDA) programmes for controlling neglected tropical diseases (NTDs). We propose a decay-adjusted spatio-temporal (DAST) model that explicitly accounts for the time-varying impact of MDA on NTD prevalence, providing a flexible and interpretable framework for estimating intervention effects from sparse survey data. Using case studies on soil-transmitted helminths and lymphatic filariasis, we show that DAST offers a practical alternative to standard geostatistical models when the objective includes quantifying MDA impact and supporting short-term programmatic forecasting. We also discuss extensions and identifiability challenges, advocating for data-driven parsimony over complexity in settings where the available data are too sparse to support the estimation of highly parameterised models.

stat.AP

A Time-Series Model for Areal Data Using Area-Specific Gaussian Processes with Spatially Correlated Hyperparameters

In many applied settings, areal data are observed repeatedly over long time periods, as commonly occurs in infectious disease surveillance and environmental or demographic monitoring. Accurate characterization of local temporal dynamics and uncertainty is important for monitoring disease trends, identifying local changes, and supporting public health decision-making. Traditional spatio-temporal models generally represent spatial, temporal, and space-time interaction components through structured random effects acting on the latent outcome process. We propose a Bayesian spatio-temporal hierarchical framework in which temporal dynamics are modeled using area-specific Gaussian processes, while spatial dependence is introduced through spatially correlated Gaussian-process covariance hyperparameters. This allows neighboring regions to share information about the characteristics of their temporal dependence while retaining area-specific temporal trajectories, providing an alternative representation of spatio-temporal dependence. Inference is performed using Markov chain Monte Carlo methods. The approach is illustrated using monthly malaria incidence data from three Mozambican provinces and evaluated using the Root Mean Squared Error, Continuous Ranked Probability Score, empirical coverage probability, and credible interval width. Compared with established spatio-temporal models, the proposed framework achieves competitive predictive accuracy and better-calibrated predictive uncertainty. Spatially structuring the temporal hyperparameters also improves predictive interval calibration relative to an equivalent model with independent temporal processes, demonstrating the value of borrowing spatial information on temporal dependence and supporting this formulation as a competitive alternative to conventional outcome-level spatial smoothing for spatio-temporal areal data in disease surveillance.

stat.ME

Integrating Temporal Disaggregation and Distributed Lag Nonlinear Models for Bayesian Spatio-Temporal Disease Mapping with High-Resolution Environmental Exposures

Environmental conditions are major drivers of malaria transmission, but epidemiological analyses are often constrained by temporal misalignment between health outcomes reported at coarse time scales and environmental exposures available at finer resolutions. Conventional approaches aggregate environmental data to match health outcomes, potentially obscuring delayed and nonlinear relationships. We propose a Bayesian spatio-temporal framework that addresses this limitation through a latent daily disease process linked to observed monthly malaria counts by temporal disaggregation. The framework integrates distributed lag nonlinear models for climatic effects, spatio-temporal random effects, and intervention covariates within a unified hierarchical model. The methodology was applied to malaria surveillance data from 161 districts in Mozambique between 2017 and 2024, integrating temperature, precipitation, relative humidity, vegetation, elevation, and malaria interventions. Compared with a conventional monthly model, the proposed framework improved predictive accuracy and uncertainty quantification while exploiting the temporal resolution of environmental data. Estimated relationships showed nonlinear associations between climatic variability and malaria incidence, including an optimal temperature range, increasing risk with positive vegetation anomalies, and nonlinear precipitation effects. By avoiding temporal aggregation of environmental exposures, the framework provides a flexible approach for investigating delayed environmental effects from routine surveillance data and can be extended to other environmentally sensitive diseases with mismatched temporal resolutions.

stat.AP

Approximate Likelihood-Based Inference for Spatial Generalized Linear Mixed Models

We study maximum likelihood estimation for spatial generalized linear mixed models with Gaussian process approximations using a stochastic Newton-Raphson algorithm. We consider two Gaussian Process approximations in this context: spectral Gaussian process approximations and stochastic partial differential equations (SPDE). We refine the stochastic maximum likelihood algorithm and we propose a new stopping criterion for efficient termination to prevent long runs of sampling in the stationary post-convergence phase and a Monte Carlo estimator of fixed effect standard errors. We run a series of simulation comparisons of spatial statistical models alongside the popular Bayesian integrated nested Laplacian approximation method which incorporates SPDE. We show that HSGP provides nominal coverage of fixed and random effect parameters with smooth latent fields but performance degrades for rough fields. SPDE in a stochastic maximum likelihood framework maintains nominal coverage and matches or improves upon the performance of Bayesian integrated nested Laplacian approximation.

stat.ME

A flexible class of latent variable models for the analysis of antibody response data

Existing approaches to modelling antibody concentration data are mostly based on finite mixture models that rely on the assumption that individuals can be divided into two distinct groups: seronegative and seropositive. Here, we challenge this dichotomous modelling assumption and propose a latent variable modelling framework in which the immune status of each individual is represented along a continuum of latent seroreactivity, ranging from minimal to strong immune activation. This formulation provides greater flexibility in capturing age-related changes in antibody distributions while preserving the full information content of quantitative measurements. We show that the proposed class of models can accommodate a large variety of model formulations, both mechanistic and regression-based, and also includes finite mixture models as a special case. We also propose a computationally efficient $L_2$-based estimator as an alternative to maximum likelihood estimation, which substantially reduces computational cost, and we establish its consistency. Through a case study on malaria serology, we demonstrate how the flexibility of the novel framework enables joint analyses across all ages while accounting for changes in transmission patterns. We conclude by outlining extensions of the proposed modelling framework and its relevance to other omics applications.

stat.ME

Mapping food insecurity in the Brazilian Amazon using a spatial item factor analysis model

Food insecurity, a latent construct defined as the lack of consistent access to sufficient and nutritious food, is a pressing global issue with serious health and social justice implications. Item factor analysis is commonly used to study such latent constructs, but it typically assumes independence between sampling units. In the context of food insecurity, this assumption is often unrealistic, as food access is linked to socio-economic conditions and social relations that are spatially structured. To address this, we propose a spatial item factor analysis model that captures spatial dependence, allowing us to predict latent factors at unsampled locations and identify food insecurity hotspots. We develop a Bayesian sampling scheme for inference and illustrate the explanatory strength of our model by analysing household perceptions of food insecurity in Ipixuna, a remote river-dependent urban centre in the Brazilian Amazon. Our approach is implemented in the R package spifa, with further details provided in the Supplementary Material. This spatial extension offers policymakers and researchers a stronger tool for understanding and addressing food insecurity to locate and prioritise areas in greatest need. Our proposed methodology can be applied more widely to other spatially structured latent constructs.

stat.ME

Understanding the effects of dichotomization of continuous outcomes on geostatistical inference

Diagnosis is often based on the exceedance or not of continuous health indicators of a predefined cut-off value, so as to classify patients into positives and negatives for the disease under investigation. In this paper, we investigate the effects of dichotomization of spatially-referenced continuous outcome variables on geostatistical inference. Although this issue has been extensively studied in other fields, dichotomization is still a common practice in epidemiological studies. Furthermore, the effects of this practice in the context of prevalence mapping have not been fully understood. Here, we demonstrate how spatial correlation affects the loss of information due to dichotomization, how linear geostatistical models can be used to map disease prevalence and thus avoid dichotomization, and finally, how dichotomization affects our predictive inference on prevalence. To pursue these objectives, we develop a metric, based on the composite likelihood, which can be used to quantify the potential loss of information after dichotomization without requiring the fitting of Binomial geostatistical models. Through a simulation study and two applications on disease mapping in Africa, we show that, as thresholds used for dichotomization move further away from the mean of the underlying process, the performance of binomial geostatistical models deteriorates substantially. We also find that dichotomization can lead to the loss of fine scale features of disease prevalence and increased uncertainty in the parameter estimates, especially in the presence of a large noise to signal ratio. These findings strongly support the conclusions from previous studies that dichotomization should be always avoided whenever feasible.

stat.AP

A Spatially Discrete Approximation to Log-Gaussian Cox Processes for Modelling Aggregated Disease Count Data

In this paper, we develop a computationally efficient discrete approximation to log-Gaussian Cox process (LGCP) models for the analysis of spatially aggregated disease count data. Our approach overcomes an inherent limitation of spatial models based on Markov structures, namely that each such model is tied to a specific partition of the study area, and allows for spatially continuous prediction. We compare the predictive performance of our modelling approach with LGCP through a simulation study and an application to primary biliary cirrhosis incidence data in Newcastle-Upon-Tyne, UK. Our results suggest that when disease risk is assumed to be a spatially continuous process, the proposed approximation to LGCP provides reliable estimates of disease risk both on spatially continuous and aggregated scales. The proposed methodology is implemented in the open-source R package SDALGCP.

stat.ME

Spatial Analysis Made Easy with Linear Regression and Kernels

Kernel methods are an incredibly popular technique for extending linear models to non-linear problems via a mapping to an implicit, high-dimensional feature space. While kernel methods are computationally cheaper than an explicit feature mapping, they are still subject to cubic cost on the number of points. Given only a few thousand locations, this computational cost rapidly outstrips the currently available computational power. This paper aims to provide an overview of kernel methods from first-principals (with a focus on ridge regression), before progressing to a review of random Fourier features (RFF), a set of methods that enable the scaling of kernel methods to big datasets. At each stage, the associated R code is provided. We begin by illustrating how the dual representation of ridge regression relies solely on inner products and permits the use of kernels to map the data into high-dimensional spaces. We progress to RFFs, showing how only a few lines of code provides a significant computational speed-up for a negligible cost to accuracy. We provide an example of the implementation of RFFs on a simulated spatial data set to illustrate these properties. Lastly, we summarise the main issues with RFFs and highlight some of the advanced techniques aimed at alleviating them.

stat.ML

A Geostatistical Framework for Combining Spatially Referenced Disease Prevalence Data from Multiple Diagnostics

Multiple diagnostic tests are often used due to limited resources or because they provide complementary information on the epidemiology of a disease under investigation. Existing statistical methods to combine prevalence data from multiple diagnostics ignore the potential over-dispersion induced by the spatial correlations in the data. To address this issue, we develop a geostatistical framework that allows for joint modelling of data from multiple diagnostics by considering two main classes of inferential problems: (1) to predict prevalence for a gold-standard diagnostic using low-cost and potentially biased alternative tests; (2) to carry out joint prediction of prevalence from multiple tests. We apply the proposed framework to two case studies: mapping Loa loa prevalence in Central and West Africa, using miscroscopy and a questionnaire-based test called RAPLOA; mapping Plasmodium falciparum malaria prevalence in the highlands of Western Kenya using polymerase chain reaction and a rapid diagnostic test. We also develop a Monte Carlo procedure based on the variogram in order to identify parsimonious geostatistical models that are compatible with the data. Our study highlights (i) the importance of accounting for diagnostic-specific residual spatial variation and (ii) the benefits accrued from joint geostatistical modelling so as to deliver more reliable and precise inferences on disease prevalence.

stat.AP

Geostatistical methods for disease mapping and visualization using data from spatio-temporally referenced prevalence surveys

In this paper we set out general principles and develop geostatistical methods for the analysis of data from spatio-temporally referenced prevalence surveys. Our objective is to provide a tutorial guide that can be used in order to identify parsimonious geostatistical models for prevalence mapping. A general variogram-based Monte Carlo procedure is proposed to check the validity of the modelling assumptions. We describe and contrast likelihood-based and Bayesian methods of inference, showing how to account for parameter uncertainty under each of the two paradigms. We also describe extensions of the standard model for disease prevalence that can be used when stationarity of the spatio-temporal covariance function is not supported by the data. We discuss how to define predictive targets and argue that exceedance probabilities provide one of the most effective ways to convey uncertainty in prevalence estimates. We describe statistical software for the visualization of spatio-temporal predictive summaries of prevalence through interactive animations. Finally, we illustrate an application to historical malaria prevalence data from 1334 surveys conducted in Senegal between 1905 and 2014.

stat.ME

On the goodness-of-fit of generalized linear geostatistical models

We propose a generalization of Zhang's coefficient of determination to generalized linear geostatistical models and illustrate its application to river-blindness mapping. The generalized coefficient of determination has a more intuitive interpretation than other measures of predictive performance and allows to assess the individual contribution of each explanatory variable and the random effects to spatial prediction. The developed methodology is also more widely applicable to any generalized linear mixed model.

stat.ME

Geostatistical inference in the presence of geomasking: a composite-likelihood approach

In almost any geostatistical analysis, one of the underlying, often implicit, modelling assump- tions is that the spatial locations, where measurements are taken, are recorded without error. In this study we develop geostatistical inference when this assumption is not valid. This is often the case when, for example, individual address information is randomly altered to provide pri- vacy protection or imprecisions are induced by geocoding processes and measurement devices. Our objective is to develop a method of inference based on the composite likelihood that over- comes the inherent computational limits of the full likelihood method as set out in Fanshawe and Diggle (2011). Through a simulation study, we then compare the performance of our proposed approach with an N-weighted least squares estimation procedure, based on a corrected version of the empirical variogram. Our results indicate that the composite-likelihood approach outper- forms the latter, leading to smaller root-mean-square-errors in the parameter estimates. Finally, we illustrate an application of our method to analyse data on malnutrition from a Demographic and Health Survey conducted in Senegal in 2011, where locations were randomly perturbed to protect the privacy of respondents.

stat.AP

Model-Based Geostatistics for Prevalence Mapping in Low-Resource Settings

In low-resource settings, prevalence mapping relies on empirical prevalence data from a finite, often spatially sparse, set of surveys of communities within the region of interest, possibly supplemented by remotely sensed images that can act as proxies for environmental risk factors. A standard geostatistical model for data of this kind is a generalized linear mixed model with binomial error distribution, logistic link and a combination of explanatory variables and a Gaussian spatial stochastic process in the linear predictor. In this paper, we first review statistical methods and software associated with this standard model, then consider several methodological extensions whose development has been motivated by the requirements of specific applications. These include: methods for combining randomised survey data with data from non-randomised, and therefore potentially biased, surveys; spatio-temporal extensions; spatially structured zero-inflation. Throughout, we illustrate the methods with disease mapping applications that have arisen through our involvement with a range of African public health programmes.

stat.AP

On The Inverse Geostatistical Problem of Inference on Missing Locations

The standard geostatistical problem is to predict the values of a spatially continuous phenomenon, $S(x)$ say, at locations $x$ using data $(y_i,x_i):i=1,..,n$ where $y_i$ is the realization at location $x_i$ of $S(x_i)$, or of a random variable $Y_i$ that is stochastically related to $S(x_i)$. In this paper we address the inverse problem of predicting the locations of observed measurements $y$. We discuss how knowledge of the sampling mechanism can and should inform a prior specification, $π(x)$ say, for the joint distribution of the measurement locations $X = \{x_i: i=1,...,n\}$, and propose an efficient Metropolis-Hastings algorithm for drawing samples from the resulting predictive distribution of the missing elements of $X$. An important feature in many applied settings is that this predictive distribution is multi-modal, which severely limits the usefulness of simple summary measures such as the mean or median. We present two simulated examples to demonstrate the importance of the specification for $π(x)$, and analyze rainfall data from Paraná State, Brazil to show how, under additional assumptions, an empirical of estimate of $π(x)$ can be used when no prior information on the sampling design is available.

stat.AP

On the Computation of Multivariate Scenario Sets for the Skew-t and Generalized Hyperbolic Families

We examine the problem of computing multivariate scenarios sets for skewed distributions. Our interest is motivated by the potential use of such sets in the "stress testing" of insurance companies and banks whose solvency is dependent on changes in a set of financial "risk factors". We define multivariate scenario sets based on the notion of half-space depth (HD) and also introduce the notion of expectile depth (ED) where half-spaces are defined by expectiles rather than quantiles. We then use the HD and ED functions to define convex scenario sets that generalize the concepts of quantile and expectile to higher dimensions. In the case of elliptical distributions these sets coincide with the regions encompassed by the contours of the density function. In the context of multivariate skewed distributions, the equivalence of depth contours and density contours does not hold in general. We consider two parametric families that account for skewness and heavy tails: the generalized hyperbolic and the skew-t distributions. By making use of a canonical form representation, where skewness is completely absorbed by one component, we show that the HD contours of these distributions are "near-elliptical" and, in the case of the skew-Cauchy distribution, we prove that the HD contours are exactly elliptical. We propose a measure of multivariate skewness as a deviation from angular symmetry and show that it can explain the quality of the elliptical approximation for the HD contours.

math.ST