SearcharxivSearch

arXiv subjects

Geir-Arne Fuglstad

Publications and source records attributed to Geir-Arne Fuglstad.

At least 19 recordsLinked to original sources

Non-stationary Spatial Modeling Using Fractional SPDEs

We construct a Gaussian random field (GRF) that combines fractional smoothness with spatially varying anisotropy. The GRF is defined through a stochastic partial differential equation (SPDE), where the range, marginal variance, and anisotropy vary spatially according to a spectral parametrization of the SPDE coefficients. Priors are constructed to reduce overfitting in this flexible covariance model, and we estimate parameters using a gradient-based optimization approach based on automatic differentiation. In a simulation study, we investigate how many observations are required to reliably estimate fractional smoothness and non-stationarity, and find that one realization containing 500 observations or more is needed in the scenario considered. We also find that the proposed penalization prevents overfitting across varying numbers of observation locations. Two case studies demonstrate that the relative importance of fractional smoothness and non-stationarity is application dependent. Non-stationarity improves predictions in an application to ocean salinity, whereas fractional smoothness improves predictions in an application to precipitation. Predictive ability is assessed using mean squared error and the continuous ranked probability score. In addition to prediction, the proposed approach can be used as a tool to explore the presence of fractional smoothness and non-stationarity.

stat.ME

Flexible covariance structures on metric graphs

Whittle-Matérn (WM) Gaussian random fields (GRFs) are defined as solutions of stochastic partial differential equations (SPDEs) and provide a natural analog of Matérn GRFs on non-Euclidean geometry where the Matérn covariance function is not valid. In particular, WM GRFs on metric graphs have been an active area of research motivated by road and river networks where spatial dependence is more naturally described by intrinsic distances in the network than by Euclidean distances. This family of GRFs is controlled by three parameters relating to marginal variance, spatial range, and smoothness, but can be extended to so-called generalized WM GRFs through spatially varying coefficients in the SPDE. Recent work has considered the use of spatially varying covariates, but the full possibilities of flexibility have not been considered. In this work, we introduce latent GRFs that describe the spatially varying coefficients of the SPDE. This flexible model is compared to less flexible models in a simulation study evaluating both the ability to estimate the covariance structure and predictive ability. An important focus is the number of observations and replications necessary to reliably recover the covariance structure. We find that the flexible model improves over less flexible models in the presence of sufficient data. We also demonstrate practical applicability on traffic counts in a part of Madrid, and observe major differences between in-sample and out-of-sample predictive abilities of the models compared.

stat.ME

ARMA approximation of a Non-separable Spatio-Temporal Model with Fractional Smoothnesses in Space and Time

The Matérn covariance model is ubiquitous in spatial modelling, but there is no default choice for spatio-temporal modelling. In this paper, we consider the recently proposed ``diffusion-based'' extension of the spatial Matérn covariance model to a spatio-temporal non-separable covariance model that allows fractional smoothnesses in space and in time. The model is described in terms of a space-time fractional stochastic partial differential equation, but currently proposed computational approaches have strong restrictions on the possible smoothnesses in time. We propose a discretization method based on rational approximations in time to handle arbitrary smoothnesses, which leads to a vector autoregressive moving average process (VARMA). We prove that the covariance function of the approximation converges pointwise, determine explicit convergence rates as a function of spatial and temporal resolutions and the accuracy of the rational approximation, and conduct numerical verification to demonstrate small pointwise error for low orders of the VARMA process. Through a simulation study, we demonstrate that the parameters can be estimated back and that correctly specifying the temporal smoothness is especially important for forecasting. The approach is illustrated for three months of daily mean temperatures in mainland France.

stat.ME

Sampling distributions for complex design variance estimators in a Fay-Herriot model

Fay-Herriot (FH) models with variance smoothing typically use chi-squared sampling distributions for the design variance estimators. This choice is only valid under strong assumptions on the population and the sampling design, and the choice of sampling distribution is understudied for complex survey designs such as the stratified two-stage clustering design used by the Demographic and Health Surveys (DHS). DHS conducts surveys in low- and middle-income countries and result in low sample sizes for unplanned domains of interest. Thus, accounting for the uncertainty in the estimated design variances is important. We derive two sampling distributions under the DHS design, a simple and a more complex, while clearly specifying and discussing the required superpopulation and design assumptions. In a simulation study, we compare the two sampling distributions to the empirical sampling distributions, and the resulting FH models with variance smoothing to the standard FH model. We find that the standard model exhibits undercoverage, while the variance smoothing models produce better credible intervals according to proper scoring rules. Interestingly, the simple sampling distribution, which is easiest to implement, performs equally as well as the more complex sampling distribution. We illustrate the proposed models by estimating height-for-age z-scores using the 2022 Kenya DHS.

stat.ME

Joint Modelling of Line and Point Data on Metric Graphs

Metric graphs are useful tools for describing spatial domains like road and river networks, where spatial dependence act along the network. We take advantage of recent developments for such Gaussian Random Fields (GRFs), and consider joint spatial modelling of observations with different spatial supports. Motivated by an application to traffic state modelling in Trondheim, Norway, we consider line-referenced data, which can be described by an integral of the GRF along a line segment on the metric graph, and point-referenced data. Through a simulation study inspired by the application, we investigate the number of replicates that are needed to estimate parameters and to predict unobserved locations. The former is assessed using bias and variability, and the latter is assessed through root mean square error (RMSE), continuous rank probability scores (CRPSs), and coverage. Joint modelling is contrasted with a simplified approach that treat line-referenced observations as point-referenced observations. The results suggest joint modelling leads to strong improvements. The application to Trondheim, Norway, combines point-referenced induction loop data and line-referenced public transportation data. To ensure positive speeds, we use a non-linear link function, which requires integrals of non-linear combinations of the linear predictor. This is made computationally feasible by a combination of the R packages inlabru and MetricGraph, and new code for processing geographical line data to work with existing graph representations and fmesher methods for dealing with line support in inlabru on objects from MetricGraph. We fit the model to two datasets where we expect different spatial dependency and compare the results.

stat.ME

Surface finite element approximation of parabolic SPDEs with Whittle--Matérn noise

We propose and analyse a new type of fully discrete surface finite element approximation of a class of linear parabolic stochastic evolution equations with additive noise. Our discretization uses a surface finite element approximation of the noise, and is tailored for equations with noise having covariance operator defined by (negative powers of) elliptic operators, like Whittle--Matérn random fields. We derive strong and pathwise convergence rates of our approximation, and verify these by numerical experiments.

math.NA

The Two Cultures of Prevalence Mapping: Small Area Estimation and Model-Based Geostatistics

In low- and middle-income countries (LMICs), accurate estimates of subnational health and demographic indicators are critical for guiding policy and identifying disparities. Many indicators of interest are proportions of binary outcomes and the task of estimating these fractions is often called prevalence mapping. In LMICs, health and vital records data are limited, so prevalence mapping relies on data from household surveys with complex sampling designs. However, estimates are often desired at spatial resolutions at which data are insufficient. We review two families of approaches to prevalence mapping: small area estimation (SAE) methods (from the survey statistics literature) and model-based geostatistics (MBG) methods (from the spatial statistics literature). SAE models can be ``area-level" or ``unit-level" and commonly use area-specific random effects and rely upon high-quality covariate data from administrative sources. Unit-level models for binary responses are relatively underdeveloped. MBG approaches explicitly specify binary response models, incorporate continuous spatial random effects, and leverage alternative data sources, e.g., satellite imagery. SAE methods often address the design by incorporating sampling weights or modeling the sampling mechanism. Two delicate issues arise when using MBG methods. First, aggregating unit level predictions to create area-level summaries requires population-level information that is rarely available. Second, MBG approaches typically assume the sampling design is ignorable. We review both approaches, and argue that binary response models can be improved using insights from both the survey sampling and the spatial statistics literature. We highlight these issues using household survey data from the Zambia 2018 Demographic Health Survey to estimate subnational HIV prevalence for woman aged 15--49.

stat.AP

Finite element approximation of parabolic SPDEs with Whittle--Matérn noise

We propose and analyse a new type of fully discrete finite element approximation of a class of linear stochastic parabolic evolution equations with additive noise. Our discretization differs from previous ones in that we use a finite element approximation of the noise, as opposed to an $L^2$ projection. This approximation is tailored for equations where the noise has covariance operator defined in terms of (negative powers of) elliptic operators, like Whittle--Matérn random fields. Strong convergence rates up to order $2$ in space and $1$ in time are shown and verified by numerical experiments in dimension $1$ and $2$.

math.NA

Space-Time Smoothing of Survey Outcomes using the R Package SUMMER

The increasing availability of complex survey data, and the continued need for estimates of demographic and health indicators at a fine spatial and temporal scale, which leads to issues of data sparsity, has led to the need for spatio-temporal smoothing methods that acknowledge the manner in which the data were collected. The open source R package SUMMER implements a variety of methods for spatial or spatio-temporal smoothing of survey data. The emphasis is on small-area estimation. We focus primarily on indicators in a low and middle-income countries context. Our methods are particularly useful for data from Demographic Health Surveys and Multiple Indicator Cluster Surveys. We build upon functions within the survey package, and use INLA for fast Bayesian computation. This paper includes a brief overview of these methods and illustrates the workflow of accessing and processing surveys, estimating subnational child mortality rates, and visualizing results with both simulated data and DHS surveys.

stat.AP

Non-stationary Spatio-Temporal Modeling Using the Stochastic Advection-Diffusion Equation

We construct flexible spatio-temporal models through stochastic partial differential equations (SPDEs) where both diffusion and advection can be spatially varying. Computations are done through a Gaussian Markov random field approximation of the solution of the SPDE, which is constructed through a finite volume method. The new flexible non-separable model is compared to a flexible separable model both for reconstruction and forecasting, and evaluated in terms of root mean square errors and continuous rank probability scores. A simulation study demonstrates that the non-separable model performs better when the data is simulated from a non-separable model with diffusion and advection. Further, we estimate surrogate models for emulating the output of a ocean model in Trondheimsfjorden, Norway, and simulate observations of autonomous underwater vehicles. The results show that the flexible non-separable model outperforms the flexible separable model for real-time prediction of unobserved locations.

stat.ME

A joint model for DHS and MICS surveys: Spatial modeling with anonymized locations

Anonymizing the GPS locations of observations can bias a spatial model's parameter estimates and attenuate spatial predictions when improperly accounted for, and is relevant in applications from public health to paleoseismology. In this work, we demonstrate that a newly introduced method for geostatistical modeling in the presence of anonymized point locations can be extended to account for more general kinds of positional uncertainty due to location anonymization, including both jittering (a form of random perturbations of GPS coordinates) and geomasking (reporting only the name of the area containing the true GPS coordinates). We further provide a numerical integration scheme that flexibly accounts for the positional uncertainty as well as spatial and covariate information. We apply the method to women's secondary education completion data in the 2018 Nigeria demographic and health survey (NDHS) containing jittered point locations, and the 2016 Nigeria multiple indicator cluster survey (NMICS) containing geomasked locations. We show that accounting for the positional uncertainty in the surveys can improve predictions in terms of their continuous rank probability score.

stat.ME

Impact of Jittering on Raster- and Distance-based Geostatistical Analyses of DHS Data

Fine-scale covariate rasters are routinely used in geostatistical models for mapping demographic and health indicators based on household surveys from the Demographic and Health Surveys (DHS) program. However, the geostatistical analyses ignore the fact that GPS coordinates in DHS surveys are jittered for privacy purposes. We demonstrate the need to account for this jittering, and we propose a computationally efficient approach that can be routinely applied. We use the new method to analyse the prevalence of completion of secondary education for 20--49 year old women in Nigeria in 2018 based on the 2018 DHS survey. The analysis demonstrates substantial changes in the estimates of spatial range and fixed effects compared to when we ignore jittering. Through a simulation study that mimics the dataset, we demonstrate that accounting for jittering reduces attenuation in the estimated coefficients for covariates and improves predictions. The results also show that the common approach of averaging covariate values in windows around the observed locations does not lead to the same improvements as accounting for jittering.

stat.AP

GeoAdjust: Adjusting for Positional Uncertainty in Geostatistial Analysis of DHS Data

The R-package GeoAdjust https://github.com/umut-altay/GeoAdjust-package implements fast empirical Bayesian geostatistical inference for household survey data from the Demographic and Health Surveys Program (DHS) using Template Model Builder (TMB). DHS household survey data is an important source of data for tracking demographic and health indicators, but positional uncertainty has been intentionally introduced in the GPS coordinates to preserve privacy. GeoAdjust accounts for such positional uncertainty in geostatistical models containing both spatial random effects and raster- and distance-based covariates. The R package supports Gaussian, binomial and Poisson likelihoods with identity link, logit link, and log link functions respectively. The user defines the desired model structure by setting a small number of function arguments, and can easily experiment with different hyperparameters for the priors. GeoAdjust is the first software package that is specifically designed to address positional uncertainty in the GPS coordinates of point referenced household survey data. The package provides inference for model parameters and can predict values at unobserved locations.

stat.CO

High Resolution Global Precipitation Downscaling with Latent Gaussian Models and Nonstationary SPDE Structure

Obtaining high-resolution maps of precipitation data can provide key insights to stakeholders to assess a sustainable access to water resources at urban scale. Mapping a nonstationary, sparse process such as precipitation at very high spatial resolution requires the interpolation of global datasets at the location where ground stations are available with statistical models able to capture complex non-Gaussian global space-time dependence structures. In this work, we propose a new approach based on capturing the spatial dependence of a latent Gaussian process via a locally deformed Stochastic Partial Differential Equation (SPDE) with a buffer allowing for a different spatial structure across land and sea. The finite volume approximation of the SPDE, coupled with Integrated Nested Laplace Approximation ensures feasible Bayesian inference for tens of millions of observations. The simulation studies showcase the improved predictability of the proposed approach against stationary and no-buffer alternatives. The proposed approach is then used to yield high resolution simulations of daily precipitation across the United States.

stat.AP

Spatially Varying Anisotropy for Gaussian Random Fields in Three-Dimensional Space

Isotropic covariance structures can be unreasonable for phenomena in three-dimensional spaces such as the ocean. In the ocean, the variability of the response may vary with depth, and ocean currents may lead to spatially varying anisotropy. We construct a class of non-stationary anisotropic Gaussian random fields (GRFs) in three dimensions through stochastic partial differential equations (SPDEs) where computations are done using Gaussian Markov random field approximations. The approach is proven in a simulation study where the amount of data required to estimate these models is explored. Then, the method is applied to construct a GRF prior on an ocean mass outside Trondheim, Norway, based on simulations from the complex numerical ocean model SINMOD. This GRF prior is compared to a stationary anisotropic GRF using in-situ measurements collected with an autonomous underwater vehicle where our approach outperforms the stationary anisotropic GRF for real-time prediction of unobserved locations.

stat.ME

Fast geostatistical inference under positional uncertainty: Analysing DHS household survey data

Household survey data from the Demographic and Health Surveys (DHS) Program is published with GPS coordinates. However, almost all geostatistical analyses of such data ignore that the published GPS coordinates are randomly displaced (jittered). In this short report, we develop a geostatistical model that accounts for the positional uncertainty when analysing DHS surveys, and provide a fast implementation using Template Model Builder. The key focus is inference with Gaussian random fields under positional uncertainty, and our approach works for both Gaussian and non-Gaussian likelihoods. A simulation study with a binomial observation model shows that the new approach performs equally or better than the common approach of ignoring jittering, both in terms of more accurate parameter estimates and improved predictive measures. We demonstrate that the improvement would be larger under stronger jittering. An analysis of contraceptive use in Kenya shows that the approach is fast and easy to use in practice.

stat.CO

Spatial Aggregation with Respect to a Population Distribution

Spatial aggregation with respect to a population distribution involves estimating aggregate quantities for a population based on an observation of individuals in a subpopulation. In this context, a geostatistical workflow must account for three major sources of `aggregation error': aggregation weights, fine scale variation, and finite population variation. However, common practice is to treat the unknown population distribution as a known population density and ignore empirical variability in outcomes. We improve common practice by introducing a `sampling frame model' that allows aggregation models to account for the three sources of aggregation error simply and transparently. We compare the proposed and the traditional approach using two simulation studies that mimic neonatal mortality rate (NMR) data from the 2014 Kenya Demographic and Health Survey (KDHS2014). For the traditional approach, undercoverage/overcoverage depends arbitrarily on the aggregation grid resolution, while the new approach exhibits low sensitivity. The differences between the two aggregation approaches increase as the population of an area decreases. The differences are substantial at the second administrative level and finer, but also at the first administrative level for some population quantities. We find differences between the proposed and traditional approach are consistent with those we observe in an application to NMR data from the KDHS2014.

stat.ME

makemyprior: Intuitive Construction of Joint Priors for Variance Parameters in R

Priors allow us to robustify inference and to incorporate expert knowledge in Bayesian hierarchical models. This is particularly important when there are random effects that are hard to identify based on observed data. The challenge lies in understanding and controlling the joint influence of the priors for the variance parameters, and makemyprior is an R package that guides the formulation of joint prior distributions for variance parameters. A joint prior distribution is constructed based on a hierarchical decomposition of the total variance in the model along a tree, and takes the entire model structure into account. Users input their prior beliefs or express ignorance at each level of the tree. Prior beliefs can be general ideas about reasonable ranges of variance values and need not be detailed expert knowledge. The constructed priors lead to robust inference and guarantee proper posteriors. A graphical user interface facilitates construction and assessment of different choices of priors through visualization of the tree and joint prior. The package aims to expand the toolbox of applied researchers and make priors an active component in their Bayesian workflow.

stat.CO