SearcharxivSearch

arXiv subjects

Benjamin M. Taylor

Publications and source records attributed to Benjamin M. Taylor.

13 recordsLinked to original sources

Spatial Latent Gaussian Modelling with Change of Support

Spatial data are often derived from multiple sources (e.g. satellites, in-situ sensors, survey samples) with different supports, but associated with the same properties of a spatial phenomenon of interest. It is common for predictors to also be measured on different spatial supports than the response variables. Although there is no standard way to work with spatial data with different supports, a prevalent approach used by practitioners has been to use downscaling or interpolation to project all the variables of analysis towards a common support, and then using standard spatial models. The main disadvantage with this approach is that simple interpolation can introduce biases and, more importantly, the uncertainty associated with the change of support is not taken into account in parameter estimation. In this article, we propose a Bayesian spatial latent Gaussian model that can handle data with different rectilinear supports in both the response variable and predictors. Our approach allows to handle changes of support more naturally according to the properties of the spatial stochastic process being used, and to take into account the uncertainty from the change of support in parameter estimation and prediction. We use spatial stochastic processes as linear combinations of basis functions where Gaussian Markov random fields define the weights. Our hierarchical modelling approach can be described by the following steps: (i) define a latent model where response variables and predictors are considered as latent stochastic processes with continuous support, (ii) link the continuous-index set stochastic processes with its projection to the support of the observed data, (iii) link the projected process with the observed data. We show the applicability of our approach by simulation studies and modelling land suitability for improved grassland in Rhondda Cynon Taf, a county borough in Wales.

stat.ME

Malaria Risk Mapping Using Routine Health System Incidence Data in Zambia

Improvements to Zambia's malaria surveillance system allow better monitoring of incidence and targetting of responses at refined spatial scales. As transmission decreases, understanding heterogeneity in risk at fine spatial scales becomes increasingly important. However, there are challenges in using health system data for high-resolution risk mapping: health facilities have undefined and overlapping catchment areas, and report on an inconsistent basis. We propose a novel inferential framework for risk mapping of malaria incidence data based on formal down-scaling of confirmed case data reported through the health system in Zambia. We combine data from large community intervention trials in 2011-2016 and model health facility catchments based upon treatment-seeking behaviours; our model for monthly incidence is an aggregated log-Gaussian Cox process, which allows us to predict incidence at fine scale. We predicted monthly malaria incidence at 5km$^2$ resolution nationally: whereas 4.8 million malaria cases were reported through the health system in 2016, we estimated that the number of cases occurring at the community level was closer to 10 million. As Zambia continues to scale up community-based reporting of malaria incidence, these outputs provide realistic estimates of community-level malaria burden as well as high resolution risk maps for targeting interventions at the sub-catchment level.

stat.AP

A Multi-Way Correlation Coefficient

Pearson's correlation is an important summary measure of the amount of dependence between two variables. It is natural to want to generalise the concept of correlation as a single number that measures the inter-relatedness of three or more variables e.g. how `correlated' are a collection of variables in which non are specifically to be treated as an `outcome'? In this short article, we introduce such a measure, and show that it reduces to the modulus of Pearson's $r$ in the two dimensional case.

stat.ME

A Model-Based General Alternative to the Standardised Precipitation Index

In this paper, we introduce two new model-based versions of the widely-used standardized precipitation index (SPI) for detecting and quantifying the magnitude of extreme hydro-climatic events. Our analytical approach is based on generalized additive models for location, scale and shape (GAMLSS), which helps as to overcome some limitations of the SPI. We compare our model-based standardised indices (MBSIs) with the SPI using precipitation data collected between January 2004 - December 2013 (522 weeks) in Caapiranga, a road-less municipality of Amazonas State. As a result, it is shown that the MBSI-1 is an index with similar properties to the SPI, but with improved methodology. In comparison to the SPI, our MBSI-1 index allows for the use of different zero-augmented distributions, it works with more flexible time-scales, can be applied to shorter records of data and also takes into account temporal dependencies in known seasonal behaviours. Our approach is implemented in an R package, mbsi, available from Github.

stat.ME

Mapping food insecurity in the Brazilian Amazon using a spatial item factor analysis model

Food insecurity, a latent construct defined as the lack of consistent access to sufficient and nutritious food, is a pressing global issue with serious health and social justice implications. Item factor analysis is commonly used to study such latent constructs, but it typically assumes independence between sampling units. In the context of food insecurity, this assumption is often unrealistic, as food access is linked to socio-economic conditions and social relations that are spatially structured. To address this, we propose a spatial item factor analysis model that captures spatial dependence, allowing us to predict latent factors at unsampled locations and identify food insecurity hotspots. We develop a Bayesian sampling scheme for inference and illustrate the explanatory strength of our model by analysing household perceptions of food insecurity in Ipixuna, a remote river-dependent urban centre in the Brazilian Amazon. Our approach is implemented in the R package spifa, with further details provided in the Supplementary Material. This spatial extension offers policymakers and researchers a stronger tool for understanding and addressing food insecurity to locate and prioritise areas in greatest need. Our proposed methodology can be applied more widely to other spatially structured latent constructs.

stat.ME

Continuous Inference for Aggregated Point Process Data

This article introduces new methods for inference with count data registered on a set of aggregation units. Such data are omnipresent in epidemiology due to confidentiality issues: it is much more common to know the county in which an individual resides, say, than know their exact location in space. Inference for aggregated data has traditionally made use of models for discrete spatial variation, for example conditional autoregressive models (CAR). We argue that such discrete models can be improved from both a scientific and inferential perspective by using spatiotemporally continuous models to directly model the aggregated counts. We introduce methods for delivering (limiting) continuous inference with spatitemporal aggregated count data in which the aggregation units might change over time and are subject to uncertainty. We illustrate our methods using two examples: from epidemiology, spatial prediction malaria incidence in Namibia; and from politics, forecasting voting under the proposed changes to parlimentary boundaries in the United Kingdom.

stat.ME

Spatial Modelling of Emergency Service Response Times

This article concerns the statistical modelling of emergency service response times. We apply advanced methods from spatial survival analysis to deliver inference for data collected by the London Fire Brigade on response times to reported dwelling fires. Existing approaches to the analysis of these data have been mainly descriptive; we describe and demonstrate the advantages of a more sophisticated approach. Our final parametric proportional hazards model includes harmonic regression terms to describe how response time varies with time-of-day and shared spatially correlated frailties on an auxiliary grid for computational efficiency. We investigate the short-term impact of fire station closures in 2014. Whilst the London Fire Brigade are working hard to keep response times down, our findings suggest there is a limit to what can be achieved logistically: the present article identifies areas around the now closed Belsize, Downham, Kingsland, Knightsbridge, Silvertown, Southwark, Wesminster and Woolwich fire stations in which there should perhaps be some concern as to the provision of fire services.

stat.AP

Auxiliary Variable Markov Chain Monte Carlo for Spatial Survival and Geostatistical Models

This article was motivated by the desire to improve Markov chain Monte Carlo methods for spatial survival models in which the locations of individuals in space are known. For a dataset comprising information on n individuals, standard methods of MCMC-based inference involve computing the inverse of an n by n matrix at each iteration. However with a judicious choice of auxiliary variables on a regular grid with m prediction points it will be shown how to fit an essentially equivalent model but with a substantially reduced computational cost. For a fixed output grid, the computational cost of the new method is reduced from O(n^3) to O(n); the cost of increasing the output grid size being O(m\log m). Furthermore, the new method simultaneously solves the problem of spatial prediction of functions of the latent field, which for standard methods usually presents a further computational challenge. We apply the new method to a spatial survival dataset previously analysed in Henderson et. al 2002 and show how the new method can be applied to spatial and spatiotemporal geostatistical datasets with the same computational benefits.

stat.ME

Spatial and Spatio-Temporal Log-Gaussian Cox Processes: Extending the Geostatistical Paradigm

In this paper we first describe the class of log-Gaussian Cox processes (LGCPs) as models for spatial and spatio-temporal point process data. We discuss inference, with a particular focus on the computational challenges of likelihood-based inference. We then demonstrate the usefulness of the LGCP by describing four applications: estimating the intensity surface of a spatial point process; investigating spatial segregation in a multi-type process; constructing spatially continuous maps of disease risk from spatially discrete data; and real-time health surveillance. We argue that problems of this kind fit naturally into the realm of geostatistics, which traditionally is defined as the study of spatially continuous processes using spatially discrete observations at a finite number of locations. We suggest that a more useful definition of geostatistics is by the class of scientific problems that it addresses, rather than by particular models or data formats.

stat.ME

INLA or MCMC? A Tutorial and Comparative Evaluation for Spatial Prediction in log-Gaussian Cox Processes

We investigate two options for performing Bayesian inference on spatial log-Gaussian Cox processes assuming a spatially continuous latent field: Markov chain Monte Carlo (MCMC) and the integrated nested Laplace approximation (INLA). We first describe the device of approximating a spatially continuous Gaussian field by a Gaussian Markov random field on a discrete lattice, and present a simulation study showing that, with careful choice of parameter values, small neighbourhood sizes can give excellent approximations. We then introduce the spatial log-Gaussian Cox process and describe MCMC and INLA methods for spatial prediction within this model class. We report the results of a simulation study in which we compare MALA and the technique of approximating the continuous latent field by a discrete one, followed by approximate Bayesian inference via INLA over a selection of 18 simulated scenarios. The results question the notion that the latter technique is both significantly faster and more robust than MCMC in this setting; 100,000 iterations of the MALA algorithm running in 20 minutes on a desktop PC delivered greater predictive accuracy than the default \verb=INLA= strategy, which ran in 4 minutes and gave comparative performance to the full Laplace approximation which ran in 39 minutes.

stat.CO

lgcp An R Package for Inference with Spatio-Temporal Log-Gaussian Cox Processes

This paper introduces an R package for spatio-temporal prediction and forecasting for log-Gaussian Cox processes. The main computational tool for these models is Markov chain Monte Carlo and the new package, lgcp, therefore also provides an extensible suite of functions for implementing MCMC algorithms for processes of this type. The modelling framework and details of inferential procedures are first presented before a tour of lgcp functionality is given via a walk-through data-analysis. Topics covered include reading in and converting data, estimation of the key components and parameters of the model, specifying output and simulation quantities, computation of Monte Carlo expectations, post-processing and simulation of data sets.

stat.CO

On Estimating the Ability of NBA Players

This paper introduces a new model and methodology for estimating the ability of NBA players. The main idea is to directly measure how good a player is by comparing how their team performs when they are on the court as opposed to when they are off it. This is achieved in a such a way as to control for the changing abilities of the other players on court at different times during a match. The new method uses multiple seasons' data in a structured way to estimate player ability in an isolated season, measuring separately defensive and offensive merit as well as combining these to give an overall rating. The use of game statistics in predicting player ability will be considered. Results using data from the 2008/9 season suggest that LeBron James, who won the NBA MVP award, was the best overall player. The best defensive player was Lamar Odom and the best rookie was Russell Westbrook, neither of whom won an NBA award that season. The results further indicate that whilst the frequently-reported game statistics provide some information on offensive ability, they do not perform well in the prediction of defensive ability.

stat.AP

An Adaptive Sequential Monte Carlo Sampler

Sequential Monte Carlo (SMC) methods are not only a popular tool in the analysis of state space models, but offer an alternative to MCMC in situations where Bayesian inference must proceed via simulation. This paper introduces a new SMC method that uses adaptive MCMC kernels for particle dynamics. The proposed algorithm features an online stochastic optimization procedure to select the best MCMC kernel and simultaneously learn optimal tuning parameters. Theoretical results are presented that justify the approach and give guidance on how it should be implemented. Empirical results, based on analysing data from mixture models, show that the new adaptive SMC algorithm (ASMC) can both choose the best MCMC kernel, and learn an appropriate scaling for it. ASMC with a choice between kernels outperformed the adaptive MCMC algorithm of Haario et al. (1998) in 5 out of the 6 cases considered.

stat.CO