SearcharxivSearch

arXiv subjects

Murali Haran

Publications and source records attributed to Murali Haran.

At least 19 recordsLinked to original sources

Estimating water levels in the High Plains Aquifer by synthesizing satellite data with groundwater well observations

The High Plains Aquifer (HPA) is a critical water resource in the Central United States, yet its depletion remains a major concern. While the Gravity Recovery and Climate Experiment (GRACE) satellite mission provides large-scale estimates of liquid water equivalent thickness (LWET), its coarse spatial resolution (approx. 24 km) limits local inference. In contrast, well observations from the National Ground-Water Monitoring Network (NGWMN) offer valuable but spatially sparse measurements of depth-to-groundwater. In this paper, we develop a downscaling framework for the satellite data that integrates the two sources and covariates using a Bayesian hierarchical framework. Our model uses a latent Gaussian Markov Random Field (GMRF) that describes the groundwater storage at a high spatial resolution. To address computational issues, we use a basis representation approach specified via the Moran basis. We find that fine-scale covariates like irrigation intensity and precipitation help refine spatial predictions. We thus provide, to our knowledge, the first statistically-rigorous approach for downscaling groundwater information based on GRACE satellite data and NGWMN groundwater measurements. Our approach yields high-resolution estimates of groundwater variations across the HPA from 2002--2022. The resulting fine-scale inference provides valuable insights into groundwater dynamics, highlighting the effects of land use and local extraction patterns.

stat.AP

Algorithms for Models with Intractable Normalizing Functions

In this paper we discuss a well known computing problem -- inference for models with intractable normalizing functions. Models with intractable normalizing functions arise in a wide variety of areas, for instance network models, models for spatial data on lattices, spatial point processes, flexible models for count data and gene expression, and models for permutations. Simulating from these models for fixed parameter values is well studied, starting with work dating back seventy years to the origin of the Metropolis algorithm. On the other hand some of the most practical and theoretically justified algorithms for inference, particularly Bayesian inference, have only been developed within the past two decades. The most computationally efficient algorithms often do not have well developed theory and few if any approaches exist for assessing the quality of approximations based on them. For many problems even the best algorithms can be computationally infeasible. Hence, this is an exciting area of research with many open problems. We explain several key algorithms, providing connections and touching upon practical advantages and disadvantages of each, with some discussion of theoretical properties where they impact practice. We discuss an approach for assessing the accuracy of approximations produced by these algorithms; this diagnostic is particularly valuable for algorithm tuning. While our focus is largely on models with intractable normalizing functions, we also discuss algorithms that are more broadly applicable to models where the entire likelihood function is intractable; these methods are of course also applicable to intractable normalizing function problems.

stat.ME

Bayesian inference for the automultinomial model with an application to landcover data

Multicategory lattice data arise in a wide variety of disciplines such as image analysis, biology, and forestry. We consider modeling such data with the automultinomial model, which can be viewed as a natural extension of the autologistic model to multicategory responses, or equivalently as an extension of the Potts model that incorporates covariate information into a pure-intercept model. The automultinomial model has the advantage of having a unique parameter that controls the spatial correlation. However, the model's likelihood involves an intractable normalizing function of the model parameters that poses serious computational problems for likelihood-based inference. We address this difficulty by performing Bayesian inference through the Double-Metropolis Hastings algorithm, and implement diagnostics to assess the convergence to the target posterior distribution. Through simulation studies and an application to land cover data, we find that the automultinomial model is flexible across a wide range of spatial correlations while maintaining a relatively simple specification. For large data sets we find it also has advantages over spatial generalized linear mixed models. To make this model practical for scientists, we provide recommendations for its specification and computational implementation.

stat.OT

Modeling discrete lattice data using the Potts and tapered Potts models

The Ising and Potts models, among the most important models in statistical physics, have been used for modeling binary and multinomial data on lattices in a wide variety of disciplines such as psychology, image analysis, biology, and forestry. However, these models have several well known shortcomings: (i) they can result in poorly fitting models, that is, simulations from fitted models often do not produce realizations that look like the observed data; (ii) phase transitions and the presence of ground states introduce significant challenges for statistical inference, model interpretation, and goodness of fit; (iii) intractable normalizing constants that are functions of the model parameters pose serious computational problems for likelihood-based inference. Here we develop a tapered version of the Ising and Potts models that addresses issues (i) and (ii). We develop efficient Markov Chain Monte Carlo Maximum Likelihood Estimation (MCMCMLE) algorithms that address issue (iii). We perform an extensive simulation study for the classical and Tapered Potts models that provide insights regarding the issues generated by the phase transition and ground states. Finally, we offer practical recommendations for modeling and computation based on applications of our approach to simulated data as well as data from the 2021 National Land Cover Database.

stat.ME

Probabilistic Downscaling for Flood Hazard Models

Riverine flooding poses significant risks. Developing strategies to manage flood risks requires flood projections with decision-relevant scales and well-characterized uncertainties, often at high spatial resolutions. However, calibrating high-resolution flood models can be computationally prohibitive. To address this challenge, we propose a probabilistic downscaling approach that maps low-resolution model projections onto higher-resolution grids. The existing literature presents two distinct types of downscaling approaches: (1) probabilistic methods, which are versatile and applicable across various physics-based models, and (2) deterministic downscaling methods, specifically tailored for flood hazard models. Both types of downscaling approaches come with their own set of mutually exclusive advantages. Here we introduce a new approach, PDFlood, that combines the advantages of existing probabilistic and flood model-specific downscaling approaches, mainly (1) spatial flooding probabilities and (2) improved accuracy from approximating physical processes. Compared to the state of the art deterministic downscaling approach for flood hazard models, PDFlood allows users to consider previously neglected uncertainties while providing comparable accuracy, thereby better informing the design of risk management strategies. While we develop PDFlood for flood models, the general concepts translate to other applications such as wildfire models.

stat.ME

Fast Bayesian inference for spatial mean-parameterized Conway-Maxwell-Poisson models

Count data with complex features arise in many disciplines, including ecology, agriculture, criminology, medicine, and public health. Zero inflation, spatial dependence, and non-equidispersion are common features in count data. There are two classes of models that allow for these features -- he mode-parameterized Conway--Maxwell--Poisson (COMP) distribution and the generalized Poisson model. However both require the use of either constraints on the parameter space or a parameterization that leads to challenges in interpretability. We propose a spatial mean-parameterized COMP model that retains the flexibility of these models while resolving the above issues. We use a Bayesian spatial filtering approach in order to efficiently handle high-dimensional spatial data and we use reversible-jump MCMC to automatically choose the basis vectors for spatial filtering. The COMP distribution poses two additional computational challenges -- an intractable normalizing function in the likelihood and no closed-form expression for the mean. We propose a fast computational approach that addresses these challenges by, respectively, introducing an efficient auxiliary variable algorithm and pre-computing key approximations for fast likelihood evaluation. We illustrate the application of our methodology to simulated and real datasets, including Texas HPV-cancer data and US vaccine refusal data.

stat.ME

A Class of Models for Large Zero-inflated Spatial Data

Spatially correlated data with an excess of zeros, usually referred to as zero-inflated spatial data, arise in many disciplines. Examples include count data, for instance, abundance (or lack thereof) of animal species and disease counts, as well as semi-continuous data like observed precipitation. Spatial two-part models are a flexible class of models for such data. Fitting two-part models can be computationally expensive for large data due to high-dimensional dependent latent variables, costly matrix operations, and slow mixing Markov chains. We describe a flexible, computationally efficient approach for modeling large zero-inflated spatial data using the projection-based intrinsic conditional autoregression (PICAR) framework. We study our approach, which we call PICAR-Z, through extensive simulation studies and two environmental data sets. Our results suggest that PICAR-Z provides accurate predictions while remaining computationally efficient. An important goal of our work is to allow researchers who are not experts in computation to easily build computationally efficient extensions to zero-inflated spatial models; this also allows for a more thorough exploration of modeling choices in two-part models than was previously possible. We show that PICAR-Z is easy to implement and extend in popular probabilistic programming languages such as nimble and stan.

stat.ME

Measuring Sample Quality in Algorithms for Intractable Normalizing Function Problems

Models with intractable normalizing functions have numerous applications. Because the normalizing constants are functions of the parameters of interest, standard Markov chain Monte Carlo cannot be used for Bayesian inference for these models. A number of algorithms have been developed for such models. Some have the posterior distribution as their asymptotic distribution. Other ``asymptotically inexact'' algorithms do not possess this property. There is limited guidance for evaluating approximations based on these algorithms. Hence it is very hard to tune them. We propose two new diagnostics that address these problems for intractable normalizing function models. Our first diagnostic, inspired by the second Bartlett identity, is in principle broadly applicable to Monte Carlo approximations beyond the normalizing function problem. We develop an approximate version of this diagnostic that is applicable to intractable normalizing function problems. Our second diagnostic is a Monte Carlo approximation to a kernel Stein discrepancy-based diagnostic introduced by Gorham and Mackey (2017). We provide theoretical justification for our methods and apply them to several algorithms in challenging simulated and real data examples including an Ising model, an exponential random graph model, and a Conway--Maxwell--Poisson regression model, obtaining interesting insights about the algorithms in these contexts.

stat.ME

Spatial distribution and determinants of childhood vaccination refusal in the United States

Parental refusal and delay of childhood vaccination has increased in recent years in the United States. This phenomenon challenges maintenance of herd immunity and increases the risk of outbreaks of vaccine-preventable diseases. We examine US county-level vaccine refusal for patients under five years of age collected during the period 2012--2015 from an administrative healthcare dataset. We model these data with a Bayesian zero-inflated negative binomial regression model to capture social and political processes that are associated with vaccine refusal, as well as factors that affect our measurement of vaccine refusal.Our work highlights fine-scale socio-demographic characteristics associated with vaccine refusal nationally, finds that spatial clustering in refusal can be explained by such factors, and has the potential to aid in the development of targeted public health strategies for optimizing vaccine uptake.

stat.AP

A Shared Component Point Process Model for Urban Policing

Newly available point-level datasets allow us to relate police use of force to other events describing police behavior. Current methods for relating two point processes typically rely on the spatial aggregation of one of the two point processes. We investigate new methods that build upon shared component models and case-control methods to retain the point-level nature of both point processes while characterizing the relationship between them. We find that the shared component approach is particularly useful in flexibly relating two point processes, and we illustrate this flexibility in simulated examples and an application to Chicago policing data.

stat.ME

Flood hazard model calibration using multiresolution model output

Riverine floods pose a considerable risk to many communities. Improving flood hazard projections has the potential to inform the design and implementation of flood risk management strategies. Current flood hazard projections are uncertain, especially due to uncertain model parameters. Calibration methods use observations to quantify model parameter uncertainty. With limited computational resources, researchers typically calibrate models using either relatively few expensive model runs at high spatial resolutions or many cheaper runs at lower spatial resolutions. This leads to an open question: Is it possible to effectively combine information from the high and low resolution model runs? We propose a Bayesian emulation-calibration approach that assimilates model outputs and observations at multiple resolutions. As a case study for a riverine community in Pennsylvania, we demonstrate our approach using the LISFLOOD-FP flood hazard model. The multiresolution approach results in improved parameter inference over the single resolution approach in multiple scenarios. Results vary based on the parameter values and the number of available models runs. Our method is general and can be used to calibrate other high dimensional computer models to improve projections.

stat.ME

A Two-Stage Cox Process Model with Spatial and Nonspatial Covariates

Rich new marked point process data allow researchers to consider disparate problems such as the factors affecting the location and type of police use of force incidents, and the characteristics that impact the location and size of forest fires. We develop a two-stage log Gaussian Cox process that models these data in terms of both spatial (community-level) and nonspatial (individual or event-level) characteristics; both types of covariates are present in the examples we consider and are not easy to incorporate via existing methods. Via simulated and real data examples we find that our model is easy to interpret and flexible, accommodating multiple types of marks and multiple types of spatial covariates. In the first example we consider, our approach allows us to study the impact of community-level socioeconomic features such as unemployment as well as event-level features such as officer tenure on force used by police, illustrated through simulated examples. In our second example we consider factors that impact the locations and severity of forest fires from the Castilla-La Mancha region of Spain between 2004-2007.

stat.ME

Fast expectation-maximization algorithms for spatial generalized linear mixed models

Spatial generalized linear mixed models (SGLMMs) are popular and flexible models for non-Gaussian spatial data. They are useful for spatial interpolations as well as for fitting regression models that account for spatial dependence, and are commonly used in many disciplines such as epidemiology, atmospheric science, and sociology. Inference for SGLMMs is typically carried out under the Bayesian framework at least in part because computational issues make maximum likelihood estimation challenging, especially when high-dimensional spatial data are involved. Here we provide a computationally efficient projection-based maximum likelihood approach and two computationally efficient algorithms for routinely fitting SGLMMs. The two algorithms proposed are both variants of expectation maximization algorithm, using either Markov chain Monte Carlo or a Laplace approximation for the conditional expectation. Our methodology is general and applies to both discrete-domain (Gaussian Markov random field) as well as continuous-domain (Gaussian process) spatial models. We show, via simulation and real data applications, that our methods perform well both in terms of parameter estimation as well as prediction. Crucially, our methodology is computationally efficient and scales well with the size of the data and is applicable to problems where maximum likelihood estimation was previously infeasible.

stat.ME

A Space-time Model for Inferring A Susceptibility Map for An Infectious Disease

Motivated by foot-and-mouth disease (FMD) outbreak data from Turkey, we develop a model to estimate disease risk based on a space-time record of outbreaks. The spread of infectious disease in geographical units depends on both transmission between neighbouring units and the intrinsic susceptibility of each unit to an outbreak. Spatially correlated susceptibility may arise from known factors, such as population density, or unknown (or unmeasured) factors such as commuter flows, environmental conditions, or health disparities. Our framework accounts for both space-time transmission and susceptibility. We model the unknown spatially correlated susceptibility as a Gaussian process. We show that the susceptibility surface can be estimated from observed, geo-located time series of infection events and use a projection-based dimension reduction approach which improves computational efficiency. In addition to identifying high risk regions from the Turkey FMD data, we also study how our approach works on the well known England-Wales measles outbreaks data; our latter study results in an estimated susceptibility surface that is strongly correlated with population size, consistent with prior analyses.

stat.AP

PICAR: An Efficient Extendable Approach for Fitting Hierarchical Spatial Models

Hierarchical spatial models are very flexible and popular for a vast array of applications in areas such as ecology, social science, public health, and atmospheric science. It is common to carry out Bayesian inference for these models via Markov chain Monte Carlo (MCMC). Each iteration of the MCMC algorithm is computationally expensive due to costly matrix operations. In addition, the MCMC algorithm needs to be run for more iterations because the strong cross-correlations among the spatial latent variables result in slow mixing Markov chains. To address these computational challenges, we propose a projection-based intrinsic conditional autoregression (PICAR) approach, which is a discretized and dimension-reduced representation of the underlying spatial random field using empirical basis functions on a triangular mesh. Our approach exhibits fast mixing as well as a considerable reduction in computational cost per iteration. PICAR is computationally efficient and scales well to high dimensions. It is also automated and easy to implement for a wide array of user-specified hierarchical spatial models. We show, via simulation studies, that our approach performs well in terms of parameter inference and prediction. We provide several examples to illustrate the applicability of our method, including (i) a high-dimensional cloud cover dataset that showcases its computational efficiency, (ii) a spatially varying coefficient model that demonstrates the ease of implementation of PICAR in the probabilistic programming languages stan and nimble, and (iii) a watershed survey example that illustrates how PICAR applies to models that are not amenable to efficient inference via existing methods.

stat.CO

Reduced-dimensional Monte Carlo Maximum Likelihood for Latent Gaussian Random Field Models

Monte Carlo maximum likelihood (MCML) provides an elegant approach to find maximum likelihood estimators (MLEs) for latent variable models. However, MCML algorithms are computationally expensive when the latent variables are high-dimensional and correlated, as is the case for latent Gaussian random field models. Latent Gaussian random field models are widely used, for example in building flexible regression models and in the interpolation of spatially dependent data in many research areas such as analyzing count data in disease modeling and presence-absence satellite images of ice sheets. We propose a computationally efficient MCML algorithm by using a projection-based approach to reduce the dimensions of the random effects. We develop an iterative method for finding an effective importance function; this is generally a challenging problem and is crucial for the MCML algorithm to be computationally feasible. We find that our method is applicable to both continuous (latent Gaussian process) and discrete domain (latent Gaussian Markov random field) models. We illustrate the application of our methods to challenging simulated and real data examples for which maximum likelihood estimation would otherwise be very challenging. Furthermore, we study an often overlooked challenge in MCML approaches to latent variable models: practical issues in calculating standard errors of the resulting estimates, and assessing whether resulting confidence intervals provide nominal coverage. Our study therefore provides useful insights into the details of implementing MCML algorithms for high-dimensional latent variable models.

stat.CO

A Fast Particle-Based Approach for Calibrating a 3-D Model of the Antarctic Ice Sheet

We consider the scientifically challenging and policy-relevant task of understanding the past and projecting the future dynamics of the Antarctic ice sheet. The Antarctic ice sheet has shown a highly nonlinear threshold response to past climate forcings. Triggering such a threshold response through anthropogenic greenhouse gas emissions would drive drastic and potentially fast sea level rise with important implications for coastal flood risks. Previous studies have combined information from ice sheet models and observations to calibrate model parameters. These studies have broken important new ground but have either adopted simple ice sheet models or have limited the number of parameters to allow for the use of more complex models. These limitations are largely due to the computational challenges posed by calibration as models become more computationally intensive or when the number of parameters increases. Here we propose a method to alleviate this problem: a fast sequential Monte Carlo method that takes advantage of the massive parallelization afforded by modern high performance computing systems. We use simulated examples to demonstrate how our sample-based approach provides accurate approximations to the posterior distributions of the calibrated parameters. The drastic reduction in computational times enables us to provide new insights into important scientific questions, for example, the impact of Pliocene era data and prior parameter information on sea level projections. These studies would be computationally prohibitive with other computational approaches for calibration such as Markov chain Monte Carlo or emulation-based methods. We also find considerable differences in the distributions of sea level projections when we account for a larger number of uncertain parameters.

stat.AP

Ice Model Calibration Using Semi-continuous Spatial Data

Rapid changes in Earth's cryosphere caused by human activity can lead to significant environmental impacts. Computer models provide a useful tool for understanding the behavior and projecting the future of Arctic and Antarctic ice sheets. However, these models are typically subject to large parametric uncertainties due to poorly constrained model input parameters that govern the behavior of simulated ice sheets. Computer model calibration provides a formal statistical framework to infer parameters using observational data, and to quantify the uncertainty in projections due to the uncertainty in these parameters. Calibration of ice sheet models is often challenging because the relevant model output and observational data take the form of semi-continuous spatial data, with a point mass at zero and a right-skewed continuous distribution for positive values. Current calibration approaches cannot handle such data. Here we introduce a hierarchical latent variable model that handles binary spatial patterns and positive continuous spatial patterns as separate components. To overcome challenges due to high-dimensionality we use likelihood-based generalized principal component analysis to impose low-dimensional structures on the latent variables for spatial dependence. We apply our methodology to calibrate a physical model for the Antarctic ice sheet and demonstrate that we can overcome the aforementioned modeling and computational challenges. As a result of our calibration, we obtain improved future ice-volume change projections.

stat.ME