Searcharxiv⌕ Search

arXiv subjects

Ben Seiyon Lee

Publications and source records attributed to Ben Seiyon Lee.

13 recordsLinked to original sources

A Scalable Variational Bayes Approach for Fitting Non-Conjugate Spatial Generalized Linear Mixed Models via Basis Expansions

Large spatial datasets with non-Gaussian responses are increasingly common in environmental monitoring, ecology, and remote sensing, yet scalable Bayesian inference for such data remains challenging. Markov chain Monte Carlo methods are often prohibitive for large datasets, and existing variational Bayes methods rely on conjugacy or strong approximations that limit their applicability and can underestimate posterior variances. A scalable variational framework that incorporates semi-implicit variational inference (SIVI) with basis representations of spatial generalized linear mixed models, which may not have conjugacy, is proposed. The proposed framework accommodates gamma, negative binomial, Poisson, Bernoulli, and Gaussian responses on continuous spatial domains. Across 20 simulation scenarios with 50,000 locations, SIVI achieves predictive accuracy and posterior distributions comparable to Metropolis--Hastings and Hamiltonian Monte Carlo while providing notable computational speedups. Applications to remotely-sensed land surface temperature and blue jay abundance further demonstrate the utility of the approach for large non-Gaussian spatial datasets.

stat.ME↗

REX-SUB: A Scalable Subsampling Strategy for Modeling Large Spatial Datasets

Recent advances in data collection technologies have led to the emergence of massive spatial datasets, with measurements obtained at millions of spatial locations. Geostatistical models typically employ Gaussian processes (GPs) to capture spatial dependence, but standard GP fitting becomes prohibitive at such scales. A promising solution is optimal subsampling, where a subset of locations is selected that optimizes a criterion. In this study, we propose a randomized exchange algorithm for subsampling (REX-SUB) which efficiently selects small subsamples that minimize prediction errors in the fitted spatial GP models. To further improve computational efficiency, we embed a scalable Vecchia approximation to the GP's joint likelihood, which takes advantage of sparsity in the precision matrix to enable fast inference on the selected subsamples. Through a simulation study and an application to a remotely sensed precipitable water dataset, we show that REX-SUB yields lower mean squared prediction errors and interval scores compared to competing subsampling strategies.

stat.ME↗

A Cubing Strategy for Identifying Stable Hyperparameter Regions for Uncertainty Quantification in Spatial Deep Learning

Spatially referenced datasets have become increasingly prevalent across many fields, largely driven by advances in data collection methods such as satellite remote sensing. In many applications, predictions at unobserved locations are accompanied by reliable uncertainty estimates. While deep learning methods provide both scalable and accurate models for spatial predictions, there remains no clear consensus for addressing uncertainty quantification in spatial deep learning. Monte Carlo (MC) dropout has become a popular approach for uncertainty quantification, yet existing implementations typically focus on tuning the dropout rate while fixing other influential hyperparameters, such as weight decay and the predictive standard deviation multiplier, often through ad-hoc or manual tuning. We propose a cubing-based diagnostic framework that recursively partitions the hyperparameter space to identify stable regions where MC dropout yields well-calibrated predictive intervals. The approach evaluates hyperparameter regions using scoring rules relative to a statistical baseline model, which serves as a calibration anchor. Through a simulation study spanning multiple spatial dependence regimes as well as a large remotely-sensed land surface temperature dataset, we demonstrate that our approach produces competitive or superior predictive intervals compared to the baseline model. Our methodology provides practitioners with a systematic procedure for incorporating uncertainty quantification into spatial deep learning models.

stat.CO↗

Spatial Extremes at Scale: A Case Study of Surface Skin Temperature and Heat Risk in the United States

Understanding and mapping extreme heat is critical for risk management and public health planning, particularly in regions with complex terrain and heterogeneous climate. We present a case study of extreme heat in the Four Corners region of the United States, using high-resolution surface skin temperature data from the North American Land Data Assimilation System to characterize spatially heterogeneous and seasonally varying extremes across complex terrain, and to assess their implications for heat-related public health risks. Spatial extremes exhibit complex dependencies across geographic regions, which require sophisticated statistical models to capture. While recent advances in spatial extreme value modeling provide flexible representations of joint tail dependencies, statistical inference remains computationally demanding, especially for datasets with a large number of locations. To address this, we propose a random scale mixture process that facilitates Bayesian inference of spatial extremes, and develop scalable inference strategies that leverage advances in spatial modeling and amortized learning. We evaluate the proposed inference methods through large-scale simulation studies, representing the first such extensive study in spatial extremes, and a high-resolution surface skin temperature application in the Four Corners region. Surface skin temperature is particularly useful as a predictor for air temperature, for studying heatwaves and related environmental phenomena, and to calculate heat indices reflecting downstream health risks at any location. Our findings provide insights into efficient, data-driven approaches for modeling spatial extremes, and serve as guidelines for practitioners in the fields of climate science, environmental risk assessment, and beyond.

stat.AP↗

Bayesian Spatiotemporal Nonstationary Model Quantifies Robust Increases in Daily Extreme Rainfall Across the Western Gulf Coast

Precipitation exceedance probabilities are widely used in engineering design, risk assessment, and floodplain management. While common approaches like NOAA Atlas 14 assume that extreme precipitation characteristics are stationary over time, this assumption may underestimate current and future hazards due to anthropogenic climate change. However, the incorporation of nonstationarity in the statistical modeling of extreme precipitation has faced practical challenges that have restricted its applications. In particular, random sampling variability challenges the reliable estimation of trends and parameters, especially when observational records are limited. To address this methodological gap, we propose the Spatially Varying Covariates Model, a hierarchical Bayesian spatial framework that integrates nonstationarity and regionalization for robust frequency analysis of extreme precipitation. This model draws from extreme value theory, spatial statistics, and Bayesian statistics, and is validated through cross-validation and multiple performance metrics. Applying this framework to a case study of daily rainfall in the Western Gulf Coast, we identify robustly increasing trends in extreme precipitation intensity and variability throughout the study area, with notable spatial heterogeneity. This flexible model accommodates stations with varying observation records, yields smooth return level estimates, and can be straightforwardly adapted to the analysis of precipitation frequencies at different durations and for other regions.

stat.AP↗

A Scalable Variational Bayes Approach to Fit High-dimensional Spatial Generalized Linear Mixed Models

Gaussian and discrete non-Gaussian spatial datasets are common across fields like public health, ecology, geosciences, and social sciences. Bayesian spatial generalized linear mixed models (SGLMMs) are a flexible class of models for analyzing such data, but they struggle to scale to large datasets. Many scalable Bayesian methods, built upon basis representations or sparse covariance matrices, still rely on posterior sampling via Markov chain Monte Carlo (MCMC). Variational Bayes (VB) methods have been applied to SGLMMs, but only for small areal datasets. We propose two computationally efficient VB approaches for analyzing moderately sized and massive (millions of locations) Gaussian and discrete non-Gaussian spatial data in the continuous spatial domain. Our methods leverage semi-parametric approximations of latent spatial processes and parallel computing to ensure computational efficiency. The proposed methods deliver inferential and predictive performance comparable to gold-standard MCMC methods while achieving computational speedups of up to 3600 times. In most cases, our VB approaches outperform state-of-the-art alternatives such as INLA and Hamiltonian Monte Carlo. We validate our methods through a comparative numerical study and applications to real-world datasets. These VB approaches can enable practitioners to model millions of discrete non-Gaussian spatial observations on standard laptops, significantly expanding access to advanced spatial modeling tools.

stat.ME↗

A Class of Models for Large Zero-inflated Spatial Data

Spatially correlated data with an excess of zeros, usually referred to as zero-inflated spatial data, arise in many disciplines. Examples include count data, for instance, abundance (or lack thereof) of animal species and disease counts, as well as semi-continuous data like observed precipitation. Spatial two-part models are a flexible class of models for such data. Fitting two-part models can be computationally expensive for large data due to high-dimensional dependent latent variables, costly matrix operations, and slow mixing Markov chains. We describe a flexible, computationally efficient approach for modeling large zero-inflated spatial data using the projection-based intrinsic conditional autoregression (PICAR) framework. We study our approach, which we call PICAR-Z, through extensive simulation studies and two environmental data sets. Our results suggest that PICAR-Z provides accurate predictions while remaining computationally efficient. An important goal of our work is to allow researchers who are not experts in computation to easily build computationally efficient extensions to zero-inflated spatial models; this also allows for a more thorough exploration of modeling choices in two-part models than was previously possible. We show that PICAR-Z is easy to implement and extend in popular probabilistic programming languages such as nimble and stan.

stat.ME↗

Flood hazard model calibration using multiresolution model output

Riverine floods pose a considerable risk to many communities. Improving flood hazard projections has the potential to inform the design and implementation of flood risk management strategies. Current flood hazard projections are uncertain, especially due to uncertain model parameters. Calibration methods use observations to quantify model parameter uncertainty. With limited computational resources, researchers typically calibrate models using either relatively few expensive model runs at high spatial resolutions or many cheaper runs at lower spatial resolutions. This leads to an open question: Is it possible to effectively combine information from the high and low resolution model runs? We propose a Bayesian emulation-calibration approach that assimilates model outputs and observations at multiple resolutions. As a case study for a riverine community in Pennsylvania, we demonstrate our approach using the LISFLOOD-FP flood hazard model. The multiresolution approach results in improved parameter inference over the single resolution approach in multiple scenarios. Results vary based on the parameter values and the number of available models runs. Our method is general and can be used to calibrate other high dimensional computer models to improve projections.

stat.ME↗

Effects of Mixed Distribution Statistical Flood Frequency Models on Dam Safety Assessments: A Case Study of the Pueblo Dam, USA

Statistical flood frequency analysis coupled with hydrograph scaling is commonly used to generate design floods to assess dam safety assessment. The safety assessments can be highly sensitive to the choice of the statistical flood frequency model. Standard dam safety assessments are typically based on a single distribution model of flood frequency, often the Log Pearson Type III or Generalized Extreme Value distributions. Floods, however, may result from multiple physical processes such as rain on snow, snowmelt or rainstorms. This can result in a mixed distribution of annual peak flows, according to the cause of each flood. Engineering design choices based on a single distribution statistical model are vulnerable to the effects of this potential structural model error. To explore the practicality and potential value of implementing mixed distribution statistical models in engineering design, we compare the goodness of fit of several single- and mixed-distribution peak flow models, as well as the contingent dam safety assessment at Pueblo, Colorado as a didactic example. Summer snowmelt and intense summer rainstorms are both key drivers of annual peak flow at Pueblo. We analyze the potential implications for the annual probability of overtopping-induced failure of the Pueblo Dam as a didactic example. We address the temporal and physical cause separation problems by building on previous work with mixed distributions. We find a Mixed Generalized Extreme Value distribution model best fits peak flows observed in the gaged record, historical floods, and paleo floods at Pueblo. Finally, we show that accounting for mixed distributions in the safety assessment at Pueblo Dam increases the assessed risk of overtopping.

stat.AP↗

A safety factor approach to designing urban infrastructure for dynamic conditions

Current approaches to design flood-sensitive infrastructure typically assume a stationary rainfall distribution and neglect many uncertainties. These assumptions are inconsistent with observations that suggest intensifying extreme precipitation events and the uncertainties surrounding projections of the coupled natural-human systems. Here we demonstrate a safety factor approach to designing urban infrastructure in a changing climate. Our results show that assuming climate stationarity and neglecting deep uncertainties can drastically underestimate flood risks and lead to poor infrastructure design choices. We find that climate uncertainty dominates the socioeconomic and engineering uncertainties that impact the hydraulic reliability in stormwater drainage systems. We quantify the upfront costs needed to achieve higher hydraulic reliability and robustness against the deep uncertainties surrounding projections of rainfall, surface runoff characteristics, and infrastructure lifetime. Depending on the location, we find that adding safety factors of 1.4 to 1.7 to the standard stormwater pipe design guidance produces robust performance to the considered deep uncertainties. The insights gained from this study highlight the need for updating traditional engineering design strategies to improve infrastructure reliability under socioeconomic and environmental changes.

stat.AP↗

PICAR: An Efficient Extendable Approach for Fitting Hierarchical Spatial Models

Hierarchical spatial models are very flexible and popular for a vast array of applications in areas such as ecology, social science, public health, and atmospheric science. It is common to carry out Bayesian inference for these models via Markov chain Monte Carlo (MCMC). Each iteration of the MCMC algorithm is computationally expensive due to costly matrix operations. In addition, the MCMC algorithm needs to be run for more iterations because the strong cross-correlations among the spatial latent variables result in slow mixing Markov chains. To address these computational challenges, we propose a projection-based intrinsic conditional autoregression (PICAR) approach, which is a discretized and dimension-reduced representation of the underlying spatial random field using empirical basis functions on a triangular mesh. Our approach exhibits fast mixing as well as a considerable reduction in computational cost per iteration. PICAR is computationally efficient and scales well to high dimensions. It is also automated and easy to implement for a wide array of user-specified hierarchical spatial models. We show, via simulation studies, that our approach performs well in terms of parameter inference and prediction. We provide several examples to illustrate the applicability of our method, including (i) a high-dimensional cloud cover dataset that showcases its computational efficiency, (ii) a spatially varying coefficient model that demonstrates the ease of implementation of PICAR in the probabilistic programming languages stan and nimble, and (iii) a watershed survey example that illustrates how PICAR applies to models that are not amenable to efficient inference via existing methods.

stat.CO↗

A Fast Particle-Based Approach for Calibrating a 3-D Model of the Antarctic Ice Sheet

We consider the scientifically challenging and policy-relevant task of understanding the past and projecting the future dynamics of the Antarctic ice sheet. The Antarctic ice sheet has shown a highly nonlinear threshold response to past climate forcings. Triggering such a threshold response through anthropogenic greenhouse gas emissions would drive drastic and potentially fast sea level rise with important implications for coastal flood risks. Previous studies have combined information from ice sheet models and observations to calibrate model parameters. These studies have broken important new ground but have either adopted simple ice sheet models or have limited the number of parameters to allow for the use of more complex models. These limitations are largely due to the computational challenges posed by calibration as models become more computationally intensive or when the number of parameters increases. Here we propose a method to alleviate this problem: a fast sequential Monte Carlo method that takes advantage of the massive parallelization afforded by modern high performance computing systems. We use simulated examples to demonstrate how our sample-based approach provides accurate approximations to the posterior distributions of the calibrated parameters. The drastic reduction in computational times enables us to provide new insights into important scientific questions, for example, the impact of Pliocene era data and prior parameter information on sea level projections. These studies would be computationally prohibitive with other computational approaches for calibration such as Markov chain Monte Carlo or emulation-based methods. We also find considerable differences in the distributions of sea level projections when we account for a larger number of uncertain parameters.

stat.AP↗

Deep uncertainties in sea-level rise and storm surge projections: Implications for coastal flood risk management

Sea-levels are rising in many areas around the world, posing risks to coastal communities and infrastructures. Strategies for managing these flood risks present decision challenges that require a combination of geophysical, economic, and infrastructure models. Previous studies have broken important new ground on the considerable tensions between the costs of upgrading infrastructure and the damages that could result from extreme flood events. However, many risk-based adaptation strategies remain silent on certain potentially important uncertainties, as well as the trade-offs between competing objectives. Here, we implement and improve on a classic decision-analytical model (van Dantzig 1956) to: (i) capture trade-offs across conflicting stakeholder objectives, (ii) demonstrate the consequences of structural uncertainties in the sea-level rise and storm surge models, and (iii) identify the parametric uncertainties that most strongly influence each objective using global sensitivity analysis. We find that the flood adaptation model produces potentially myopic solutions when formulated using traditional mean-centric decision theory. Moving from a single-objective problem formulation to one with multi-objective trade-offs dramatically expands the decision space, and highlights the need for compromise solutions to address stakeholder preferences. We find deep structural uncertainties that have large effects on the model outcome, with the storm surge parameters accounting for the greatest impacts. Global sensitivity analysis effectively identifies important parameter interactions that local methods overlook, and which could have critical implications for flood adaptation strategies.

stat.AP↗