SearcharxivSearch

arXiv subjects

Arnab Hazra

Publications and source records attributed to Arnab Hazra.

At least 19 recordsLinked to original sources

Neural Architectures for Amortized Bayesian Inference: Statistical Foundations and Empirical Assessments

Since the turn of the century, approximate Bayesian inference has steadily evolved as new computational techniques have been incorporated to handle increasingly complex, large-scale predictive problems. The recent success of deep neural networks and foundation models has now given rise to a new paradigm in statistical modeling, in which Bayesian inference can be amortized through large-scale learned predictors. In amortized inference, substantial computation is required at the beginning to train a neural network, but it can subsequently produce approximate posteriors or predictions at much lower computational cost across a wide range of tasks. While the typical Bayesian inference procedures are computationally expensive due to repeated likelihood calculations and Monte Carlo steps for each new dataset, amortized inference provides a much lower computational cost at deployment. Despite the growing popularity of amortized inference, its statistical interpretation and position within Bayesian inference remain poorly explored. In this paper, we present a statistical perspective on several major neural architectures, including feedforward networks, Deep Sets, and Transformers, and examine how they naturally support amortized Bayesian inference. We explore how these models perform structured approximation and also probabilistic reasoning in ways that yield controlled generalization error throughout a wide range of deployment scenarios, and how these properties can be harnessed for Bayesian computation. Via simulation studies, we evaluate the accuracy, robustness, and uncertainty quantification of amortized inference across varying sample sizes, varying noise distributional families, varying sparsity levels, and multimodality, highlighting its strengths and limitations.

stat.ML

Minimum Density Power Divergence Estimation for the Gamma Distribution with Applications to Robust Rainfall Modeling

Statistical modeling of rainfall amounts is of considerable importance in meteorology, hydrology, and agriculture. The gamma distribution remains one of the most popular choices for modeling rainfall data due to its flexibility and ability to capture the skewness of rainfall observations. Rainfall datasets often contain atypical observations due to measurement errors and extreme weather events, making maximum likelihood estimation (MLE) highly sensitive to contamination. In this paper, we develop a robust estimation framework for the two-parameter gamma distribution based on the minimum density power divergence estimator (MDPDE). Explicit estimating equations are derived, and several theoretical properties of the proposed estimators are established. In particular, closed-form expressions for the asymptotic covariance matrix are obtained, and robustness is investigated through influence function analysis and asymptotic relative efficiency. The finite-sample performance of the estimators is examined through simulation studies under both pure and contaminated gamma models. The proposed methodology is further implemented to analyze detrended areally weighted monsoon rainfall data from the 36 meteorological subdivisions of India for the period 1951--2014. The results demonstrate that the MDPDE provides a useful compromise between robustness and efficiency, yielding more stable inference than MLE in the presence of outliers while maintaining high efficiency for uncontaminated data.

stat.ME

Estimating Causal Attribution of Anthropogenic Forcing on High-Temperature Extremes Using a Latent Gaussian Spatial Model

Climate change has become a significant global concern due to its capacity to cause substantial disruption to daily life by increasing the frequency and intensity of extreme weather events. Given the rising trend of human interventions in the climate system over recent decades, this study aims to quantify the relative contribution of anthropogenic forcing to the increasing likelihood of climate extremes, with a particular emphasis on high-temperature extremes. Our analysis focuses on annual temperature maxima from the IPSL-CM6A model in the CMIP6 experiment. We propose a novel causal inference framework that focuses on differences in return levels derived from annual temperature maxima between the factual and counterfactual worlds. While jointly modeling the annual maxima from the two worlds using a bivariate generalized extreme value distribution, we model the spatially-varying coefficients using a latent Gaussian framework. Specifically, given that the data are available over a $1^\circ \times 1^\circ$ grid, we employ the multivariate intrinsic conditional autoregressive model for the latent layer in the proposed hierarchical model, ensuring proper posterior distributions. We implement a recently developed highly-efficient approximate Bayesian inference technique, `Max-and-smooth', that uses a Laplace approximation of the likelihood and then performs Gibbs sampling based on the approximate posterior. The results include posterior estimates of the causal effect of anthropogenic forcing on high-temperature extremes, along with the trends in this effect, over the factual world. Furthermore, we estimate credible regions for a significant causal effect to facilitate hotspot detection across the mainland United States.

stat.AP

Scalable Bayesian inference for high-dimensional mixed-type multivariate spatial data

Spatial generalized linear mixed-effects models are popularly used to analyze spatially indexed univariate responses. However, with modern technology, it is common to observe vector-valued mixed-type responses, e.g., a combination of binary, count, or continuous types, at each location. Methods for jointly modeling such mixed-type multivariate spatial responses are rare. Using multivariate Gaussian processes (GPs) in the latent layer, we present a class of Bayesian spatial methods applicable to any combination of exponential family responses. Since multivariate GP-based methods can suffer from computational bottlenecks when the number of spatial locations is high, we further employ a computationally efficient Vecchia approximation for fast posterior inference and prediction. Key theoretical properties of the proposed model, such as identifiability and the structure of the induced covariance, are established. Our approach employs a Markov chain Monte Carlo-based inference method that uses elliptical slice sampling within a blocked Metropolis-within-Gibbs sampling framework. We illustrate the efficacy of the proposed method through simulation studies and a real-data application on joint modeling of wildfire counts and burnt areas across the United States.

stat.ME

Large Wave Direction Data Modeling Using Wrapped Spatial Gaussian Markov Random Fields

Statistical modeling of dependent directional data remains relatively underexplored, particularly in high-dimensional spatial settings. Existing approaches for spatial angular data primarily rely on wrapped Gaussian process (WGP) models, which provide a coherent framework for capturing spatial dependence on the circle. However, WGP-based methods become computationally challenging when the spatial domain is large, and observations are available at high resolution. This limitation is especially relevant in the analysis of large-scale geological and climate phenomena, such as tsunamis and hurricanes, where directional measurements (e.g., wave or wind directions) may be available over an entire ocean basin. To address these challenges, we propose a wrapped Gaussian Markov random field (WGMRF) model for large spatial directional datasets. By exploiting the sparse precision structure inherent in Gaussian Markov random fields, the proposed approach achieves substantial computational gains while preserving flexible spatial dependence on the circular scale. We discuss key properties of the model, including its identifiability and dependence characteristics. The model fitting involves standard Markov chain Monte Carlo techniques. Through extensive simulation studies and an application to the wave direction data across the Indian Ocean during the 2004 Indian Ocean Tsunami, we compare the proposed method with both a non-spatial wrapped Gaussian model and a low-rank WGP alternative. The results demonstrate that the WGMRF offers improved predictive performance and scalability in large-domain applications.

stat.ME

Regression modeling of multivariate precipitation extremes under regular variation

Motivated by the EVA2025 data challenge, where we participated as the team DesiBoys, we propose a regression strategy within the framework of regular variation to estimate the occurrences and intensities of high precipitation extremes derived from different climate runs of the CESM2 Large Ensemble Community Project (LENS2). Our approach first empirically estimates the target quantities at sub-asymptotic (lower threshold) levels and sets them as response variables within a simple regression framework arising from the theoretical expressions of joint regular variation. Although a seasonal pattern is evident in the data, the precipitation intensities do not exhibit any significant long-term trends across years. Besides, we can safely assume the data to be independent across different climate model runs, thereby simplifying the modeling framework. Once the regression parameters are estimated, we employ a standard prediction approach to infer precipitation levels at very high quantiles. We calculate the confidence intervals using a nonparametric block bootstrap procedure. While a likelihood-based inference grounded in multivariate extreme value theory may provide more accurate estimates and confidence intervals, it would involve a significantly higher computational burden. Our proposed simple and computationally straightforward two-stage approach provides reasonable estimates for the desired quantities, securing us a joint second position in the final rankings of the EVA2025 conference data challenge competition.

stat.ME

A Bayesian latent Gaussian conditional autoregressive copula model for analyzing spatially-varying trends in rainfall

Assessing the availability of rainfall water plays a crucial role in rainfed agriculture. Given the substantial proportion of agricultural practices in India being rainfed and considering the potential trends in rainfall amounts across years due to climate change, we build a statistical model for analyzing monsoon total rainfall data for 34 meteorological subdivisions of mainland India available for 1951-2014. Here, we model the marginal distributions using a gamma regression model and the dependence through a Gaussian conditional autoregressive (CAR) copula model. Due to the natural variation in the monsoon total rainfall received across various dry through wet regions of the country, we allow the parameters of the marginal distributions to be spatially varying, under a latent Gaussian model framework. The neighborhood structure of the regions determines the dependence structure of both the likelihood and the prior layers, where we explore both CAR and intrinsic CAR structures for the priors. The proposed methodology also effectively imputes the missing data. We use the Markov chain Monte Carlo algorithms to draw Bayesian inferences. In simulation studies, the proposed model outperforms several competitors that do not allow a dependence structure at the data or prior layers. Implementing the proposed method for the Indian areal rainfall dataset, we draw inferences about the model parameters and discuss the potential effect of climate change on rainfall across India. While the assessment of the impact of climate change on rainfall motivates our study, the proposed methodology can be easily adapted to other contexts dealing with non-Gaussian non-stationary areal datasets where data from single or multiple temporal covariates are also available, and it is appropriate to assume their coefficients to be spatially varying.

stat.AP

A Review of Statistical and Machine Learning Approaches for Coral Bleaching Assessment

Coral bleaching is a major concern for marine ecosystems; more than half of the world's coral reefs have either bleached or died over the past three decades. Increasing sea surface temperatures, along with various spatiotemporal environmental factors, are considered the primary reasons behind coral bleaching. The statistical and machine learning communities have focused on multiple aspects of the environment in detail. However, the literature on various stochastic modeling approaches for assessing coral bleaching is extremely scarce. Data-driven strategies are crucial for effective reef management, and this review article provides an overview of existing statistical and machine learning methods for assessing coral bleaching. Statistical frameworks, including simple regression models, generalized linear models, generalized additive models, Bayesian regression models, spatiotemporal models, and resilience indicators, such as Fisher's Information and Variance Index, are commonly used to explore how different environmental stressors influence coral bleaching. On the other hand, machine learning methods, including random forests, decision trees, support vector machines, and spatial operators, are more popular for detecting nonlinear relationships, analyzing high-dimensional data, and allowing integration of heterogeneous data from diverse sources. In addition to summarizing these models, we also discuss potential data-driven future research directions, with a focus on constructing statistical and machine learning models in specific contexts related to coral bleaching.

stat.AP

A semiparametric generalized exponential regression model with a principled distance-based prior

The generalized exponential distribution is a well-known probability model in lifetime data analysis and several other research areas, including precipitation modeling. Despite having broad applications for independently and identically distributed observations, its uses as a generalized linear model for non-identically distributed data are limited. This paper introduces a semiparametric Bayesian generalized exponential (GE) regression model. Our proposed approach involves modeling the GE rate parameter within a generalized additive model framework. An important feature is the integration of a principled distance-based prior for the GE shape parameter; this allows the model to shrink to an exponential regression model that retains the advantages of the exponential family. We draw inferences using the Markov chain Monte Carlo algorithm and discuss some theoretical results pertaining to Bayesian asymptotics. Extensive simulations demonstrate that the proposed model outperforms simpler alternatives. The Western Ghats mountain range holds critical importance in regulating monsoon rainfall across Southern India, profoundly impacting regional agriculture. Here, we analyze daily wet-day rainfall data for the monsoon months between 1901--2022 for the Northern, Middle, and Southern Western Ghats regions. Applying the proposed model to analyze the rainfall data over 122 years provides insights into model parameters, short-term temporal patterns, and the impact of climate change. We observe a significant decreasing trend in wet-day rainfall for the Southern Western Ghats region.

stat.AP

Approximate Bayesian inference for high-resolution spatial disaggregation using alternative data sources

This paper addresses the challenge of obtaining precise demographic information at a fine-grained spatial level, a necessity for planning localized public services such as water distribution networks, or understanding local human impacts on the ecosystem. While population sizes are commonly available for large administrative areas, such as wards in India, practical applications often demand knowledge of population density at smaller spatial scales. We explore the integration of alternative data sources, specifically satellite-derived products, including land cover, land use, street density, building heights, vegetation coverage, and drainage density. Using a case study focused on Bangalore City, India, with a ward-level population dataset for 198 wards and satellite-derived sources covering 786,702 pixels at a resolution of 30mX30m, we propose a semiparametric Bayesian spatial regression model for obtaining pixel-level population estimates. Given the high dimensionality of the problem, exact Bayesian inference is deemed impractical; we discuss an approximate Bayesian inference scheme based on the recently proposed max-and-smooth approach, a combination of Laplace approximation and Markov chain Monte Carlo. A simulation study validates the reasonable performance of our inferential approach. Mapping pixel-level estimates to the ward level demonstrates the effectiveness of our method in capturing the spatial distribution of population sizes. While our case study focuses on a demographic application, the methodology developed here readily applies to count-type spatial datasets from various scientific disciplines, where high-resolution alternative data sources are available.

stat.ME

Estimating Changepoints in Extremal Dependence, Applied to Aviation Stock Prices During COVID-19 Pandemic

The dependence in the tails of the joint distribution of two random variables is generally assessed using $χ$-measure, the limiting conditional probability of one variable being extremely high given the other variable is also extremely high. This work is motivated by the structural changes in $χ$-measure between the daily rate of return (RoR) of the two Indian airlines, IndiGo and SpiceJet, during the COVID-19 pandemic. We model the daily maximum and minimum RoR vectors (potentially transformed) using the bivariate Hüsler-Reiss (BHR) distribution. To estimate the changepoint in the $χ$-measure of the BHR distribution, we explore two changepoint detection procedures based on the Likelihood Ratio Test (LRT) and Modified Information Criterion (MIC). We obtain critical values and power curves of the LRT and MIC test statistics for low through high values of $χ$-measure. We also explore the consistency of the estimators of the changepoint based on LRT and MIC numerically. In our data application, for RoR maxima and minima, the most prominent changepoints detected by LRT and MIC are close to the announcement of the first phases of lockdown and unlock, respectively, which are realistic; thus, our study would be beneficial for portfolio optimization in the case of future pandemic situations.

stat.AP

Efficient Modeling of Spatial Extremes over Large Geographical Domains

Various natural phenomena exhibit spatial extremal dependence at short spatial distances. However, existing models proposed in the spatial extremes literature often assume that extremal dependence persists across the entire domain. This is a strong limitation when modeling extremes over large geographical domains, and yet it has been mostly overlooked in the literature. We here develop a more realistic Bayesian framework based on a novel Gaussian scale mixture model, with the Gaussian process component defined by a stochastic partial differential equation yielding a sparse precision matrix, and the random scale component modeled as a low-rank Pareto-tailed or Weibull-tailed spatial process determined by compactly-supported basis functions. We show that our proposed model is approximately tail-stationary and that it can capture a wide range of extremal dependence structures. Its inherently sparse structure allows fast Bayesian computations in high spatial dimensions based on a customized Markov chain Monte Carlo algorithm prioritizing calibration in the tail. We fit our model to analyze heavy monsoon rainfall data in Bangladesh. Our study shows that our model outperforms natural competitors and that it fits precipitation extremes well. We finally use the fitted model to draw inference on long-term return levels for marginal precipitation and spatial aggregates.

stat.ME

Flexible Modeling of Nonstationary Extremal Dependence using Spatially-Fused LASSO and Ridge Penalties

Statistical modeling of a nonstationary spatial extremal dependence structure is challenging. Max-stable processes are common choices for modeling spatially-indexed block maxima, where an assumption of stationarity is usual to make inference feasible. However, this assumption is often unrealistic for data observed over a large or complex domain. We propose a computationally-efficient method for estimating extremal dependence using a globally nonstationary, but locally-stationary, max-stable process by exploiting nonstationary kernel convolutions. We divide the spatial domain into a fine grid of subregions, assign each of them its own dependence parameters, and use LASSO ($L_1$) or ridge ($L_2$) penalties to obtain spatially-smooth parameter estimates. We then develop a novel data-driven algorithm to merge homogeneous neighboring subregions. The algorithm facilitates model parsimony and interpretability. To make our model suitable for high-dimensional data, we exploit a pairwise likelihood to draw inferences and discuss computational and statistical efficiency. An extensive simulation study demonstrates the superior performance of our proposed model and the subregion-merging algorithm over the approaches that either do not model nonstationarity or do not update the domain partition. We apply our proposed method to model monthly maximum temperatures at over 1400 sites in Nepal and the surrounding Himalayan and sub-Himalayan regions; we again observe significant improvements in model fit compared to a stationary process and a nonstationary process without subregion-merging. Furthermore, we demonstrate that the estimated merged partition is interpretable from a geographic perspective and leads to better model diagnostics by adequately reducing the number of subregion-specific parameters.

stat.ME

Computationally Scalable Bayesian SPDE Modeling for Censored Spatial Responses

Observations of groundwater pollutants, such as arsenic or Perfluorooctane sulfonate (PFOS), are riddled with left censoring. These measurements have impact on the health and lifestyle of the populace. Left censoring of these spatially correlated observations are usually addressed by applying Gaussian processes (GPs), which have theoretical advantages. However, this comes with a challenging computational complexity of $\mathcal{O}(n^3)$, which is impractical for large datasets. Additionally, a sizable proportion of the data being left-censored creates further bottlenecks, since the likelihood computation now involves an intractable high-dimensional integral of the multivariate Gaussian density. In this article, we tackle these two problems simultaneously by approximating the GP with a Gaussian Markov random field (GMRF) approach that exploits an explicit link between a GP with Matérn correlation function and a GMRF using stochastic partial differential equations (SPDEs). We introduce a GMRF-based measurement error into the model, which alleviates the likelihood computation for the censored data, drastically improving the speed of the model while maintaining admirable accuracy. Our approach demonstrates robustness and substantial computational scalability, compared to state-of-the-art methods for censored spatial responses across various simulation settings. Finally, the fit of this fully Bayesian model to the concentration of PFOS in groundwater available at 24,959 sites across California, where 46.62\% responses are censored, produces prediction surface and uncertainty quantification in real time, thereby substantiating the applicability and scalability of the proposed method. Code for implementation is made available via GitHub.

stat.ME

Robust statistical modeling of monthly rainfall: The minimum density power divergence approach

Statistical modeling of monthly, seasonal, or annual rainfall data is an important research area in meteorology. These models play a crucial role in rainfed agriculture, where a proper assessment of the future availability of rainwater is necessary. The rainfall amount during a rainy month or a whole rainy season} can take any positive value and some simple (one or two-parameter) probability models supported over the positive real line that are generally used for rainfall modeling are exponential, gamma, Weibull, lognormal, Pearson Type-V/VI, log-logistic, etc., where the unknown model parameters are routinely estimated using the maximum likelihood estimator (MLE). However, the presence of outliers or extreme observations is a common issue in rainfall data and the MLEs being highly sensitive to them often leads to spurious inference. Here, we discuss a robust parameter estimation approach based on the minimum density power divergence estimator (MDPDE). We fit the above four parametric models to the detrended areally-weighted monthly rainfall data from the 36 meteorological subdivisions of India for the years 1951-2014 and compare the fits based on MLE and the proposed optimum MDPDE; the superior performance of MDPDE is showcased for several cases. For all month-subdivision combinations, we discuss the best-fit models and median rainfall amounts.

stat.AP

Minimum Density Power Divergence Estimation for the Generalized Exponential Distribution

Statistical modeling of rainfall data is an active research area in agro-meteorology. The most common models fitted to such datasets are exponential, gamma, log-normal, and Weibull distributions. As an alternative to some of these models, the generalized exponential (GE) distribution was proposed by Gupta and Kundu (2001, Exponentiated Exponential Family: An Alternative to Gamma and Weibull Distributions, Biometrical Journal). Rainfall (specifically for short periods) datasets often include outliers, and thus, a proper robust parameter estimation procedure is necessary. Here, we use the popular minimum density power divergence estimation (MDPDE) procedure developed by Basu et al. (1998, Robust and Efficient Estimation by Minimising a Density Power Divergence, Biometrika) for estimating the GE parameters. We derive the analytical expressions for the estimating equations and asymptotic distributions. We analytically compare MDPDE with maximum likelihood estimation in terms of robustness, through an influence function analysis. Besides, we study the asymptotic relative efficiency of MDPDE analytically for different parameter settings. We apply the proposed technique to some simulated datasets and two rainfall datasets from Texas, United States. The results indicate superior performance of MDPDE compared to the other existing estimation techniques in most of the scenarios.

stat.ME

An utopic adventure in the modelling of conditional univariate and multivariate extremes

The EVA 2023 data competition consisted of four challenges, ranging from interval estimation for very high quantiles of univariate extremes conditional on covariates, point estimation of unconditional return levels under a custom loss function, to estimation of the probabilities of tail events for low and high-dimensional multivariate data. We tackle these tasks by revisiting the current and existing literature on conditional univariate and multivariate extremes. We propose new cross-validation methods for covariate-dependent models, validation metrics for exchangeable multivariate models, formulae for the joint probability of exceedance for multivariate generalized Pareto vectors and a composition sampling algorithm for generating multivariate tail events for the latter. We highlight overarching themes ranging from model validation at extremely high quantile levels to building custom estimation strategies that leverage model assumptions.

stat.AP

Exploring the Efficacy of Statistical and Deep Learning Methods for Large Spatial Datasets: A Case Study

Increasingly large and complex spatial datasets pose massive inferential challenges due to high computational and storage costs. Our study is motivated by the KAUST Competition on Large Spatial Datasets 2023, which tasked participants with estimating spatial covariance-related parameters and predicting values at testing sites, along with uncertainty estimates. We compared various statistical and deep learning approaches through cross-validation and ultimately selected the Vecchia approximation technique for model fitting. To overcome the constraints in the R package GpGp, which lacked support for fitting zero-mean Gaussian processes and direct uncertainty estimation-two things that are necessary for the competition, we developed additional \texttt{R} functions. Besides, we implemented certain subsampling-based approximations and parametric smoothing for skewed sampling distributions of the estimators. Our team DesiBoys secured victory in two out of four sub-competitions, validating the effectiveness of our proposed strategies. Moreover, we extended our evaluation to a large real spatial satellite-derived dataset on total precipitable water, where we compared the predictive performances of different models using multiple diagnostics.

stat.CO