Searcharxiv⌕ Search

arXiv subjects

Soutir Bandyopadhyay

Publications and source records attributed to Soutir Bandyopadhyay.

18 recordsLinked to original sources

Gaussian Process Decorrelation for Spatiotemporal Deep Learning-Based Snow Water Equivalent Prediction

In the Western United States, snowmelt is essential to the agricultural industry in addition to being a key source of municipal drinking water. Consequently, accurate snowpack forecasting is critical for water policy and management. Automated Snow Telemetry (SNOTEL) stations provide accurate daily measurements of snow water equivalent (SWE) that exhibit strong correlations in space and in time. We tackle the problem of predicting future SWE values across the SNOTEL network. Specifically, we use a Gaussian Process-based linear transformation to remove spatial correlations before training a long short-term memory (LSTM) neural network on the decorrelated SWE data. This approach allows the LSTM to learn a clean temporal signal at each station. We show that this separation of spatial and temporal components yields better predictive success than multiple baseline models. Furthermore, we incorporate conformal prediction to quantify uncertainty in the resulting SWE forecasts, providing a distribution-free approach to illustrate a potential framework for establishing predictive intervals for spatiotemporal data. Together, accurate point forecasts and distribution-free uncertainty quantification provide a framework for SWE accumulation forecasting on subseasonal scales or projecting SWE with future data while motivating and supporting future work in predicting a large-scale, spatiotemporally complete SWE map.

cs.LG↗

Spatial function-on-function quantile regression

This paper introduces a novel penalized spatial function-on-function quantile regression framework for analyzing spatially indexed functional data, bridging a critical gap between spatial functional models and quantile regression. Our work makes three key contributions. First, we propose the first spatial function-on-function quantile regression model that jointly accounts for spatial correlation across curves through a functional spatial autoregressive structure while allowing inference on arbitrary conditional quantiles of the functional response. Unlike traditional mean-based alternatives, this approach successfully captures state-dependent volatility and distributional dynamics beyond the conditional mean. Second, we develop a two-stage instrumental-variable estimation strategy to address endogeneity induced by the functional spatial lag. By utilizing tensor-product B-spline expansions with tensor-product roughness penalties, our method ensures optimal smoothness without the destructive information loss inherent in principal component truncation. Third, for fixed spline dimensions, we establish $\sqrt n$-asymptotic normality of the spline coefficient estimators and the induced finite-rank Gaussian-process limits for the reconstructed coefficient surfaces. Extensive Monte Carlo experiments and a high-resolution analysis of Italian PM$_{2.5}$ air quality data demonstrate that spatial function-on-function quantile regression significantly outperforms non-spatial and mean-based competitors, providing a robust and informative tool for environmental risk management and complex functional data analysis. Our method has been implemented in the SpatialFoFReg R package.

stat.ME↗

Frequency Domain Resampling for Gridded Spatial Data

In frequency domain analysis for spatial data, spectral averages based on the periodogram often play an important role in understanding spatial covariance structure, but also have complicated sampling distributions due to complex variances from aggregated periodograms. In order to non-parametrically approximate these sampling distributions for purposes of inference, resampling can be useful, but previous developments in spatial bootstrap have faced challenges in the scope of their validity, specifically due to issues in capturing the complex variances of spatial spectral averages. As a consequence, existing frequency domain bootstraps for spatial data are highly restricted in application to only special processes (e.g. Gaussian) or certain spatial statistics. To address this limitation and to approximate a wide range of spatial spectral averages, we propose a practical hybrid-resampling approach that combines two different resampling techniques in the forms of spatial subsampling and spatial bootstrap. Subsampling helps to capture the variance of spectral averages while bootstrap captures the distributional shape. The hybrid resampling procedure can then accurately quantify uncertainty in spectral inference under mild spatial assumptions. Moreover, compared to the more studied time series setting, this work fills a gap in the theory of subsampling/bootstrap for spatial data regarding spectral average statistics.

math.ST↗

Local Variable and Neighborhood Selection for Firearm Fatality in the Southeast USA

A major public health concern in the United States (US) is gun-related deaths. The number of gun injuries largely varies spatially because of county-wise heterogeneity of race, sex, age, and income distributions. But still, a major challenge is to locally identify the influential socio-economic factors behind these firearm fatality incidents. For a diverging number of predictors, a rich literature exists regarding SCAD under the independence framework; however, a vacuum remains when discussing local variable selection for spatially correlated, over-dispersed data. This research presents a two-step localized variable selection and inference framework for spatially indexed gunshot fatality data. In the first step, we select variables locally using the SCAD penalty for specific locations where the number of gunshot incidents exceeds a threshold. For these locations, after selecting the predictors, we proceed to the next step, which involves examining the directional variation in the latent spatial neighborhood structure. We further discuss the theoretical properties of this county-specific local variable selection under infill asymptotics. This method has threefold advantages: (i) this method selects the variables locally, (ii) this method provides inference about directional variation of a selected predictor, and (iii) instead of assuming the spatial neighborhood structure in an ad hoc manner, this method identifies the specific type of spatial neighborhood structure that is most appropriate for modeling the random effects.

stat.ME↗

Cumulative Logit Ordinal Regression with Proportional Odds under Nonignorable Missing Response -- Application to Phase III Trial

Missing data are inevitable in clinical trials, and trials that produce categorical ordinal responses are not exempted from this. Typically, missing values in the data occur due to different missing mechanisms, such as missing completely at random, missing at random, and missing not at random. Under a specific missing data regime, when the conditional distribution of the missing data is dependent on the ordinal response variable itself along with other predictor variables, then the missing data mechanism is called nonignorable. In this article we propose an expectation maximization based algorithm for fitting a proportional odds regression model when the missing responses are nonignorable. We report results from an extensive simulation study to illustrate the methodology and its finite sample properties. We also apply the proposed method to a recently completed Phase III psoriasis study using an investigational compound. The corresponding SAS program is provided.

stat.ME↗

Modeling Spatial Extremes using Non-Gaussian Spatial Autoregressive Models via Convolutional Neural Networks

Data derived from remote sensing or numerical simulations often have a regular gridded structure and are large in volume, making it challenging to find accurate spatial models that can fill in missing grid cells or simulate the process effectively, especially in the presence of spatial heterogeneity and heavy-tailed marginal distributions. To overcome this issue, we present a spatial autoregressive modeling framework, which maps observations at a location and its neighbors to independent random variables. This is a highly flexible modeling approach and well-suited for non-Gaussian fields, providing simpler interpretability. In particular, we consider the SAR model with Generalized Extreme Value distribution innovations to combine the observation at a central grid location with its neighbors, capturing extreme spatial behavior based on the heavy-tailed innovations. While these models are fast to simulate by exploiting the sparsity of the key matrices in the computations, the maximum likelihood estimation of the parameters is prohibitive due to the intractability of the likelihood, making optimization challenging. To overcome this, we train a convolutional neural network on a large training set that covers a useful parameter space, and then use the trained network for fast parameter estimation. Finally, we apply this model to analyze annual maximum precipitation data from ERA-Interim-driven Weather Research and Forecasting (WRF) simulations, allowing us to explore its spatial extreme behavior across North America.

stat.ML↗

A Statistical Framework for District Energy Long-term Electric Load Forecasting

An accurate forecast of electric demand is essential for the optimal design of a generation system. For district installations, the projected lifespan may extend one or two decades. The reliance on a single-year forecast, combined with a fixed load growth rate, is the current industry standard, but does not support a multi-decade investment. Existing work on long-term forecasting focuses on annual growth rate and/or uses time resolution that is coarser than hourly. To address the gap, we propose multiple statistical forecast models, verified over as long as an 11-year horizon. Combining demand data, weather data, and occupancy trends results in a hybrid statistical model, i.e., generalized additive model (GAM) with a seasonal autoregressive integrated moving average (SARIMA) of the GAM residuals, a multiple linear regression (MLR) model, and a GAM with ARIMA errors model. We evaluate accuracy based on: (i) annual growth rates of monthly peak loads; (ii) annual growth rates of overall energy consumption; (iii) preservation of daily, weekly, and month-to-month trends that occur within each year, known as the 'seasonality' of the data; and, (iv) realistic representation of demand for a full range of weather and occupancy conditions. For example, the models yield an 11-year forecast from a one-year training data set with a normalized root mean square error of 9.091%, a six-year forecast from a one-year training data set with a normalized root mean square error of 8.949%, and a one-year forecast from a 1.2-year training data set with a normalized root mean square error of 6.765%.

stat.AP↗

Adapting Quantile Mapping to Bias Correct Solar Radiation Data

Bias correction is a common pre-processing step applied to climate model data before it is used for further analysis. This article introduces an efficient adaptation of a well-established bias-correction method - quantile mapping - for global horizontal irradiance (GHI) that ensures corrected data is physically plausible through incorporating measurements of clearsky GHI. The proposed quantile mapping method is fit on reanalysis data to first bias correct for regional climate models (RCMs) and is tested on RCMs forced by general circulation models (GCMs) to understand existing biases directly from GCMs. Additionally, we adapt a functional analysis of variance methodology that analyzes sources of remaining biases after implementing the proposed quantile mapping method and considered biases by climate region. This analysis is applied to four sets of climate model output from NA-CORDEX and compared against data from the National Solar Radiation Database produced by the National Renewable Energy Lab.

stat.AP↗

Temporal and spatial downscaling for solar radiation

Global and regional climate model projections are useful for gauging future patterns of climate variables, including solar radiation, but data from these models is often too coarse to assess local impacts. Within the context of solar radiation, the changing climate may have an effect on photovoltaic (PV) production, especially as the PV industry moves to extend plant lifetimes to 50 years. Predicting PV production while taking into account a changing climate requires data at a resolution that is useful for building PV plants. Although temporal and spatial downscaling of solar radiation data is widely studied, we present a novel method to downscale solar radiation data from daily averages to hourly profiles, while maintaining spatial correlation of parameters characterizing the diurnal profile of solar radiation. The method focuses on the use of a diurnal template which can be shifted and scaled according to the time or year and location and the use of thin plate splines for spatial downscaling. This analysis is applied to data from the National Solar Radiation Database housed at the National Renewable Energy Lab and a case study of the mentioned methods over several sub-regions of continental United States is presented.

stat.AP↗

Regridding Uncertainty for Statistical Downscaling of Solar Radiation

Initial steps in statistical downscaling involve being able to compare observed data from regional climate models (RCMs). This prediction requires (1) regridding RCM output from their native grids and at differing spatial resolutions to a common grid in order to be comparable to observed data and (2) bias correcting RCM data, via quantile mapping, for example, for future modeling and analysis. The uncertainty associated with (1) is not always considered for downstream operations in (2). This work examines this uncertainty, which is not often made available to the user of a regridded data product. This analysis is applied to RCM solar radiation data from the NA-CORDEX data archive and observed data from the National Solar Radiation Database housed at the National Renewable Energy Lab. A case study of the mentioned methods over California is presented.

stat.AP↗

Fast parameter estimation of Generalized Extreme Value distribution using Neural Networks

The heavy-tailed behavior of the generalized extreme-value distribution makes it a popular choice for modeling extreme events such as floods, droughts, heatwaves, wildfires, etc. However, estimating the distribution's parameters using conventional maximum likelihood methods can be computationally intensive, even for moderate-sized datasets. To overcome this limitation, we propose a computationally efficient, likelihood-free estimation method utilizing a neural network. Through an extensive simulation study, we demonstrate that the proposed neural network-based method provides Generalized Extreme Value (GEV) distribution parameter estimates with comparable accuracy to the conventional maximum likelihood method but with a significant computational speedup. To account for estimation uncertainty, we utilize parametric bootstrapping, which is inherent in the trained network. Finally, we apply this method to 1000-year annual maximum temperature data from the Community Climate System Model version 3 (CCSM3) across North America for three atmospheric concentrations: 289 ppm $\mathrm{CO}_2$ (pre-industrial), 700 ppm $\mathrm{CO}_2$ (future conditions), and 1400 ppm $\mathrm{CO}_2$, and compare the results with those obtained using the maximum likelihood approach.

stat.ML↗

Adapting conditional simulation using circulant embedding for irregularly spaced spatial data

Computing an ensemble of random fields using conditional simulation is an ideal method for retrieving accurate estimates of a field conditioned on available data and for quantifying the uncertainty of these realizations. Methods for generating random realizations, however, are computationally demanding, especially when the estimates are conditioned on numerous observed data and for large domains. In this article, a \textit{new}, \textit{approximate} conditional simulation approach is applied that builds on \textit{circulant embedding} (CE), a fast method for simulating stationary Gaussian processes. The standard CE is restricted to simulating stationary Gaussian processes (possibly anisotropic) on regularly spaced grids. In this work we explore two possible algorithms, namely local Kriging and nearest neighbor Kriging, that extend CE for irregularly spaced data points. We establish the accuracy of these methods to be suitable for practical inference and the speedup in computation allows for generating conditional fields close to an interactive time frame. The methods are motivated by the U.S. Geological Survey's software \textit{ShakeMap}, which provides near real-time maps of shaking intensity after the occurrence of a significant earthquake. An example for the 2019 event in Ridgecrest, California is used to illustrate our method.

stat.ME↗

Robust Density Power Divergence Estimates for Panel Data Models

The panel data regression models have become one of the most widely applied statistical approaches in different fields of research, including social, behavioral, environmental sciences, and econometrics. However, traditional least-squares-based techniques frequently used for panel data models are vulnerable to the adverse effects of the data contamination or outlying observations that may result in biased and inefficient estimates and misleading statistical inference. In this study, we propose a minimum density power divergence estimation procedure for panel data regression models with random effects to achieve robustness against outliers. The robustness, as well as the asymptotic properties of the proposed estimator, are rigorously established. The finite-sample properties of the proposed method are investigated through an extensive simulation study and an application to climate data in Oman. Our results demonstrate that the proposed estimator exhibits improved performance over some traditional and robust methods in the presence of data contamination.

stat.ME↗

A robust specification test in linear panel data models

The presence of outlying observations may adversely affect statistical testing procedures that result in unstable test statistics and unreliable inferences depending on the distortion in parameter estimates. In spite of the fact that the adverse effects of outliers in panel data models, there are only a few robust testing procedures available for model specification. In this paper, a new weighted likelihood based robust specification test is proposed to determine the appropriate approach in panel data including individual-specific components. The proposed test has been shown to have the same asymptotic distribution as that of most commonly used Hausman's specification test under null hypothesis of random effects specification. The finite sample properties of the robust testing procedure are illustrated by means of Monte Carlo simulations and an economic-growth data from the member countries of the Organisation for Economic Co-operation and Development. Our records reveal that the robust specification test exhibit improved performance in terms of size and power of the test in the presence of contamination.

stat.ME↗

Data Driven Robust Estimation Methods for Fixed Effects Panel Data Models

The panel data regression models have gained increasing attention in different areas of research including but not limited to econometrics, environmental sciences, epidemiology, behavioral and social sciences. However, the presence of outlying observations in panel data may often lead to biased and inefficient estimates of the model parameters resulting in unreliable inferences when the least squares (LS) method is applied. We propose extensions of the M-estimation approach with a data-driven selection of tuning parameters to achieve desirable level of robustness against outliers without loss of estimation efficiency. The consistency and asymptotic normality of the proposed estimators have also been proved under some mild regularity conditions. The finite sample properties of the existing and proposed robust estimators have been examined through an extensive simulation study and an application to macroeconomic data. Our findings reveal that the proposed methods often exhibits improved estimation and prediction performances in the presence of outliers and are consistent with the traditional LS method when there is no contamination.

stat.ME↗

Robust Estimation for Linear Panel Data Models

In different fields of applications including, but not limited to, behavioral, environmental, medical sciences and econometrics, the use of panel data regression models has become increasingly popular as a general framework for making meaningful statistical inferences. However, when the ordinary least squares (OLS) method is used to estimate the model parameters, presence of outliers may significantly alter the adequacy of such models by producing biased and inefficient estimates. In this work we propose a new, weighted likelihood based robust estimation procedure for linear panel data models with fixed and random effects. The finite sample performances of the proposed estimators have been illustrated through an extensive simulation study as well as with an application to blood pressure data set. Our thorough study demonstrates that the proposed estimators show significantly better performances over the traditional methods in the presence of outliers and produce competitive results to the OLS based estimates when no outliers are present in the data set.

stat.ME↗

Rapid Numerical Approximation Method for Integrated Covariance Functions Over Irregular Data Regions

In many practical applications, spatial data are often collected at areal levels (i.e., block data) and the inferences and predictions about the variable at points or blocks different from those at which it has been observed typically depend on integrals of the underlying continuous spatial process. In this paper we describe a method based on Fourier transform by which multiple integrals of covariance functions over irregular data regions may be numerically approximated with the same level of accuracy to traditional methods, but at a greatly reduced computational expense.

stat.CO↗

A frequency domain empirical likelihood method for irregularly spaced spatial data

This paper develops empirical likelihood methodology for irregularly spaced spatial data in the frequency domain. Unlike the frequency domain empirical likelihood (FDEL) methodology for time series (on a regular grid), the formulation of the spatial FDEL needs special care due to lack of the usual orthogonality properties of the discrete Fourier transform for irregularly spaced data and due to presence of nontrivial bias in the periodogram under different spatial asymptotic structures. A spatial FDEL is formulated in the paper taking into account the effects of these factors. The main results of the paper show that Wilks' phenomenon holds for a scaled version of the logarithm of the proposed empirical likelihood ratio statistic in the sense that it is asymptotically distribution-free and has a chi-squared limit. As a result, the proposed spatial FDEL method can be used to build nonparametric, asymptotically correct confidence regions and tests for covariance parameters that are defined through spectral estimating equations, for irregularly spaced spatial data. In comparison to the more common studentization approach, a major advantage of our method is that it does not require explicit estimation of the standard error of an estimator, which is itself a very difficult problem as the asymptotic variances of many common estimators depend on intricate interactions among several population quantities, including the spectral density of the spatial process, the spatial sampling density and the spatial asymptotic structure. Results from a numerical study are also reported to illustrate the methodology and its finite sample properties.

math.ST↗