SearcharxivSearch

arXiv subjects

Debjoy Thakur

Publications and source records attributed to Debjoy Thakur.

7 recordsLinked to original sources

Variational Approximated Restricted Maximum Likelihood Estimation for Spatial Data

This research considers a scalable inference for spatial data modeled through Gaussian intrinsic conditional autoregressive (ICAR) structures. The classical estimation method, restricted maximum likelihood (REML), requires repeated inversion and factorization of large, sparse precision matrices, which makes this computation costly. To sort this problem out, we propose a variational restricted maximum likelihood (VREML) framework that approximates the intractable marginal likelihood using a Gaussian variational distribution. By constructing an evidence lower bound (ELBO) on the restricted likelihood, we derive a computationally efficient coordinate-ascent algorithm for jointly estimating the spatial random effects and variance components. In this article, we theoretically establish the monotone convergence of ELBO and mathematically exhibit that the variational family is exact under Gaussian ICAR settings, which is an indication of nullifying approximation error at the posterior level. We empirically establish the supremacy of our VREML over MLE and INLA.

stat.ML

Local Variable and Neighborhood Selection for Firearm Fatality in the Southeast USA

A major public health concern in the United States (US) is gun-related deaths. The number of gun injuries largely varies spatially because of county-wise heterogeneity of race, sex, age, and income distributions. But still, a major challenge is to locally identify the influential socio-economic factors behind these firearm fatality incidents. For a diverging number of predictors, a rich literature exists regarding SCAD under the independence framework; however, a vacuum remains when discussing local variable selection for spatially correlated, over-dispersed data. This research presents a two-step localized variable selection and inference framework for spatially indexed gunshot fatality data. In the first step, we select variables locally using the SCAD penalty for specific locations where the number of gunshot incidents exceeds a threshold. For these locations, after selecting the predictors, we proceed to the next step, which involves examining the directional variation in the latent spatial neighborhood structure. We further discuss the theoretical properties of this county-specific local variable selection under infill asymptotics. This method has threefold advantages: (i) this method selects the variables locally, (ii) this method provides inference about directional variation of a selected predictor, and (iii) instead of assuming the spatial neighborhood structure in an ad hoc manner, this method identifies the specific type of spatial neighborhood structure that is most appropriate for modeling the random effects.

stat.ME

Multi-Resolution Analysis of Variable Selection for Road Safety in St. Louis and Its Neighboring Area

Generally, Lasso, Adaptive Lasso, and SCAD are standard approaches in variable selection in the presence of a large number of predictors. In recent years, during intensity function estimation for spatial point processes with a diverging number of predictors, many researchers have considered these penalized methods. But we have discussed a multi-resolution perspective for the variable selection method for spatial point process data. Its advantage is twofold: it not only efficiently selects the predictors but also provides the idea of which points are liable for selecting a predictor at a specific resolution. Actually, our research is motivated by the crime and accident occurrences in St. Louis and its neighborhoods. It is more relevant to select predictors at the local level, and thus we get the idea of which set of predictors is relevant for the occurrences of crime or accident in which parts of St. Louis. We describe the simulation results to justify the accuracy of local-level variable selection during intensity function estimation.

stat.ME

A Subsampling Based Neural Network for Spatial Data

The application of deep neural networks in geospatial data has become a trending research problem in the present day. A significant amount of statistical research has already been introduced, such as generalized least square optimization by incorporating spatial variance-covariance matrix, considering basis functions in the input nodes of the neural networks, and so on. However, for lattice data, there is no available literature about the utilization of asymptotic analysis of neural networks in regression for spatial data. This article proposes a consistent localized two-layer deep neural network-based regression for spatial data. We have proved the consistency of this deep neural network for bounded and unbounded spatial domains under a fixed sampling design of mixed-increasing spatial regions. We have proved that its asymptotic convergence rate is faster than that of \cite{zhan2024neural}'s neural network and an improved generalization of \cite{shen2023asymptotic}'s neural network structure. We empirically observe the rate of convergence of discrepancy measures between the empirical probability distribution of observed and predicted data, which will become faster for a less smooth spatial surface. We have applied our asymptotic analysis of deep neural networks to the estimation of the monthly average temperature of major cities in the USA from its satellite image. This application is an effective showcase of non-linear spatial regression. We demonstrate our methodology with simulated lattice data in various scenarios.

stat.ML

Geo-Spatial Cluster based Hybrid Spatio-Temporal Copula Interpolation

In the absence of Gaussianity assumptions without disturbing spatial continuity interpolating along the whole spatial surface for different time lags is challenging. The past researchers pay enough attention to Spatio-temporal interpolation ignoring the dynamic behavior of a spatial mean function, threshold distance, and direction of maintaining spatial continuity. Therefore, we employ hierarchical spatial clustering (HSC) to preserve local spatial stationarity. This research work introduces a hybrid extreme valued copula-based Spatio-temporal interpolation algorithm. Spatial dependence is captured by a blended extreme valued probability distribution (BEVD). Temporal dependency is modeled by the Bi-directional long short-time memory (BLSTM) at different temporal granularities, 1 month, 2 months, and 3 months. Spatio-temporal dependence is modeled by the Gumbel-Hougaard copula (GH). We apply the proposed Spatio-temporal interpolation approach to the air pollution data (Outdoor Particulate Matter (PM) concentration) of Delhi, collected from the website of the Central Pollution Control Board, India as a crucial circumstantial study. This article describes a probabilistic-recurrent neural networking algorithm for Spatio-temporal interpolation. This Spatio-temporal hybrid copula interpolation algorithm outperforms and is efficient enough to detect spatial trends and temporal influence. From the entire research, we notice that PM concentration in a year reaches a maximum, generally in November and December. The northern and central part of Del-hi is the most sensitive regarding air pollution.

stat.AP

Spatial Cluster-based Copula Model to Interpolate Skewed Conditional Spatial Random Field

Interpolating a skewed conditional spatial random field with missing data is cumbersome in the absence of Gaussianity assumptions. Maintaining spatial homogeneity and continuity around the observed random spatial point is also challenging, especially when interpolating along a spatial surface, focusing on the boundary points as a neighborhood. Otherwise, the point far away from one may appear the closest to another. As a result, importing the hierarchical clustering concept on the spatial random field is as convenient as developing the copula with the interface of the Expectation-Maximization algorithm and concurrently utilizing the idea of the Bayesian framework. This paper introduces a spatial cluster-based C-vine copula and a modified Gaussian kernel to derive a novel spatial probability distribution. Another investigation in this paper uses an algorithm in conjunction with a different parameter estimation technique to make spatial-based copula interpolation more compatible and efficient. We apply the proposed spatial interpolation approach to the air pollution of Delhi as a crucial circumstantial study to demonstrate this newly developed novel spatial estimation technique.

stat.ME

Statistical assessment of spatio-temporal impact of lockdown on air pollution using different modelling approaches in India

One of the main contributors to air pollution is particulate matter (PMxy), which causes several COVID-19 related diseases such as respiratory problems and cardiovascular disorders. Therefore, the spatial and temporal trend analysis of particulate matter and the mass concentration of all aerosol particles less than 2.5 m in diameter (PM2.5) has become critical to control the risk factors of co-morbidity of a patient. Lockdown plays a significant role in maintaining COVID-19 cases as well as air pollution, including particulate matter. This study aims to analyse the effect of the lockdown on controlling air pollution in metropolitan cities in India through various statistical modelling approaches. Most research articles in the literature assume a linear relationship between responses and covariates and take independent and identically distributed error terms in the model, which may not be appropriate for analysing such air pollution data. In this study, we performed a pattern analysis of daily PM2.5 emissions in various major activity zones during 2019 and 2020. By measuring the lockdown effect, we also considered seasonal influence.

stat.AP