SearcharxivSearch

arXiv subjects

Indrila Ganguly

Publications and source records attributed to Indrila Ganguly.

4 recordsLinked to original sources

Short-term forecasting of wildfire spread: A network epidemiology approach

Wildfire spread poses substantial environmental and public-health risks, motivating interpretable models for short-term forecasting. We develop a statistical framework that combines cellular automata with ideas from network epidemiology to model wildfire evolution across a spatial lattice. Each grid cell is classified as available, burning, or consumed. State transitions distinguish spread from burning neighbors, intrinsic ignition, and cessation of burning, with transition rates linked to meteorological and environmental covariates. A likelihood-based estimation procedure yields transition-specific covariate effects and probabilistic forecasts of cell states. We assess forecasting performance in a simulation study and in applications to the 2018 California wildfires and the 2019-2020 Australian wildfires. We also compare the method with a published forecasting approach using the 2017 Haypress fire. The results show strong short-term discrimination in many settings, with reduced accuracy at longer forecast horizons and during abrupt fire expansion. The framework provides an interpretable basis for studying wildfire dynamics and identifies opportunities to improve ignition forecasts through richer spatial and observation models.

stat.AP

Sequential Testing for Assessing the Incremental Value of Biomarkers Under Biorepository Specimen Constraints with Robustness to Model Misspecification

In cancer biomarker development, a key objective is to evaluate whether a new biomarker, when combined with an established one, improves early cancer detection compared to using the established biomarker alone. Incremental value is often quantified by changes at specific points on the ROC curve, such as an increase in sensitivity at a fixed specificity, which is especially relevant in early cancer detection. Our research is motivated by the Early Detection Research Network (EDRN) biorepository studies, which aim to validate multiple cancer biomarkers across laboratories using specimens obtained from a centralized biorepository, under the constraint of limited specimen availability. To address this challenge, we propose a two-stage group sequential hypothesis testing framework for assessing incremental effects, allowing early stopping for futility or efficacy to conserve valuable samples. Our asymptotic results are derived under a logistic working model and remain valid even under model misspecification, ensuring robustness and broad applicability. We further integrate a rotating group membership design to facilitate validation of multiple candidate biomarkers across laboratories. Through extensive simulations, we demonstrate valid type I error control and efficient utilization of specimens. Finally, we apply our method to data from a multicenter EDRN pancreatic cancer reference set study and show how the proposed approach identifies promising candidate biomarkers that provide incremental performance when combined with CA19-9, while enabling efficient evaluation of a large number of such candidates.

stat.ME

Scalable Resampling in Massive Generalized Linear Models via Subsampled Residual Bootstrap

Residual bootstrap is a classical method for statistical inference in regression settings. With massive data sets becoming increasingly common, there is a demand for computationally efficient alternatives to residual bootstrap. We propose a simple and versatile scalable algorithm called subsampled residual bootstrap (SRB) for generalized linear models (GLMs), a large class of regression models that includes the classical linear regression model as well as other widely used models such as logistic, Poisson and probit regression. We prove consistency and distributional results that establish that the SRB has the same theoretical guarantees under the GLM framework as the classical residual bootstrap, while being computationally much faster. We demonstrate the empirical performance of SRB via simulation studies and a real data analysis of the Forest Covertype data from the UCI Machine Learning Repository.

stat.ME

Robust Estimation of Average Treatment Effects from Panel Data

In order to evaluate the impact of a policy intervention on a group of units over time, it is important to correctly estimate the average treatment effect (ATE) measure. Due to lack of robustness of the existing procedures of estimating ATE from panel data, in this paper, we introduce a robust estimator of the ATE and the subsequent inference procedures using the popular approach of minimum density power divergence inference. Asymptotic properties of the proposed ATE estimator are derived and used to construct robust test statistics for testing parametric hypotheses related to the ATE. Besides asymptotic analyses of efficiency and powers, extensive simulation studies are conducted to study the finite-sample performances of our proposed estimation and testing procedures under both pure and contaminated data. The robustness of the ATE estimator is further investigated theoretically through the influence functions analyses. Finally our proposal is applied to study the long-term economic effects of the 2004 Indian Ocean earthquake and tsunami on the (per-capita) gross domestic products (GDP) of five mostly affected countries, namely Indonesia, Sri Lanka, Thailand, India and Maldives.

stat.ME