Searcharxiv⌕ Search

arXiv · 2609.36328

Multilevel regression trees with application to wildfires in the American west

Abstract

We propose a Bayesian regression tree model fit within a multilevel structure and apply it to historic wildfire data in the western United States. Sharing of information between related groups (ecoregions) combined with highly interpretable regression trees allows for better predictions and understanding of climate and land cover variables predictive of wildfires. By doing a simulation study with a range of performance metrics, we demonstrate our method produces tree posteriors most structurally similar to assumed true trees, while simultaneously achieving good out-of-sample predictive performance. Applied to a large wildfire data set, we explore variable splits within regression trees corresponding to each ecoregion in detail, taking into account known features of each location. Shared hyperparameters between trees provide highly useful understanding of both variable and split value importance in predicting wildfires among all ecoregions, with no direct parallel in comparable models. Namely, we highlight potential evaporation, temperature, and evergreen forest land cover as variables most associated with historic wildfires, with some observable patterns in split values most commonly chosen across the groups. We propose a new algorithm based on parallel tempering, conditioning on shared hyperparameters at the true posterior temperature, improving Markov chain mixing, a known bottleneck in Bayesian CART models.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

John Henry V. Gray, Tianjian Zhou, Benjamin A. Shaby. 2026-09-28. Multilevel regression trees with application to wildfires in the American west. https://arxiv.org/abs/2609.36328

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Geographic Disparities in Hospice Quality and Family Caregiver Experience: The Roles of Ownership, Social Vulnerability, and Workforce Capacity

Hospice quality should be interpreted in relation to both provider organization and the local conditions under which care is delivered. This study develops a provider-county performance assessment framework by linking national Centers for Medicare & Medicaid Services (CMS) hospice data and Consumer Assessment of Healthcare Providers and Systems (CAHPS) Hospice Survey outcomes with county measures of rurality, social vulnerability, health burden, and workforce and health-resource context. The adjusted analysis includes 2,928 providers in 1,078 counties and combines geographic mapping, blockwise regression, six secondary CAHPS outcomes, and eight sensitivity analyses. Adding county context increased adjusted R-squared from 0.077 to 0.202. After full adjustment, for-profit hospices had overall caregiver ratings 3.832 percentage points lower than nonprofit hospices. This negative association appeared across all six secondary CAHPS domains and remained significant in every sensitivity specification. Higher county social vulnerability was also associated with poorer caregiver experience, although its magnitude depended partly on the specification of community health burden. These findings show that county context materially improves hospice performance assessment but does not eliminate the ownership difference. The framework supports context-aware monitoring, peer comparison, and targeted quality improvement.

stat.AP↗

Constrained convex clustering for interpretable spatial domain detection in spot-based spatial transcriptomics

Popular technologies for generating spatially resolved transcriptomic data measure gene expression at the resolution of a "spot", i.e., a small tissue region 55 microns in diameter. Each spot can contain many cells of different types. In typical analyses, researchers are interested in using these data to identify and profile discrete spatial domains in tissue. In this paper, we propose a new method, DUET, which simultaneously identifies discrete spatial domains and estimates each spot's expected cell-type proportion. This allows the identified spatial domains to be characterized in terms of the underlying expected cell-type proportions, which affords interpretability and biological insight. DUET utilizes a constrained version of model-based convex clustering, and as such, can accommodate Poisson, negative binomial, normal, and other types of expression data. Moreover, our convex clustering-type criterion allows for both the number of clusters and degree of spatial smoothness to be controlled by a single tuning parameter, which can be chosen in a data-driven fashion. Through simulation studies and a real data application, we show that DUET can achieve better clustering and deconvolution performance than some existing methods.

stat.AP↗

Extended State-dependent Hawkes Process for Limit Order Books: Mathematical Foundation and the Reproduction of Volatility Signature Plots

This paper proposes an Extended State-Dependent Hawkes Process (ExsdHawkes) to model the intricate dynamics of Limit Order Books (LOBs). Our theoretical contribution lies in relaxing traditional constraints by allowing for state disappearances---a phenomenon frequently observed in high-frequency trading. We mathematically prove, using Karush--Kuhn--Tucker (KKT) conditions, that the maximum likelihood estimation remains separable, justifying an efficient two-step procedure. In the empirical section, we apply our model to three months of high-frequency tick data of Mitsubishi UFJ Financial Group (8306). We demonstrate that ExsdHawkes successfully replicates the characteristic upward slope of the volatility signature plot by capturing the ``local super-criticality'' triggered during disequilibrium states. Crucially, we clarify that the transition out of equilibrium is deterministically triggered by Aggressive Market Orders (AMS/AMB), while Marketable Limit Orders (MLO) function as a critical liquidity-depletion catalyst within the expanded spread. Comparative analysis reveals that models lacking physical constraints (e.g., standard SD-Hawkes) suffer from explosive spectral radii and fail to maintain simulation stability. Our findings suggest that physical consistency is not merely a mathematical nicety, but a prerequisite for accurately modeling macro-level volatility. By enforcing the physical geometry to `pause' the residual accumulation during inadmissible periods, ExsdHawkes maintains statistical integrity where unconstrained models succumb to structural bias and simulation instability.

stat.AP↗