SearcharxivSearch

arXiv subjects

Andrew O. Finley

Publications and source records attributed to Andrew O. Finley.

At least 19 recordsLinked to original sources

stLMM: Bayesian Spatial and Space-Time Linear Mixed Models for Small-Area Ecological Estimation

stLMM is an R package for Bayesian linear mixed models with spatial, temporal, and space-time latent effects. It provides a common formula interface for independent and identically distributed (iid) grouped effects, autoregressive (AR) temporal effects, Gaussian process (GP) and nearest-neighbor Gaussian process (NNGP) point-referenced effects, conditional autoregressive (CAR) and directed acyclic graph autoregressive (DAGAR) areal effects, separable areal space-time effects, and structured varying coefficients. The package is designed for ecological small-area estimation workflows in which analysts must move between direct-estimate and unit-level models, combine sampling variances or residual-variance models with spatial and temporal borrowing, and retain missing response rows as prediction targets. A shared sparse-precision implementation underlies the model terms. Structured latent effects are collapsed during fitting, then recovered or retained for fitted values, diagnostics, prediction, and posterior summaries. This gives users one posterior-draw workflow for propagating uncertainty from model fitting through prediction and aggregation. The package is demonstrated with a Washington county biomass example from the package article series, using Forest Inventory and Analysis (FIA) data and tree canopy cover to estimate county-year biomass means. The full article series provides reproducible source code, data, diagnostics, and related model variants.

stat.ME

Clustering the Nearest Neighbor Gaussian Process

Gaussian processes are ubiquitous as the primary tool for modeling spatial data. However, the Gaussian process is limited by its $\mathcal{O}(n^3)$ cost, making direct parameter fitting algorithms infeasible for the scale of modern data collection initiatives. The Nearest Neighbor Gaussian Process (NNGP) was introduced as a scalable approximation to dense Gaussian processes which has been successful for $n\sim 10^6$ observations. This project introduces the $\textit{clustered Nearest Neighbor Gaussian Process}$ (cNNGP) which reduces the computational and storage cost of the NNGP. The accuracy of parameter estimation and reduction in computational and memory storage requirements are demonstrated with simulated data, where the cNNGP provided comparable inference to that obtained with the NNGP, in a fraction of the sampling time. To showcase the method's performance, we modeled biomass over the state of Maine using data collected by the Global Ecosystem Dynamics Investigation (GEDI) to generate wall-to-wall predictions over the state. In 16% of the time, the cNNGP produced nearly indistinguishable inference and biomass prediction maps to those obtained with the NNGP.

stat.ME

Hierarchical models for small area estimation using zero-inflated forest inventory variables: comparison and implementation

National Forest Inventory (NFI) data are typically limited to sparse networks of sample locations due to cost constraints. While design-based estimators provide reliable forest parameter estimates for large areas, there is increasing interest in model-based small area estimation (SAE) methods to improve precision for smaller spatial, temporal, or biophysical domains. SAE methods can be broadly categorized into area- and unit-level models, with unit-level models offering greater flexibility, making them the focus of this study. Ensuring valid inference requires satisfying model distributional assumptions, which is particularly challenging for NFI variables that exhibit positive support and zero-inflation, such as forest biomass, carbon, and volume. Here, we evaluate nine candidate estimators, including two-stage unit-level hierarchical Bayesian models, single-stage Bayesian models, and two-stage frequentist models, for estimating forest biomass at the county level in Nevada and Washington, United States. Estimator performance is assessed using repeated sampling from simulated populations and unit-level cross-validation with FIA data. Results show that small area estimators incorporating a two-stage approach to account for zero-inflation, county-specific random intercepts and residual variances, and spatial random effects yield the most accurate and well-calibrated county-level estimates, with spatial effects providing the greatest benefits when spatial autocorrelation is present in the underlying population.

stat.AP

Small area estimation of growing stock timber volume, basal area, mean stem diameter, and stem density for mountain forests in Austria

Regression models were evaluated to estimate stand-level growing stock volume (GSV), quadratic mean diameter (QMD), basal area (BA), and stem density (N) in the Brixen im Thale forest district of Austria. Field measurements for GSV, QMD, and BA were collected on 146 inventory plots using a handheld mobile personal laser scanning system. Predictor variables were derived from airborne laser scanning (ALS)-derived normalized digital surface and terrain models. The objective was to generate stand-level estimates and associated uncertainty for GSV, QMD, BA, and N across 824 stands. A unit-level small area estimation framework was used to generate stand-level posterior predictive distributions by aggregating predictions from finer spatial scales. Both univariate and multivariate models, with and without spatially varying intercepts, were considered. Predictive performance was assessed via spatially blocked cross-validation, focusing on bias, accuracy, and precision. Despite exploratory analysis suggesting advantages of complex multivariate spatial models, simpler univariate spatial -- and in some cases, non-spatial -- models exhibited comparable predictive performance.

stat.AP

Multivariate spatial models for small area estimation of species-specific forest inventory parameters

National Forest Inventories (NFIs) provide statistically reliable information on forest resources at national and other large spatial scales. As forest management and conservation needs become increasingly complex, NFIs are being called upon to provide forest parameter estimates at spatial scales smaller than current design-based estimation procedures can provide. This is particularly true when estimates are desired by species or species groups. Here we propose a multivariate spatial model for small area estimation of species-specific forest inventory parameters. The hierarchical Bayesian modeling framework accounts for key complexities in species-specific forest inventory data, such as zero-inflation, correlations among species, and residual spatial autocorrelation. Importantly, by fitting the model directly to the individual plot-level data, the framework enables estimates of species-level forest parameters, with associated uncertainty, across any user-defined small area of interest. A simulation study revealed minimal bias and higher accuracy of the proposed model-based approach compared to design-based estimator. We applied the model to estimate species-specific county-level aboveground biomass for the 20 most abundant tree species in the southern United States using Forest Inventory and Analysis (FIA) data. Model-based biomass estimates had high correlations with design-based estimates, yet the model-based estimates tended to have a slight positive bias relative to design-based estimates. Importantly, the proposed model provided large gains in precision across all 20 species. On average across species, 91.5% of county-level biomass estimates had higher precision compared to the design-based estimates. The proposed framework improves the ability of NFI data users to generate species-level forest parameter estimates with reasonable precision at management-relevant spatial scales.

stat.AP

Spatial-temporal prediction of forest attributes using latent Gaussian models and inventory data

The USDA Forest Inventory and Analysis (FIA) program conducts a national forest inventory for the United States through a network of permanent field plots. FIA produces estimates of area averages and totals for plot-measured forest variables through design-based inference, assuming a fixed population and a probability sample of field plot locations. The fixed-population assumption and characteristics of the FIA sampling scheme make it difficult to estimate change in forest variables over time using design-based inference. We propose spatial-temporal models based on Gaussian processes as a flexible tool for forest inventory data, capable of inferring forest variables and change thereof over arbitrary spatial and temporal domains. It is shown to be beneficial for the covariance function governing the latent Gaussian process to account for variation at multiple scales, separating spatially local variation from ecosystem-scale variation. We demonstrate a model for forest biomass density, inferring 20 years of biomass change within two US National Forests.

stat.AP

Gridding and Parameter Expansion for Scalable Latent Gaussian Models of Spatial Multivariate Data

Scalable spatial GPs for massive datasets can be built via sparse Directed Acyclic Graphs (DAGs) where a small number of directed edges is sufficient to flexibly characterize spatial dependence. The DAG can be used to devise fast algorithms for posterior sampling of the latent process, but these may exhibit pathological behavior in estimating covariance parameters. In this article, we introduce gridding and parameter expansion methods to improve the practical performance of MCMC algorithms in terms of effective sample size per unit time (ESS/s). Gridding is a model-based strategy that reduces the number of expensive operations necessary during MCMC on irregularly spaced data. Parameter expansion reduces dependence in posterior samples in spatial regression for high resolution data. These two strategies lead to computational gains in the big data settings on which we focus. We consider popular constructions of univariate spatial processes based on Matérn covariance functions and multivariate coregionalization models for Gaussian outcomes in extensive analyses of synthetic datasets comparing with alternative methods. We demonstrate effectiveness of our proposed methods in a forestry application using remotely sensed data from NASA's Goddard LiDAR, Hyper-Spectral, and Thermal imager (G-LiHT).

stat.ME

Leveraging national forest inventory data to estimate forest carbon density status and trends for small areas

National forest inventory (NFI) data are often costly to collect, which inhibits efforts to estimate parameters of interest for small spatial, temporal, or biophysical domains. Traditionally, design-based estimators are used to estimate status of forest parameters of interest, but are unreliable for small areas where data are sparse. Additionally, design-based estimates constructed directly from the survey data are often unavailable when sample sizes are especially small. Traditional model-based small area estimation approaches, such as the Fay-Herriot (FH) model, rely on these direct estimates for inference; hence, missing direct estimates preclude the use of such approaches. Here, we detail a Bayesian spatio-temporal small area estimation model that efficiently leverages sparse NFI data to estimate status and trends for forest parameters. The proposed model bypasses the use of direct estimates and instead uses plot-level NFI measurements along with auxiliary data including remotely sensed tree canopy cover. We produce forest carbon estimates from the United States NFI over 14 years across the contiguous US (CONUS) and conduct a simulation study to assess our proposed model's accuracy, precision, and bias, compared to that of a design-based estimator. The proposed model provides improved precision and accuracy over traditional estimation methods, and provides useful insights into county-level forest carbon dynamics across the CONUS.

stat.AP

Bayesian Modeling of Incompatible Spatial Data: A Case Study Involving Post-Adrian Storm Forest Damage Assessment

Modeling incompatible spatial data, i.e., data with different spatial resolutions, is a pervasive challenge in remote sensing data analysis. Typical approaches to addressing this challenge aggregate information to a common coarse resolution, i.e., compatible resolutions, prior to modeling. Such pre-processing aggregation simplifies analysis, but potentially causes information loss and hence compromised inference and predictive performance. To avoid losing potential information provided by finer spatial resolution data and improve predictive performance, we propose a new Bayesian method that constructs a latent spatial process model at the finest spatial resolution. This model is tailored to settings where the outcome variable is measured on a coarser spatial resolution than predictor variables -- a configuration seen increasingly when high spatial resolution remotely sensed predictors are used in analysis. A key contribution of this work is an efficient algorithm that enables full Bayesian inference using finer resolution data while optimizing computational and storage costs. The proposed method is applied to a forest damage assessment for the 2018 Adrian storm in Carinthia, Austria, that uses high-resolution laser imaging detection and ranging (LiDAR) measurements and relatively coarse resolution forest inventory measurements. Extensive simulation studies demonstrate the proposed approach substantially improves inference for small prediction units.

stat.ME

Spatio-temporal areal models to support small area estimation: An application to national-scale forest carbon monitoring

National Forest Inventory (NFI) programs can provide vital information on the status, trend, and change in forest parameters. These programs are being increasingly asked to provide forest parameter estimates for spatial and temporal extents smaller than their current design and accompanying design-based methods can deliver with desired levels of uncertainty. Many NFI designs and estimation methods focus on status and are not well equipped to provide acceptable estimates for trend and change parameters, especially over small spatial domains and/or short time periods. Fine-scale space-time indexed estimates are critical to a variety of environmental, ecological, and economic monitoring efforts. Estimates for forest carbon status, trend, and change are of particular importance to international initiatives to track carbon dynamics. Model-based small area estimation (SAE) methods for NFI and similar ecological monitoring data typically pursue inference on status within small spatial domains, with few demonstrated methods that account for spatio-temporal dependence needed for trend and change estimation. We propose a spatio-temporal Bayesian model framework that delivers statistically valid estimates with full uncertainty quantification for status, trend, and change. The framework accommodates a variety of space and time dependency structures, and we detail model configurations for different settings. Through analysis of simulated datasets, we compare the relative performance of candidate models and a traditional direct estimator. We then apply candidate models to a large-scale NFI dataset to demonstrate the utility of the proposed framework for providing unique quantification of forest carbon dynamics in the contiguous United States. We also provide computationally efficient algorithms, software, and data to reproduce our results and for benchmarking.

stat.AP

Calibrating satellite maps with field data for improved predictions of forest biomass

Spatially explicit quantification of forest biomass is important for forest-health monitoring and carbon accounting. Direct field measurements of biomass are laborious and expensive, typically limiting their spatial and temporal sampling density and therefore the precision and resolution of the resulting inference. Satellites can provide biomass predictions at a far greater density, but these predictions are often biased relative to field measurements and exhibit heterogeneous errors. We developed and implemented a coregionalization model between sparse field measurements and a predictive satellite map to deliver improved predictions of biomass density at a 1-by-1 km resolution throughout the Pacific states of California, Oregon and Washington. The model accounts for zero-inflation in the field measurements and the heterogeneous errors in the satellite predictions. A stochastic partial differential equation approach to spatial modeling is applied to handle the magnitude of the satellite data. The spatial detail rendered by the model is much finer than would be possible with the field measurements alone, and the model provides substantial noise-filtering and bias-correction to the satellite map.

stat.AP

spAbundance: An R package for single-species and multi-species spatially explicit abundance models

Numerous modeling techniques exist to estimate abundance of plant and wildlife species. These methods seek to estimate abundance while accounting for multiple complexities found in ecological data, such as observational biases, spatial autocorrelation, and species correlations. There is, however, a lack of user-friendly and computationally efficient software to implement the various models, particularly for large data sets. We developed the spAbundance R package for fitting spatially-explicit Bayesian single-species and multi-species hierarchical distance sampling models, N-mixture models, and generalized linear mixed models. The models within the package can account for spatial autocorrelation using Nearest Neighbor Gaussian Processes and accommodate species correlations in multi-species models using a latent factor approach, which enables model fitting for data sets with large numbers of sites and/or species. We provide three vignettes and three case studies that highlight spAbundance functionality. We used spatially-explicit multi-species distance sampling models to estimate density of 16 bird species in Florida, USA, an N-mixture model to estimate Black-throated Blue Warbler (Setophaga caerulescens) abundance in New Hampshire, USA, and a spatial linear mixed model to estimate forest aboveground biomass across the continental USA. spAbundance provides a user-friendly, formula-based interface to fit a variety of univariate and multivariate spatially-explicit abundance models. The package serves as a useful tool for ecologists and conservation practitioners to generate improved inference and predictions on the spatial drivers of populations and communities.

stat.AP

Model-assisted estimation of domain totals, areas, and densities in two-stage sample survey designs

Model-assisted, two-stage forest survey sampling designs provide a means to combine airborne remote sensing data, collected in a sampling mode, with field plot data to increase the precision of national forest inventory estimates, while maintaining important properties of design-based inventories, such as unbiased estimation and quantification of uncertainty. In this study, we present a comprehensive set of model-assisted estimators for domain-level attributes in a two-stage sampling design, including new estimators for densities, and compare the performance of these estimators with standard poststratified estimators. Simulation was used to assess the statistical properties (bias, variability) of these estimators, with both simple random and systematic sampling configurations, and indicated that 1) all estimators were generally unbiased. and 2) the use of lidar in a sampling mode increased the precision of the estimators at all assessed field sampling intensities, with particularly marked increases in precision at lower field sampling intensities. Variance estimators are generally unbiased for model-assisted estimators without poststratification, while model-assisted estimators with poststratification were increasingly biased as field sampling intensity decreased. In general, these results indicate that airborne remote sensing, collected in a sampling mode, can be used to increase the efficiency of national forest inventories.

stat.AP

Models to support forest inventory and small area estimation using sparsely sampled LiDAR: A case study involving G-LiHT LiDAR in Tanana, Alaska

A two-stage hierarchical Bayesian model is developed and implemented to estimate forest biomass density and total given sparsely sampled LiDAR and georeferenced forest inventory plot measurements. The model is motivated by the United States Department of Agriculture (USDA) Forest Service Forest Inventory and Analysis (FIA) objective to provide biomass estimates for the remote Tanana Inventory Unit (TIU) in interior Alaska. The proposed model yields stratum-level biomass estimates for arbitrarily sized areas. Model-based estimates are compared with the TIU FIA design-based post-stratified estimates. Model-based small area estimates (SAEs) for two experimental forests within the TIU are compared with each forest's design-based estimates generated using a dense network of independent inventory plots. Model parameter estimates and biomass predictions are informed using FIA plot measurements, LiDAR data that are spatially aligned with a subset of the FIA plots, and complete coverage remotely detected data used to define landuse/landcover stratum and percent forest canopy cover. Results support a model-based approach to estimating forest variables when inventory data are sparse or resources limit collection of enough data to achieve desired accuracy and precision using design-based methods.

stat.AP

A spatial mixture model for spaceborne lidar observations over mixed forest and non-forest land types

The Global Ecosystem Dynamics Investigation (GEDI) is a spaceborne lidar instrument that collects near-global measurements of forest structure. While expansive in scope, GEDI samples are spatially sparse and cover a small fraction of the land surface. Converting the sparse samples into spatially complete predictive maps is of practical importance for many ecological studies. A complicating factor is that GEDI collects measurements over forested and non-forested land alike, with no automatic labeling of the land type. Such classification is important, as it categorically influences the probability distribution of the spatial process and the ecological interpretation of the observations/predictions. We implement a spatial mixture model, separating the spatial domain into two latent classes. The latent classes are governed by a Bernoulli spatial process and within each class the process is governed by a separate spatial model. Model predictions take the form of scalar predictions as well as discrete labeling of the class membership. Inference is conducted through a Bayesian paradigm, yielding rich quantification of prediction and uncertainty. We demonstrate the method using GEDI data over Wollemi National Park. When compared to a single spatial model, the mixture model achieves much higher posterior predictive densities on the true value. When compared to a random forest model, a common algorithmic approach in the remote sensing community, the random forest achieves better absolute prediction accuracy for prediction locations far from observed training data locations, but at the expense of location-specific assessments of uncertainty. The unsupervised binary classifications of the mixture model appear broadly ecologically interpretable as forest and non-forest when compared to optical imagery, but further comparison to ground-truth data is required.

stat.AP

Spatial prediction of diameter distributions for the alpine protection forests in Ebensee, Austria, using ALS/PLS and spatial distributional regression models

A spatial distributional regression model is presented to predict the forest structural diversity in terms of the distributions of the stem diameter at breast height (DBH) in the protection forests in Ebensee, Austria. In total 36,338 sample trees were measured via a handheld mobile personal laser scanning system (PLS) on 273 sample plots each having a 20 m radius. Recent airborne laser scanning (ALS) data was used to derive regression covariates from the normalized digital vegetation height model (DVHM) and the digital terrain model (DTM). Candidate models were constructed that differed in their linear predictors of the two gamma distribution parameters. Non-linear smoothing splines outperformed linear parametric slope coefficients, and the best implementation of spatial structured effects was achieved by a Gaussian process smooth. Model fitting and posterior parameter inference was achieved by using full Bayesian methodology and MCMC sampling algorithms implemented in the R-package BAMLSS. Spatial predictions of stem count proportions per DBH classes revealed that regeneration of smaller trees was lacking in certain areas of the protection forest landscape.

stat.AP

Quantifying and correcting geolocation error in spaceborne LiDAR forest canopy observations using high spatial accuracy ALS: A Bayesian model approach

Geolocation error in spaceborne sampling light detection and ranging (LiDAR) measurements of forest structure can compromise forest attribute estimates and degrade integration with georeferenced field measurements or other remotely sensed data. Data integration is especially problematic when geolocation error is not well quantified. We propose a general model that uses airborne laser scanning (ALS) data to quantify and correct geolocation error in spaceborne sampling LiDAR. To illustrate the model, LiDAR data from NASA Goddard's LiDAR Hyperspectral & Thermal Imager (G-LiHT) was used with a subset of LiDAR data from NASA's Global Ecosystem Dynamics Investigation (GEDI). The model accommodates multiple canopy height metrics derived from a simulated GEDI footprint kernel using spatially coincident G-LiHT, and incorporates both additive and multiplicative mapping between the canopy height metrics generated from both datasets. A Bayesian implementation provides probabilistic uncertainty quantification in both parameter and geolocation error estimates. Results show a systematic geolocation error of 9.62 m in the southwest direction. In addition, estimated geolocation errors within GEDI footprints were highly variable, with results showing a ~0.45 probability the true footprint center is within 20 m. Estimating and correcting geolocation error via the model outlined here can help inform subsequent efforts to integrate spaceborne LiDAR data, like GEDI, with other georeferenced data.

stat.AP

Modeling complex species-environment relationships through spatially-varying coefficient occupancy models

Occupancy models are frequently used by ecologists to quantify spatial variation in species distributions while accounting for observational biases in the collection of detection-nondetection data. However, the common assumption that a single set of regression coefficients can adequately explain species-environment relationships is often unrealistic, especially across large spatial domains. Here we develop single-species (i.e., univariate) and multi-species (i.e., multivariate) spatially-varying coefficient (SVC) occupancy models to account for spatially-varying species-environment relationships. We employ Nearest Neighbor Gaussian Processes and Polya-Gamma data augmentation in a hierarchical Bayesian framework to yield computationally efficient Gibbs samplers, which we implement in the spOccupancy R package. For multi-species models, we use spatial factor dimension reduction to efficiently model datasets with large numbers of species (e.g., > 10). The hierarchical Bayesian framework readily enables generation of posterior predictive maps of the SVCs, with fully propagated uncertainty. We apply our SVC models to quantify spatial variability in the relationships between maximum breeding season temperature and occurrence probability of 21 grassland bird species across the U.S. Jointly modeling species generally outperformed single-species models, which all revealed substantial spatial variability in species occurrence relationships with maximum temperatures. Our models are particularly relevant for quantifying species-environment relationships using detection-nondetection data from large-scale monitoring programs, which are becoming increasingly prevalent for answering macroscale ecological questions regarding wildlife responses to global change.

stat.AP