SearcharxivSearch

arXiv subjects

Bruce D. Cook

Publications and source records attributed to Bruce D. Cook.

8 recordsLinked to original sources

Model-assisted estimation of domain totals, areas, and densities in two-stage sample survey designs

Model-assisted, two-stage forest survey sampling designs provide a means to combine airborne remote sensing data, collected in a sampling mode, with field plot data to increase the precision of national forest inventory estimates, while maintaining important properties of design-based inventories, such as unbiased estimation and quantification of uncertainty. In this study, we present a comprehensive set of model-assisted estimators for domain-level attributes in a two-stage sampling design, including new estimators for densities, and compare the performance of these estimators with standard poststratified estimators. Simulation was used to assess the statistical properties (bias, variability) of these estimators, with both simple random and systematic sampling configurations, and indicated that 1) all estimators were generally unbiased. and 2) the use of lidar in a sampling mode increased the precision of the estimators at all assessed field sampling intensities, with particularly marked increases in precision at lower field sampling intensities. Variance estimators are generally unbiased for model-assisted estimators without poststratification, while model-assisted estimators with poststratification were increasingly biased as field sampling intensity decreased. In general, these results indicate that airborne remote sensing, collected in a sampling mode, can be used to increase the efficiency of national forest inventories.

stat.AP

Models to support forest inventory and small area estimation using sparsely sampled LiDAR: A case study involving G-LiHT LiDAR in Tanana, Alaska

A two-stage hierarchical Bayesian model is developed and implemented to estimate forest biomass density and total given sparsely sampled LiDAR and georeferenced forest inventory plot measurements. The model is motivated by the United States Department of Agriculture (USDA) Forest Service Forest Inventory and Analysis (FIA) objective to provide biomass estimates for the remote Tanana Inventory Unit (TIU) in interior Alaska. The proposed model yields stratum-level biomass estimates for arbitrarily sized areas. Model-based estimates are compared with the TIU FIA design-based post-stratified estimates. Model-based small area estimates (SAEs) for two experimental forests within the TIU are compared with each forest's design-based estimates generated using a dense network of independent inventory plots. Model parameter estimates and biomass predictions are informed using FIA plot measurements, LiDAR data that are spatially aligned with a subset of the FIA plots, and complete coverage remotely detected data used to define landuse/landcover stratum and percent forest canopy cover. Results support a model-based approach to estimating forest variables when inventory data are sparse or resources limit collection of enough data to achieve desired accuracy and precision using design-based methods.

stat.AP

Quantifying and correcting geolocation error in spaceborne LiDAR forest canopy observations using high spatial accuracy ALS: A Bayesian model approach

Geolocation error in spaceborne sampling light detection and ranging (LiDAR) measurements of forest structure can compromise forest attribute estimates and degrade integration with georeferenced field measurements or other remotely sensed data. Data integration is especially problematic when geolocation error is not well quantified. We propose a general model that uses airborne laser scanning (ALS) data to quantify and correct geolocation error in spaceborne sampling LiDAR. To illustrate the model, LiDAR data from NASA Goddard's LiDAR Hyperspectral & Thermal Imager (G-LiHT) was used with a subset of LiDAR data from NASA's Global Ecosystem Dynamics Investigation (GEDI). The model accommodates multiple canopy height metrics derived from a simulated GEDI footprint kernel using spatially coincident G-LiHT, and incorporates both additive and multiplicative mapping between the canopy height metrics generated from both datasets. A Bayesian implementation provides probabilistic uncertainty quantification in both parameter and geolocation error estimates. Results show a systematic geolocation error of 9.62 m in the southwest direction. In addition, estimated geolocation errors within GEDI footprints were highly variable, with results showing a ~0.45 probability the true footprint center is within 20 m. Estimating and correcting geolocation error via the model outlined here can help inform subsequent efforts to integrate spaceborne LiDAR data, like GEDI, with other georeferenced data.

stat.AP

Conjugate Nearest Neighbor Gaussian Process Models for Efficient Statistical Interpolation of Large Spatial Data

A key challenge in spatial statistics is the analysis for massive spatially-referenced data sets. Such analyses often proceed from Gaussian process specifications that can produce rich and robust inference, but involve dense covariance matrices that lack computationally exploitable structures. The matrix computations required for fitting such models involve floating point operations in cubic order of the number of spatial locations and dynamic memory storage in quadratic order. Recent developments in spatial statistics offer a variety of massively scalable approaches. Bayesian inference and hierarchical models, in particular, have gained popularity due to their richness and flexibility in accommodating spatial processes. Our current contribution is to provide computationally efficient exact algorithms for spatial interpolation of massive data sets using scalable spatial processes. We combine low-rank Gaussian processes with efficient sparse approximations. Following recent work by [1], we model the low-rank process using a Gaussian predictive process (GPP) and the residual process as a sparsity-inducing nearest-neighbor Gaussian process (NNGP). A key contribution here is to implement these models using exact conjugate Bayesian modeling to avoid expensive iterative algorithms. Through the simulation studies, we evaluate performance of the proposed approach and the robustness of our models, especially for long range prediction. We implement our approaches for remotely sensed light detection and ranging (LiDAR) data collected over the US Forest Service Tanana Inventory Unit (TIU) in a remote portion of Interior Alaska.

stat.ME

Spatial Factor Models for High-Dimensional and Large Spatial Data: An Application in Forest Variable Mapping

Gathering information about forest variables is an expensive and arduous activity. As such, directly collecting the data required to produce high-resolution maps over large spatial domains is infeasible. Next generation collection initiatives of remotely sensed Light Detection and Ranging (LiDAR) data are specifically aimed at producing complete-coverage maps over large spatial domains. Given that LiDAR data and forest characteristics are often strongly correlated, it is possible to make use of the former to model, predict, and map forest variables over regions of interest. This entails dealing with the high-dimensional ($\sim$$10^2$) spatially dependent LiDAR outcomes over a large number of locations (~10^5-10^6). With this in mind, we develop the Spatial Factor Nearest Neighbor Gaussian Process (SF-NNGP) model, and embed it in a two-stage approach that connects the spatial structure found in LiDAR signals with forest variables. We provide a simulation experiment that demonstrates inferential and predictive performance of the SF-NNGP, and use the two-stage modeling strategy to generate complete-coverage maps of forest variables with associated uncertainty over a large region of boreal forests in interior Alaska.

stat.AP

Geostatistical estimation of forest biomass in interior Alaska combining Landsat-derived tree cover, sampled airborne lidar and field observations

The goal of this research was to develop and examine the performance of a geostatistical coregionalization modeling approach for combining field inventory measurements, strip samples of airborne lidar and Landsat-based remote sensing data products to predict aboveground biomass (AGB) in interior Alaska's Tanana Valley. The proposed modeling strategy facilitates pixel-level mapping of AGB density predictions across the entire spatial domain. Additionally, the coregionalization framework allows for statistically sound estimation of total AGB for arbitrary areal units within the study area---a key advance to support diverse management objectives in interior Alaska. This research focuses on appropriate characterization of prediction uncertainty in the form of posterior predictive coverage intervals and standard deviations. Using the framework detailed here, it is possible to quantify estimation uncertainty for any spatial extent, ranging from pixel-level predictions of AGB density to estimates of AGB stocks for the full domain. The lidar-informed coregionalization models consistently outperformed their counterpart lidar-free models in terms of point-level predictive performance and total AGB precision. Additionally, the inclusion of Landsat-derived forest cover as a covariate further improved estimation precision in regions with lower lidar sampling intensity. Our findings also demonstrate that model-based approaches that do not explicitly account for residual spatial dependence can grossly underestimate uncertainty, resulting in falsely precise estimates of AGB. On the other hand, in a geostatistical setting, residual spatial structure can be modeled within a Bayesian hierarchical framework to obtain statistically defensible assessments of uncertainty for AGB estimates.

stat.AP

Joint hierarchical models for sparsely sampled high-dimensional LiDAR and forest variables

Recent advancements in remote sensing technology, specifically Light Detection and Ranging (LiDAR) sensors, provide the data needed to quantify forest characteristics at a fine spatial resolution over large geographic domains. From an inferential standpoint, there is interest in prediction and interpolation of the often sparsely sampled and spatially misaligned LiDAR signals and forest variables. We propose a fully process-based Bayesian hierarchical model for above ground biomass (AGB) and LiDAR signals. The process-based framework offers richness in inferential capabilities, e.g., inference on the entire underlying processes instead of estimates only at pre-specified points. Key challenges we obviate include misalignment between the AGB observations and LiDAR signals and the high-dimensionality in the model emerging from LiDAR signals in conjunction with the large number of spatial locations. We offer simulation experiments to evaluate our proposed models and also apply them to a challenging dataset comprising LiDAR and spatially coinciding forest inventory variables collected on the Penobscot Experimental Forest (PEF), Maine. Our key substantive contributions include AGB data products with associated measures of uncertainty for the PEF and, more broadly, a methodology that should find use in a variety of current and upcoming forest variable mapping efforts using sparsely sampled remotely sensed high-dimensional data.

stat.AP

Dynamic spatial regression models for space-varying forest stand tables

Many forest management planning decisions are based on information about the number of trees by species and diameter per unit area. This information is commonly summarized in a stand table, where a stand is defined as a group of forest trees of sufficiently uniform species composition, age, condition, or productivity to be considered a homogeneous unit for planning purposes. Typically information used to construct stand tables is gleaned from observed subsets of the forest selected using a probability-based sampling design. Such sampling campaigns are expensive and hence only a small number of sample units are typically observed. This data paucity means that stand tables can only be estimated for relatively large areal units. Contemporary forest management planning and spatially explicit ecosystem models require stand table input at higher spatial resolution than can be affordably provided using traditional approaches. We propose a dynamic multivariate Poisson spatial regression model that accommodates both spatial correlation between observed diameter distributions and also correlation between tree counts across diameter classes within each location. To improve fit and prediction at unobserved locations, diameter specific intensities can be estimated using auxiliary data such as management history or remotely sensed information. The proposed model is used to analyze a diverse forest inventory dataset collected on the United States Forest Service Penobscot Experimental Forest in Bradley, Maine. Results demonstrate that explicitly modeling the residual spatial structure via a multivariate Gaussian process and incorporating information about forest structure from LiDAR covariates improve model fit and can provide high spatial resolution stand table maps with associated estimates of uncertainty.

stat.AP