SearcharxivSearch

arXiv subjects

Emily L. Kang

Publications and source records attributed to Emily L. Kang.

9 recordsLinked to original sources

A Bayesian Framework for Built-in Input Dimension Reduction for Gaussian Process Modeling

Gaussian process (GP) modeling is widely used in computational science and engineering. However, fitting a GP to high-dimensional inputs remains challenging due to the curse of dimensionality. While various methods have been proposed to reduce input dimensionality, they typically follow a two-stage approach, performing dimension reduction and GP fitting separately. We introduce a Bayesian framework that seamlessly integrates dimensionality reduction with GP modeling and inference. Our approach, built on a hierarchical Bayesian model with priors on the Stiefel manifold, enforces orthonormality on the projection matrix and enables posterior inference via Hamiltonian Monte Carlo with geodesic flow. Additionally, we extend this framework by incorporating Deep Gaussian Processes (DGP) with built-in dimension reduction, providing a more flexible and powerful tool for complex datasets. Through extensive numerical studies, we demonstrate that while the proposed Bayesian method incurs higher computational costs, it improves predictive performance and uncertainty quantification, providing a principled and robust alternative to existing methods.

stat.ML

PPD-CPP: Pointwise predictive density calibrated-power prior in dynamically borrowing historical information

Incorporating historical or real-world data into analyses of treatment effects for rare diseases has become increasingly popular. A major challenge, however, lies in determining the appropriate degree of congruence between historical and current data. In this study, we devote ourselves to the capacity of historical data in replicating the current data, and propose a new congruence measure/estimand $p_{CM}$. $p_{CM}$ quantifies the heterogeneity between two datasets following the idea of the marginal posterior predictive $p$-value, and its asymptotic properties were derived. Building upon $p_{CM}$, we develop the pointwise predictive density calibrated-power prior (PPD-CPP) to dynamically leverage historical information. PPD-CPP achieves the borrowing consistency and allows modeling the power parameter either as a fixed scalar or case-specific quantity informed by covariates. Simulation studies were conducted to demonstrate the performance of these methods and the methodology was illustrated using the Mother's Gift study and \textit{Ceriodaphnia dubia} toxicity test.

stat.ME

Recursive Nearest Neighbor Co-Kriging Models for Big Multiple Fidelity Spatial Data Sets

Big datasets are gathered daily from different remote sensing platforms. Recently, statistical co-kriging models, with the help of scalable techniques, have been able to combine such datasets by using spatially varying bias corrections. The associated Bayesian inference for these models is usually facilitated via Markov chain Monte Carlo (MCMC) methods which present (sometimes prohibitively) slow mixing and convergence because they require the simulation of high-dimensional random effect vectors from their posteriors given large datasets. To enable fast inference in big data spatial problems, we propose the recursive nearest neighbor co-kriging (RNNC) model. Based on this model, we develop two computationally efficient inferential procedures: a) the collapsed RNNC which reduces the posterior sampling space by integrating out the latent processes, and b) the conjugate RNNC, an MCMC free inference which significantly reduces the computational time without sacrificing prediction accuracy. An important highlight of conjugate RNNC is that it enables fast inference in massive multifidelity data sets by avoiding expensive integration algorithms. The efficient computational and good predictive performances of our proposed algorithms are demonstrated on benchmark examples and the analysis of the High-resolution Infrared Radiation Sounder data gathered from two NOAA polar orbiting satellites in which we managed to reduce the computational time from multiple hours to just a few minutes.

stat.CO

Bayesian Latent Variable Co-kriging Model in Remote Sensing for Observations with Quality Flagged

Remote sensing data products often include quality flags that inform users whether the associated observations are of good, acceptable or unreliable qualities. However, such information on data fidelity is not considered in remote sensing data analyses. Motivated by observations from the Atmospheric Infrared Sounder (AIRS) instrument on board NASA's Aqua satellite, we propose a latent variable co-kriging model with separable Gaussian processes to analyze large quality-flagged remote sensing data sets together with their associated quality information. We augment the posterior distribution by an imputation mechanism to decompose large covariance matrices into separate computationally efficient components taking advantage of their input structure. Within the augmented posterior, we develop a Markov chain Monte Carlo (MCMC) procedure that mostly consists of direct simulations from conditional distributions. In addition, we propose a computationally efficient recursive prediction procedure. We apply the proposed method to air temperature data from the AIRS instrument. We show that incorporating quality flag information in our proposed model substantially improves the prediction performance compared to models that do not account for quality flags.

stat.AP

Hierarchical Bayesian Nearest Neighbor Co-Kriging Gaussian Process Models; An Application to Intersatellite Calibration

Recent advancements in remote sensing technology and the increasing size of satellite constellations allows massive geophysical information to be gathered daily on a global scale by numerous platforms of different fidelity. The auto-regressive co-kriging model is a suitable framework to analyse such data sets because it accounts for cross-dependencies among different fidelity satellite outputs. However, its implementation in multifidelity large spatial data-sets is practically infeasible because its computational complexity increases cubically with the total number of observations. In this paper, we propose a nearest neighbour co-kriging Gaussian process that couples the auto-regressive model and nearest neighbour GP by using augmentation ideas; reducing the computational complexity to be linear with the total number of spatial observed locations. The latent process of the nearest neighbour GP is augmented in a manner which allows the specification of semi-conjugate priors. This facilitates the design of an efficient MCMC sampler involving mostly direct sampling updates which can be implemented in parallel computational environments. The good predictive performance of the proposed method is demonstrated in a simulation study. We use the proposed method to analyze High-resolution Infrared Radiation Sounder data gathered from two NOAA polar orbiting satellites.

stat.CO

A Fused Gaussian Process Model for Very Large Spatial Data

With the development of new remote sensing technology, large or even massive spatial datasets covering the globe become available. Statistical analysis of such data is challenging. This article proposes a semiparametric approach to model large or massive spatial datasets. In particular, a Gaussian process with additive components is proposed, with its covariance structure consisting of two components: one component is flexible without assuming a specific parametric covariance function but is able to achieve dimension reduction; the other is parametric and simultaneously induces sparsity. The inference algorithm for parameter estimation and spatial prediction is devised. The resulting spatial prediction methodology that we call fused Gaussian process (FGP), is applied to simulated data and a massive satellite dataset. The results demonstrate the computational and inferential benefits of FGP over competing methods and show that FGP is robust against model misspecification and captures spatial nonstationarity. The supplemental materials are available online.

stat.ME

An Additive Approximate Gaussian Process Model for Large Spatio-Temporal Data

Motivated by a large ground-level ozone dataset, we propose a new computationally efficient additive approximate Gaussian process. The proposed method incorporates a computational-complexity-reduction method and a separable covariance function, which can flexibly capture various spatio-temporal dependence structure. The first component is able to capture nonseparable spatio-temporal variability while the second component captures the separable variation. Based on a hierarchical formulation of the model, we are able to utilize the computational advantages of both components and perform efficient Bayesian inference. To demonstrate the inferential and computational benefits of the proposed method, we carry out extensive simulation studies assuming various scenarios of underlying spatio-temporal covariance structure. The proposed method is also applied to analyze large spatio-temporal measurements of ground-level ozone in the Eastern United States.

stat.ME

Spatio-Temporal Data Fusion for Massive Sea Surface Temperature Data from MODIS and AMSR-E Instruments

Remote sensing data have been widely used to study various geophysical processes. With the advances in remote-sensing technology, massive amount of remote sensing data are collected in space over time. Different satellite instruments typically have different footprints, measurement-error characteristics, and data coverages. To combine datasets from different satellite instruments, we propose a dynamic fused Gaussian process (DFGP) model that enables fast statistical inference such as filtering and smoothing for massive spatio-temporal datasets in a data-fusion context. Based upon a spatio-temporal-random-effects model, the DFGP methodology represents the underlying true process with two components: a linear combination of a small number of basis functions and random coefficients with a general covariance matrix, together with a linear combination of a large number of basis functions and Markov random coefficients. To model the underlying geophysical process at different spatial resolutions, we rely on the change-of-support property, which also allows efficient computations in the DFGP model. To estimate model parameters, we devise a computationally efficient stochastic expectation-maximization (SEM) algorithm to ensure its scalability for massive datasets. The DFGP model is applied to a total of 3.7 million sea surface temperature datasets in the tropical Pacific Ocean for a one-week time period in 2010 from MODIS and AMSR-E instruments.

stat.ME

Spatial Statistical Downscaling for Constructing High-Resolution Nature Runs in Global Observing System Simulation Experiments

Observing system simulation experiments (OSSEs) have been widely used as a rigorous and cost-effective way to guide development of new observing systems, and to evaluate the performance of new data assimilation algorithms. Nature runs (NRs), which are outputs from deterministic models, play an essential role in building OSSE systems for global atmospheric processes because they are used both to create synthetic observations at high spatial resolution, and to represent the "true" atmosphere against which the forecasts are verified. However, most NRs are generated at resolutions coarser than actual observations. Here, we propose a principled statistical downscaling framework to construct high-resolution NRs via conditional simulation from coarse-resolution numerical model output. We use nonstationary spatial covariance function models that have basis function representations. This approach not only explicitly addresses the change-of-support problem, but also allows fast computation with large volumes of numerical model output. We also propose a data-driven algorithm to select the required basis functions adaptively, in order to increase the flexibility of our nonstationary covariance function models. In this article we demonstrate these techniques by downscaling a coarse-resolution physical NR at a native resolution of $1^{\circ} \text{ latitude} \times 1.25^{\circ} \text{ longitude}$ of global surface $\text{CO}_2$ concentrations to 655,362 equal-area hexagons.

stat.ME