SearcharxivSearch

arXiv subjects

Hsin-Cheng Huang

Publications and source records attributed to Hsin-Cheng Huang.

11 recordsLinked to original sources

Reconstructing East Asian Temperatures from 1368 to 1911 Using Historical Documents, Climate Models, and Data Assimilation

We propose a novel approach for reconstructing annual temperatures in East Asia from 1368 to 1911, leveraging the Reconstructed East Asian Climate Historical Encoded Series (REACHES). The lack of instrumental data during this period poses significant challenges to understanding past climate conditions. REACHES digitizes historical documents from the Ming and Qing dynasties of China, converting qualitative descriptions into a four-level ordinal temperature scale. However, these index-based data are biased toward abnormal or extreme weather phenomena, leading to data gaps that likely correspond to normal conditions. To address this bias and reconstruct historical temperatures at any point within East Asia, including locations without direct historical data, we employ a three-tiered statistical framework. First, we perform kriging to interpolate temperature data across East Asia, adopting a zero-mean assumption to handle missing information. Next, we utilize the Last Millennium Ensemble (LME) reanalysis data and apply quantile mapping to calibrate the kriged REACHES data to Celsius temperature scales. Finally, we introduce a novel Bayesian data assimilation method that integrates the kriged Celsius data with LME simulations to enhance reconstruction accuracy. We model the LME data at each geographic location using a flexible nonstationary autoregressive time series model and employ regularized maximum likelihood estimation with a fused lasso penalty. The resulting dynamic distribution serves as a prior, which is refined via Kalman filtering by incorporating the kriged Celsius REACHES data to yield posterior temperature estimates. This comprehensive integration of historical documentation, contemporary climate models, and advanced statistical methods improves the accuracy of historical temperature reconstructions and provides a crucial resource for future environmental and climate studies.

stat.AP

Assessing Spatial Stationarity and Segmenting Spatial Processes into Stationary Components

In this research, we propose a novel technique for visualizing nonstationarity in geostatistics, particularly when confronted with a single realization of data at irregularly spaced locations. Our method hinges on formulating a statistic that tracks a stable microergodic parameter of the exponential covariance function, allowing us to address the intricate challenges of nonstationary processes that lack repeated measurements. We implement the fused lasso technique to elucidate nonstationary patterns at various resolutions. For prediction purposes, we segment the spatial domain into stationary sub-regions via Voronoi tessellations. Additionally, we devise a robust test for stationarity based on contrasting the sample means of our proposed statistics between two selected Voronoi subregions. The effectiveness of our method is demonstrated through simulation studies and its application to a precipitation dataset in Colorado.

stat.ME

Spatially Adaptive Calibrations of AirBox PM$_{2.5}$ Data

Two networks are available to monitor PM$_{2.5}$ in Taiwan, including the Taiwan Air Quality Monitoring Network (TAQMN) and the AirBox network. The TAQMN, managed by Taiwan's Environmental Protection Administration (EPA), provides high-quality PM$_{2.5}$ measurements at $77$ monitoring stations. More recently, the AirBox network was launched, consisting of low-cost, small internet-of-things (IoT) microsensors (i.e., AirBoxes) at thousands of locations. While the AirBox network provides broad spatial coverage, its measurements are not reliable and require calibrations. However, applying a universal calibration procedure to all AirBoxes does not work well because the calibration curves vary with several factors, including the chemical compositions of PM$_{2.5}$, which are not homogeneous in space. Therefore, different calibrations are needed at different locations with different local environments. Unfortunately, most AirBoxes are not close to EPA stations, making the calibration task challenging. In this article, we propose a spatial model with spatially varying coefficients to account for heteroscedasticity in the data. Our method gives adaptive calibrations of AirBoxes according to their local conditions and provides accurate PM$_{2.5}$ concentrations at any location in Taiwan, incorporating two types of measurements. In addition, the proposed method automatically calibrates measurements from a new AirBox once it is added to the network. We illustrate our approach using hourly PM$_{2.5}$ data in the year 2020. After the calibration, the results show that the PM$_{2.5}$ prediction improves about 37% to 67% in root mean-squared prediction error for matching EPA data. In particular, once the calibration curves are established, we can obtain reliable PM$_{2.5}$ values at any location in Taiwan, even if we ignore EPA data.

stat.ME

Inference of Random Effects for Linear Mixed-Effects Models with a Fixed Number of Clusters

We consider a linear mixed-effects model with a clustered structure, where the parameters are estimated using maximum likelihood (ML) based on possibly unbalanced data. Inference with this model is typically done based on asymptotic theory, assuming that the number of clusters tends to infinity with the sample size. However, when the number of clusters is fixed, classical asymptotic theory developed under a divergent number of clusters is no longer valid and can lead to erroneous conclusions. In this paper, we establish the asymptotic properties of the ML estimators of random-effects parameters under a general setting, which can be applied to conduct valid statistical inference with fixed numbers of clusters. Our asymptotic theorems allow both fixed effects and random effects to be misspecified, and the dimensions of both effects to go to infinity with the sample size.

math.ST

Vector Autoregressive Models with Spatially Structured Coefficients for Time Series on a Spatial Grid

We propose a parsimonious spatiotemporal model for time series data on a spatial grid. Our model is capable of dealing with high-dimensional time series data that may be collected at hundreds of locations and capturing the spatial non-stationarity. In essence, our model is a vector autoregressive model that utilizes the spatial structure to achieve parsimony of autoregressive matrices at two levels. The first level ensures the sparsity of the autoregressive matrices using a lagged-neighborhood scheme. The second level performs a spatial clustering of the non-zero autoregressive coefficients such that nearby locations share similar coefficients. This model is interpretable and can be used to identify geographical subregions, within each of which, the time series share similar dynamical behavior with homogeneous autoregressive coefficients. The model parameters are obtained using the penalized maximum likelihood with an adaptive fused Lasso penalty. The estimation procedure is easy to implement and can be tailored to the need of a modeler. We illustrate the performance of the proposed estimation algorithm in a simulation study. We apply our model to a wind speed time series dataset generated from a climate model over Saudi Arabia to illustrate its usefulness. Limitations and possible extensions of our method are also discussed.

stat.ME

False Discovery Rates to Detect Signals from Incomplete Spatially Aggregated Data

There are a number of ways to test for the absence/presence of a spatial signal in a completely observed fine-resolution image. One of these is a powerful nonparametric procedure called Enhanced False Discovery Rate (EFDR). A drawback of EFDR is that it requires the data to be defined on regular pixels in a rectangular spatial domain. Here, we develop an EFDR procedure for possibly incomplete data defined on irregular small areas. Motivated by statistical learning, we use conditional simulation (CS) to condition on the available data and simulate the full rectangular image at its finest resolution many times (M, say). EFDR is then applied to each of these simulations resulting in M estimates of the signal and M statistically dependent p-values. Averaging over these estimates yields a single, combined estimate of a possible signal, but inference is needed to determine whether there really is a signal present. We test the original null hypothesis of no signal by combining the M p-values into a single p-value using copulas and a composite likelihood. If the null hypothesis of no signal is rejected, we use the combined estimate. We call this new procedure EFDR-CS and, to demonstrate its effectiveness, we show results from a simulation study; an experiment where we introduce aggregation and incompleteness into temperature-change data in the Asia-Pacific; and an application to total-column carbon dioxide from satellite remote sensing data over a region of the Middle East, Afghanistan, and the western part of Pakistan.

stat.ME

Regularized Spatial Maximum Covariance Analysis

In climate and atmospheric research, many phenomena involve more than one meteorological spatial processes covarying in space. To understand how one process is affected by another, maximum covariance analysis (MCA) is commonly applied. However, the patterns obtained from MCA may sometimes be difficult to interpret. In this paper, we propose a regularization approach to promote spatial features in dominant coupled patterns by introducing smoothness and sparseness penalties while accounting for their orthogonalities. We develop an efficient algorithm to solve the resulting optimization problem by using the alternating direction method of multipliers. The effectiveness of the proposed method is illustrated by several numerical examples, including an application to study how precipitations in east Africa are affected by sea surface temperatures in the Indian Ocean.

stat.ME

Mixed domain asymptotics for a stochastic process model with time trend and measurement error

We consider a stochastic process model with time trend and measurement error. We establish consistency and derive the limiting distributions of the maximum likelihood (ML) estimators of the covariance function parameters under a general asymptotic framework, including both the fixed domain and the increasing domain frameworks, even when the time trend model is misspecified or its complexity increases with the sample size. In particular, the convergence rates of the ML estimators are thoroughly characterized in terms of the growing rate of the domain and the degree of model misspecification/complexity.

math.ST

Regularized Principal Component Analysis for Spatial Data

In many atmospheric and earth sciences, it is of interest to identify dominant spatial patterns of variation based on data observed at $p$ locations and $n$ time points with the possibility that $p>n$. While principal component analysis (PCA) is commonly applied to find the dominant patterns, the eigenimages produced from PCA may exhibit patterns that are too noisy to be physically meaningful when $p$ is large relative to $n$. To obtain more precise estimates of eigenimages, we propose a regularization approach incorporating smoothness and sparseness of eigenimages, while accounting for their orthogonality. Our method allows data taken at irregularly spaced or sparse locations. In addition, the resulting optimization problem can be solved using the alternating direction method of multipliers, which is easy to implement, and applicable to a large spatial dataset. Furthermore, the estimated eigenfunctions provide a natural basis for representing the underlying spatial process in a spatial random-effects model, from which spatial covariance function estimation and spatial prediction can be efficiently performed using a regularized fixed-rank kriging method. Finally, the effectiveness of the proposed method is demonstrated by several numerical examples

stat.ME

Multi-Resolution Spatial Random-Effects Models for Irregularly Spaced Data

The spatial random-effects model is flexible in modeling spatial covariance functions, and is computationally efficient for spatial prediction via fixed rank kriging. However, the success of this model depends on an appropriate set of basis functions. In this research, we propose a class of basis functions extracted from thin-plate splines. These functions are ordered in terms of their degrees of smoothness with a higher-order function corresponding to larger-scale features and a lower-order one corresponding to smaller-scale details, leading to a parsimonious representation for a nonstationary spatial covariance function. Consequently, only a small to moderate number of functions are needed in a spatial random-effects model. The proposed class of basis functions has several advantages over commonly used ones. First, we do not need to concern about the allocation of the basis functions, but simply select the total number of functions corresponding to a resolution. Second, only a small number of basis functions is usually required, which facilitates computation. Third, estimation variability of model parameters can be considerably reduced, and hence more precise covariance function estimates can be obtained. Fourth, the proposed basis functions depend only on the data locations but not the measurements taken at those locations, and are applicable regardless of whether the data locations are sparse or irregularly spaced. In addition, we derive a simple close-form expression for the maximum likelihood estimates of model parameters in the spatial random-effects model. Some numerical examples are provided to demonstrate the effectiveness of the proposed method.

stat.ME

Asymptotic theory of generalized information criterion for geostatistical regression model selection

Information criteria, such as Akaike's information criterion and Bayesian information criterion are often applied in model selection. However, their asymptotic behaviors for selecting geostatistical regression models have not been well studied, particularly under the fixed domain asymptotic framework with more and more data observed in a bounded fixed region. In this article, we study the generalized information criterion (GIC) for selecting geostatistical regression models under a more general mixed domain asymptotic framework. Via uniform convergence developments of some statistics, we establish the selection consistency and the asymptotic loss efficiency of GIC under some regularity conditions, regardless of whether the covariance model is correctly or wrongly specified. We further provide specific examples with different types of explanatory variables that satisfy the conditions. For example, in some situations, GIC is selection consistent, even when some spatial covariance parameters cannot be estimated consistently. On the other hand, GIC fails to select the true polynomial order consistently under the fixed domain asymptotic framework. Moreover, the growth rate of the domain and the degree of smoothness of candidate regressors in space are shown to play key roles for model selection.

math.ST