SearcharxivSearch

arXiv subjects

Kin Wai Chan

Publications and source records attributed to Kin Wai Chan.

8 recordsLinked to original sources

Tight differencing in spectral density estimation with centrosymmetric kernels

Mean-robust estimation of spectral density and long-run variance is crucial for many statistical inference procedures. However, existing methods often degrade when serially dependent data exhibit volatile, time-varying trends, or sudden jumps, particularly in small samples. While differencing and kernel averaging are standard tools for achieving mean robustness and consistency, they are not inherently compatible. Combining them can compromise optimality. Specifically, tight differencing, an operation of taking small-lag differences to enhance local de-trending, introduces strong correlations that distort the high-order properties of kernel-averaged estimators. To resolve this incompatibility, we introduce a novel class of centrosymmetric kernels explicitly designed to integrate with tight differencing. We demonstrate that the optimal tight difference sequence for serially dependent data differs from classical sequences designed for independent data. Notably, these proposed optimal sequences are data-independent and can be applied directly without pre-fitting. Finally, the proposed estimators are demonstrated to be useful across various statistical inference tasks, including tests for stationarity and white noise.

stat.ME

Tail postcoloring in long-run variance estimation of time series

Prewhitening is a common approach to deal with strong autocorrelation. In this article, we propose a new approach called tail postcoloring, motivated by it. It uses parametric models to project, or color back, the neglected tail autocovariances in nonparametric estimators onto the final estimator. This approach bridges the non-parametric variance estimator and the parametric coloring model through a scaling factor. It automatically switches between these two arms using a bandwidth parameter, without the need to transform the entire dataset into residuals, as in the standard prewhitening approach. When the coloring model is well-specified, a parametric rate can be achieved. In finite samples, it is also more robust to misspecification of the coloring model compared to the whitening model in the standard approach. Besides, it avoids severe potential variance inflation or power reduction caused by the recoloring factor in the standard approach. We show that multiple parametric models can be used to construct a multiply robust tail postcolored estimator. It also naturally works for multivariate time series. A real-data example in Markov chain Monte Carlo output analysis is provided.

stat.ME

Online Generalized Method of Moments for Time Series

Online learning has gained popularity in recent years due to the urgent need to analyse large-scale streaming data, which can be collected in perpetuity and serially dependent. This motivates us to develop the online generalized method of moments (OGMM), an explicitly updated estimation and inference framework in the time series setting. The OGMM inherits many properties of offline GMM, such as its broad applicability to many problems in econometrics and statistics, natural accommodation for over-identification, and achievement of semiparametric efficiency under temporal dependence. As an online method, the key gain relative to offline GMM is the vast improvement in time complexity and memory requirement. Building on the OGMM framework, we propose improved versions of online Sargan--Hansen and structural stability tests following recent work in econometrics and statistics. Through Monte Carlo simulations, we observe encouraging finite-sample performance in online instrumental variables regression, online over-identifying restrictions test, online quantile regression, and online anomaly detection. Interesting applications of OGMM to stochastic volatility modelling and inertial sensor calibration are presented to demonstrate the effectiveness of OGMM.

stat.ME

Principles of Statistical Inference in Online Problems

To investigate a dilemma of statistical and computational efficiency faced by long-run variance estimators, we propose a decomposition of kernel weights in a quadratic form and some online inference principles. These proposals allow us to characterize efficient online long-run variance estimators. Our asymptotic theory and simulations show that this principle-driven approach leads to online estimators with a uniformly lower mean squared error than all existing works. We also discuss practical enhancements such as mini-batch and automatic updates to handle fast streaming data and optimal parameters tuning. Beyond variance estimation, we consider the proposals in the context of online quantile regression, online change point detection, Markov chain Monte Carlo convergence diagnosis, and stochastic approximation. Substantial improvements in computational cost and finite-sample statistical properties are observed when we apply our principle-driven variance estimator to original and modified inference procedures.

stat.ME

A General Framework For Constructing Locally Self-Normalized Multiple-Change-Point Tests

We propose a general framework to construct self-normalized multiple-change-point tests with time series data. The only building block is a user-specified one-change-point detecting statistic, which covers a wide class of popular methods, including cumulative sum process, outlier-robust rank statistics and order statistics. Neither robust and consistent estimation of nuisance parameters, selection of bandwidth parameters, nor pre-specification of the number of change points is required. The finite-sample performance shows that our proposal is size-accurate, robust against misspecification of the alternative hypothesis, and more powerful than existing methods. Case studies of NASDAQ option volume and Shanghai-Hong Kong Stock Connect turnover are provided.

stat.ME

General and Feasible Tests with Multiply-Imputed Datasets

Multiple imputation (MI) is a technique especially designed for handling missing data in public-use datasets. It allows analysts to perform incomplete-data inference straightforwardly by using several already imputed datasets released by the dataset owners. However, the existing MI tests require either a restrictive assumption on the missing-data mechanism, known as equal odds of missing information (EOMI), or an infinite number of imputations. Some of them also require analysts to have access to restrictive or non-standard computer subroutines. Besides, the existing MI testing procedures cover only Wald's tests and likelihood ratio tests but not Rao's score tests, therefore, these MI testing procedures are not general enough. In addition, the MI Wald's tests and MI likelihood ratio tests are not procedurally identical, so analysts need to resort to distinct algorithms for implementation. In this paper, we propose a general MI procedure, called stacked multiple imputation (SMI), for performing Wald's tests, likelihood ratio tests and Rao's score tests by a unified algorithm. SMI requires neither EOMI nor an infinite number of imputations. It is particularly feasible for analysts as they just need to use a complete-data testing device for performing the corresponding incomplete-data test.

stat.ME

Optimal Difference-based Variance Estimators in Time Series: A General Framework

Variance estimation is important for statistical inference. It becomes non-trivial when observations are masked by serial dependence structures and time-varying mean structures. Existing methods either ignore or sub-optimally handle these nuisance structures. This paper develops a general framework for the estimation of the long-run variance for time series with non-constant means. The building blocks are difference statistics. The proposed class of estimators is general enough to cover many existing estimators. Necessary and sufficient conditions for consistency are investigated. The first asymptotically optimal estimator is derived. Our proposed estimator is theoretically proven to be invariant to arbitrary mean structures, which may include trends and a possibly divergent number of discontinuities.

stat.ME

Multiple Improvements of Multiple Imputation Likelihood Ratio Tests

Multiple imputation (MI) inference handles missing data by imputing the missing values $m$ times, and then combining the results from the $m$ complete-data analyses. However, the existing method for combining likelihood ratio tests (LRTs) has multiple defects: (i) the combined test statistic can be negative, but its null distribution is approximated by an $F$-distribution; (ii) it is not invariant to re-parametrization; (iii) it fails to ensure monotonic power owing to its use of an inconsistent estimator of the fraction of missing information (FMI) under the alternative hypothesis; and (iv) it requires nontrivial access to the LRT statistic as a function of parameters instead of data sets. We show, using both theoretical derivations and empirical investigations, that essentially all of these problems can be straightforwardly addressed if we are willing to perform an additional LRT by stacking the $m$ completed data sets as one big completed data set. This enables users to implement the MI LRT without modifying the complete-data procedure. A particularly intriguing finding is that the FMI can be estimated consistently by an LRT statistic for testing whether the $m$ completed data sets can be regarded effectively as samples coming from a common model. Practical guidelines are provided based on an extensive comparison of existing MI tests. Issues related to nuisance parameters are also discussed.

math.ST