SearcharxivSearch

arXiv subjects

Kung-Sik Chan

Publications and source records attributed to Kung-Sik Chan.

16 recordsLinked to original sources

Adaptive Change Point Detection with Matrix Time Series: Leveraging Structured Mean Shifts

In high-dimensional time series, the component processes are often assembled into a matrix to display their interrelationship. We focus on detecting mean shifts with unknown change point locations in these matrix time series. Series that are activated by a change may cluster along certain rows (columns), which forms mode-specific change point alignment. Leveraging mode-specific change point alignments may substantially enhance the power for change point detection. Yet, there may be no mode-specific alignments in the change point structure. We propose a powerful test to detect mode-specific change points, yet robust to non-mode-specific changes. We show the validity of using the multiplier bootstrap to compute the p-value of the proposed methods, and derive non-asymptotic bounds on the size and power of the tests. We also propose a parallel bootstrap, a computationally efficient approach for computing the p-value of the proposed adaptive test. In particular, we show the consistency of the proposed test, under mild regularity conditions. To obtain the theoretical results, we derive new, sharp bounds on Gaussian approximation and multiplier bootstrap approximation, which are of independent interest for high dimensional problems with diverging sparsity.

stat.ME

Spectral Change Point Estimation for High Dimensional Time Series by Sparse Tensor Decomposition

Multivariate time series may be subject to partial structural changes over certain frequency band, for instance, in neuroscience. We study the change point detection problem with high dimensional time series, within the framework of frequency domain. The overarching goal is to locate all change points and delineate which series are activated by the change, over which frequencies. In practice, the number of activated series per change and frequency could span from a few to full participation. We solve the problem by first computing a CUSUM tensor based on spectra estimated from blocks of the time series. A frequency-specific projection approach is applied for dimension reduction. The projection direction is estimated by a proposed tensor decomposition algorithm that adjusts to the sparsity level of changes. Finally, the projected CUSUM vectors across frequencies are aggregated for change point detection. We provide theoretical guarantees on the number of estimated change points and the convergence rate of their locations. We derive error bounds for the estimated projection direction for identifying the frequency-specific series activated in a change. We provide data-driven rules for the choice of parameters. The efficacy of the proposed method is illustrated by simulation and a stock returns application.

stat.ME

On a fast consistent selection of nested models with possibly unnormalized probability densities

Models with unnormalized probability density functions are ubiquitous in statistics, artificial intelligence and many other fields. However, they face significant challenges in model selection if the normalizing constants are intractable. Existing methods to address this issue often incur high computational costs, either due to numerical approximations of normalizing constants or evaluation of bias corrections in information criteria. In this paper, we propose a novel and fast selection criterion, MIC, for nested models of possibly dependent data, allowing direct data sampling from a possibly unnormalized probability density function. With a suitable multiplying factor depending only on the sample size and the model complexity, MIC gives a consistent selection under mild regularity conditions and is computationally efficient. Extensive simulation studies and real-data applications demonstrate the efficacy of MIC in the selection of nested models with unnormalized probability densities.

stat.ME

Mixture Matrix-valued Autoregressive Model

Time series of matrix-valued data are increasingly available in various areas including economics, finance, social science, among others. These data may shed light on the inter-dynamical relationships between two sets of attributes, for instance, countries and economic indices. The matrix autoregressive (MAR) model provides a parsimonious approach for analyzing such data. However, the MAR model, being a linear model with parametric constraints, cannot capture the nonlinear patterns in the data, such as regime shifts in the dynamics. We propose a mixture matrix autoregressive (MMAR) model for analyzing potential regime shifts in the dynamics between two attributes, for instance, due to recession versus expansion, or stable period versus pandemic. We propose an EM algorithm for maximum likelihood estimation. We derive some theoretical properties of the proposed method including consistency and asymptotic distribution, and illustrate its performance via simulations and real applications.

stat.ME

Individualized Prediction Bands in Causal Inference with Continuous Treatments

Individualized treatments are crucial for optimal decision making and treatment allocation, specifically in personalized medicine based on the estimation of an individual's dose-response curve across a continuum of treatment levels, e.g., drug dosage. Current works focus on conditional mean and median estimates, which are useful but do not provide the full picture. We propose viewing causal inference with a continuous treatment as a covariate shift. This allows us to leverage existing weighted conformal prediction methods with both quantile and point estimates to compute individualized uncertainty quantification for dose-response curves. Our method, individualized prediction bands (IPB), is demonstrated via simulations and a real data analysis, which demonstrates the additional medical expenditure caused by continued smoking for selected individuals. The results demonstrate that IPB provides an effective solution to a gap in individual dose-response uncertainty quantification literature.

stat.ME

Highest Probability Density Conformal Regions

This paper proposes a new method for finding the highest predictive density set or region, within the heteroscedastic regression framework. This framework enjoys the property that any highest predictive density set is a translation of some scalar multiple of a highest density set for the standardized regression error, with the same prediction accuracy. The proposed method leverages this property to efficiently compute conformal prediction regions, using signed conformal inference, kernel density estimation, in conjunction with any conditional mean, and scale estimators. While most conformal prediction methods output prediction intervals, this method adapts to the target. When the target is multi-modal, the proposed method outputs an approximation of the smallest multi-modal set. When the target is uni-modal, the proposed method outputs an approximation of the smallest interval. Under mild regularity conditions, we show that these conformal prediction sets are asymptotically close to the true smallest prediction sets. Because of the conformal guarantee, even in finite sample sizes the method has guaranteed coverage. With simulations and a real data analysis we demonstrate that the proposed method is better than existing methods when the target is multi-modal, and gives similar results when the target is uni-modal. Supplementary materials, including proofs and additional images, are available online.

stat.ME

Flexible Conformal Highest Predictive Conditional Density Sets

We introduce our method, conformal highest conditional density sets (CHCDS), that forms conformal prediction sets using existing estimated conditional highest density predictive regions. We prove the validity of the method, and that conformal adjustment is negligible under some regularity conditions. In particular, if we correctly specify the underlying conditional density estimator, the conformal adjustment will be negligible. The conformal adjustment, however, always provides guaranteed nominal unconditional coverage, even when the underlying model is incorrectly specified. We compare the proposed method via simulation and a real data analysis to other existing methods. Our numerical results show that CHCDS is better than existing methods in scenarios where the error term is multi-modal, and just as good as existing methods when the error terms are unimodal.

stat.ME

Conformal Multi-Target Hyperrectangles

We propose conformal hyperrectangular prediction regions for multi-target regression. We propose split conformal prediction algorithms for both point and quantile regression to form hyperrectangular prediction regions, which allow for easy marginal interpretation and do not require covariance estimation. In practice, it is preferable that a prediction region is balanced, that is, having identical marginal prediction coverage, since prediction accuracy is generally equally important across components of the response vector. The proposed algorithms possess two desirable properties, namely, tight asymptotic overall nominal coverage as well as asymptotic balance, that is, identical asymptotic marginal coverage, under mild conditions. We then compare our methods to some existing methods on both simulated and real data sets. Our simulation results and real data analysis show that our methods outperform existing methods while achieving the desired nominal coverage and good balance between dimensions.

stat.ME

Testing for threshold regulation in presence of measurement error with an application to the PPP hypothesis

Regulation is an important feature characterising many dynamical phenomena and can be tested within the threshold autoregressive setting, with the null hypothesis being a global non-stationary process. Nonetheless, this setting is debatable since data are often corrupted by measurement errors. Thus, it is more appropriate to consider a threshold autoregressive moving-average model as the general hypothesis. We implement this new setting with the integrated moving-average model of order one as the null hypothesis. We derive a Lagrange multiplier test which has an asymptotically similar null distribution and provide the first rigorous proof of tightness pertaining to testing for threshold nonlinearity against difference stationarity, which is of independent interest. Simulation studies show that the proposed approach enjoys less bias and higher power in detecting threshold regulation than existing tests when there are measurement errors. We apply the new approach to the daily real exchange rates of Eurozone countries. It lends support to the purchasing power parity hypothesis, via a nonlinear mean-reversion mechanism triggered upon crossing a threshold located in the extreme upper tail. Furthermore, we analyse the Eurozone series and propose a threshold autoregressive moving-average specification, which sheds new light on the purchasing power parity debate.

stat.ME

Testing for threshold effects in the TARMA framework

We present supremum Lagrange Multiplier tests to compare a linear ARMA specification against its threshold ARMA extension. We derive the asymptotic distribution of the test statistics both under the null hypothesis and contiguous local alternatives. Moreover, we prove the consistency of the tests. The Monte Carlo study shows that the tests enjoy good finite-sample properties, are robust against model mis-specification and their performance is not affected if the order of the model is unknown. The tests present a low computational burden and do not suffer from some of the drawbacks that affect the quasi-likelihood ratio setting. Lastly, we apply our tests to a time series of standardized tree-ring growth indexes and this can lead to new research in climate studies.

stat.ME

A Conversation with Howell Tong

The following conversation is partly based on an interview that took place in the Hong Kong University of Science and Technology in July 2013.

stat.OT

Testing for shielding of special nuclear weapon materials

Nuclear-weapon-material detection via gamma-ray sensing is routinely applied, for example, in monitoring cross-border traffic. Natural or deliberate shielding both attenuates and distorts the shape of the gamma-ray spectra of specific radionuclides, thereby making such routine applications challenging. We develop a Lagrange multiplier (LM) test for shielding. A strong advantage of the LM test is that it only requires fitting a much simpler model that assumes no shielding. We show that, under the null hypothesis and some mild regularity conditions and as the detection time increases, LM test statistic for (composite) shielding is asymptotically Chi-square with the degree of freedom equal to the presumed number of shielding materials. We also derive the local power of the LM test. Extensive simulation studies suggest that the test is robust to the number and nature of the intervening materials, which owes to the fact that common intervening materials have broadly similar attenuation functions.

stat.AP

Inference of seasonal long-memory aggregate time series

Time-series data with regular and/or seasonal long-memory are often aggregated before analysis. Often, the aggregation scale is large enough to remove any short-memory components of the underlying process but too short to eliminate seasonal patterns of much longer periods. In this paper, we investigate the limiting correlation structure of aggregate time series within an intermediate asymptotic framework that attempts to capture the aforementioned sampling scheme. In particular, we study the autocorrelation structure and the spectral density function of aggregates from a discrete-time process. The underlying discrete-time process is assumed to be a stationary Seasonal AutoRegressive Fractionally Integrated Moving-Average (SARFIMA) process, after suitable number of differencing if necessary, and the seasonal periods of the underlying process are multiples of the aggregation size. We derive the limit of the normalized spectral density function of the aggregates, with increasing aggregation. The limiting aggregate (seasonal) long-memory model may then be useful for analyzing aggregate time-series data, which can be estimated by maximizing the Whittle likelihood. We prove that the maximum Whittle likelihood estimator (spectral maximum likelihood estimator) is consistent and asymptotically normal, and study its finite-sample properties through simulation. The efficacy of the proposed approach is illustrated by a real-life internet traffic example.

math.ST

Semiparametric zero-inflated modeling in multi-ethnic study of atherosclerosis (MESA)

We analyze the Agatston score of coronary artery calcium (CAC) from the Multi-Ethnic Study of Atherosclerosis (MESA) using the semiparametric zero-inflated modeling approach, where the observed CAC scores from this cohort consist of high frequency of zeroes and continuously distributed positive values. Both partially constrained and unconstrained models are considered to investigate the underlying biological processes of CAC development from zero to positive, and from small amount to large amount. Different from existing studies, a model selection procedure based on likelihood cross-validation is adopted to identify the optimal model, which is justified by comparative Monte Carlo studies. A shrinkaged version of cubic regression spline is used for model estimation and variable selection simultaneously. When applying the proposed methods to the MESA data analysis, we show that the two biological mechanisms influencing the initiation of CAC and the magnitude of CAC when it is positive are better characterized by an unconstrained zero-inflated normal model. Our results are significantly different from those in published studies, and may provide further insights into the biological mechanisms underlying CAC development in humans. This highly flexible statistical framework can be applied to zero-inflated data analyses in other areas.

stat.AP

Reduced rank regression via adaptive nuclear norm penalization

Adaptive nuclear-norm penalization is proposed for low-rank matrix approximation, by which we develop a new reduced-rank estimation method for the general high-dimensional multivariate regression problems. The adaptive nuclear norm of a matrix is defined as the weighted sum of the singular values of the matrix. For example, the pre-specified weights may be some negative power of the singular values of the data matrix (or its projection in regression setting). The adaptive nuclear norm is generally non-convex under the natural restriction that the weight decreases with the singular value. However, we show that the proposed non-convex penalized regression method has a global optimal solution obtained from an adaptively soft-thresholded singular value decomposition. This new reduced-rank estimator is computationally efficient, has continuous solution path and possesses better bias-variance property than its classical counterpart. The rank consistency and prediction/estimation performance bounds of the proposed estimator are established under high-dimensional asymptotic regime. Simulation studies and an application in genetics demonstrate that the proposed estimator has superior performance to several existing methods. The adaptive nuclear-norm penalization can also serve as a building block to study a broad class of singular value penalties.

stat.ME