SearcharxivSearch

arXiv subjects

Zhaoyuan Li

Publications and source records attributed to Zhaoyuan Li.

9 recordsLinked to original sources

Sequential Change Detection in Correlation Structures with Window-Limited Statistics

We consider detecting change points in the correlation structure of streaming data with minimum assumptions posed on the underlying data distribution. Detection statistics are constructed for dense and sparse change settings, based on $\ell_1$ and $\ell_{\infty}$ norms of the squared difference of vectorized pre- and post-change correlation matrices, respectively. We also propose a novel threshold determination algorithm based on sign-flip permutations that enhances the efficiency of our procedure, particularly when the data dimension is large compared to the window size. Theoretical guarantees of the proposed methods are provided in terms of average run length in the no-change regime and expected detection delay in the post-change regime. We evaluate the performance of the proposed methods across a wide range of simulated datasets and demonstrate their effectiveness, with small detection delays that are comparable to the exact optimal CUSUM test. Finally, we demonstrate the effectiveness of our methods on real-world datasets, including El Ni{ñ}o event forecasting, where we achieve a state-of-the-art hit rate exceeding 0.86 with near-zero false alarms, as well as seismic event detection.

stat.ME

Efficient Change Point Detection and Estimation in High-Dimensional Correlation Matrices

This paper considers the problems of detecting a change point and estimating the location in the correlation matrices of a sequence of high-dimensional vectors, where the dimension is large enough to be comparable to the sample size or even much larger. A new break test is proposed based on signflip parallel analysis to detect the existence of change points. Furthermore, a two-step approach combining a signflip permutation dimension reduction step and a CUSUM statistic is proposed to estimate the change point's location and recover the support of changes. The consistency of the estimator is constructed. Simulation examples and real data applications illustrate the superior empirical performance of the proposed methods. Especially, the proposed methods outperform existing ones for non-Gaussian data and the change point in the extreme tail of a sequence and become more accurate as the dimension p increases. Supplementary materials for this article are available online.

stat.ME

Unified and robust Lagrange multiplier type tests for cross-sectional independence in large panel data models

This paper revisits the Lagrange multiplier type test for the null hypothesis of no cross-sectional dependence in large panel data models. We propose a unified test procedure and its power enhancement version, which show robustness for a wide class of panel model contexts. Specifically, the two procedures are applicable to both heterogeneous and fixed effects panel data models with the presence of weakly exogenous as well as lagged dependent regressors, allowing for a general form of nonnormal error distribution. With the tools from Random Matrix Theory, the asymptotic validity of the test procedures is established under the simultaneous limit scheme where the number of time periods and the number of cross-sectional units go to infinity proportionally. The derived theories are accompanied by detailed Monte Carlo experiments, which confirm the robustness of the two tests and also suggest the validity of the power enhancement technique.

econ.EM

On John's test for sphericity in large panel data models

This paper studies John's test for sphericity of the error terms in large panel data models, where the number of cross-section units $n$ is large enough to be comparable to the number of times series observations $T$, or even larger. Based on recent random matrix theory results, John's test's asymptotic normality properties are established under both the null and the alternative hypotheses. These asymptotics are valid for general populations, i.e., not necessarily Gaussian provided certain finite moments. A fantastic phenomenon found in the paper is that John's test for panel data models possesses a powerful dimension-proof property. It keeps the same null distribution under different $(n,T)$-asymptotics, i.e., the small or medium panel regime $n/T\to 0$ as $T\to \infty$, the large panel regime $n/T\to c \in (0,\infty)$ as $ T\to \infty$, and the ultra-large panel regime $n/T\to \infty (T^δ/n =O_p(1), 1<δ<2)$ as $T\to \infty$. Moreover, John's test is always consistent except under the alternative of bounded-norm covariance with the large panel regime $n/T\to c \in (0,\infty)$.

math.ST

Extension of the Lagrange multiplier test for error cross-section independence to large panels with non normal errors

This paper reexamines the seminal Lagrange multiplier test for cross-section independence in a large panel model where both the number of cross-sectional units n and the number of time series observations T can be large. The first contribution of the paper is an enlargement of the test with two extensions: firstly the new asymptotic normality is derived in a simultaneous limiting scheme where the two dimensions (n, T) tend to infinity with comparable magnitudes; second, the result is valid for general error distribution (not necessarily normal). The second contribution of the paper is a new test statistic based on the sum of the fourth powers of cross-section correlations from OLS residuals, instead of their squares used in the Lagrange multiplier statistic. This new test is generally more powerful, and the improvement is particularly visible against alternatives with weak or sparse cross-section dependence. Both simulation study and real data analysis are proposed to demonstrate the advantages of the enlarged Lagrange multiplier test and the power enhanced test in comparison with the existing procedures.

econ.EM

Network-based Approach and Climate Change Benefits for Forecasting the Amount of Indian Monsoon Rainfall

The Indian summer monsoon rainfall (ISMR) has a decisive influence on India's agricultural output and economy. Extreme deviations from the normal seasonal amount of rainfall can cause severe droughts or floods, affecting Indian food production and security. Despite the development of sophisticated statistical and dynamical climate models, a long-term and reliable prediction of the ISMR has remained a challenging problem. Towards achieving this goal, here we construct a series of dynamical and physical climate networks based on the global near surface air temperature field. We uncover that some characteristics of the directed and weighted climate networks can serve as efficient long-term predictors for ISMR forecasting. The developed prediction method produces a forecast skill of 0.5 with a 5-month lead-time in advance by using the previous calendar year's data. The skill of our ISMR forecast, is comparable to the current state-of-the-art models, however, with quite a short (i.e., within one month) lead-time. We discuss the underlying mechanism of our predictor and associate it with network-delayed-ENSO and ENSO-monsoon connections. Moreover, our approach allows predicting the all India rainfall, as well as forecasting the different Indian homogeneous regions' rainfall, which is crucial for agriculture in India. We reveal that global warming affects the climate network by enhancing cross-equatorial teleconnections between Southwest Atlantic, Western part of the Indian Ocean, and North Asia-Pacific with significant impacts on the precipitation in India. We find a hotspots area in the mid-latitude South Atlantic, which is the basis for our predictor. Remarkably, the significant warming trend in this area yields an improvement of the prediction skill.

physics.ao-ph

Testing for Heteroscedasticity in High-dimensional Regressions

Testing heteroscedasticity of the errors is a major challenge in high-dimensional regressions where the number of covariates is large compared to the sample size. Traditional procedures such as the White and the Breusch-Pagan tests typically suffer from low sizes and powers. This paper proposes two new test procedures based on standard OLS residuals. Using the theory of random Haar orthogonal matrices, the asymptotic normality of both test statistics is obtained under the null when the degree of freedom tends to infinity. This encompasses both the classical low-dimensional setting where the number of variables is fixed while the sample size tends to infinity, and the proportional high-dimensional setting where these dimensions grow to infinity proportionally. These procedures thus offer a wide coverage of dimensions in applications. To our best knowledge, this is the first procedures in the literature for testing heteroscedasticity which are valid for medium and high-dimensional regressions. The superiority of our proposed tests over the existing methods are demonstrated by extensive simulations and by several real data analyses as well.

stat.ME

On Two Simple and Effective Procedures for High Dimensional Classification of General Populations

In this paper, we generalize two criteria, the determinant-based and trace-based criteria proposed by Saranadasa (1993), to general populations for high dimensional classification. These two criteria compare some distances between a new observation and several different known groups. The determinant-based criterion performs well for correlated variables by integrating the covariance structure and is competitive to many other existing rules. The criterion however requires the measurement dimension be smaller than the sample size. The trace-based criterion in contrast, is an independence rule and effective in the "large dimension-small sample size" scenario. An appealing property of these two criteria is that their implementation is straightforward and there is no need for preliminary variable selection or use of turning parameters. Their asymptotic misclassification probabilities are derived using the theory of large dimensional random matrices. Their competitive performances are illustrated by intensive Monte Carlo experiments and a real data analysis.

stat.ME

On estimation of the noise variance in high-dimensional probabilistic principal component analysis

In this paper, we develop new statistical theory for probabilistic principal component analysis models in high dimensions. The focus is the estimation of the noise variance, which is an important and unresolved issue when the number of variables is large in comparison with the sample size. We first unveil the reasons of a widely observed downward bias of the maximum likelihood estimator of the variance when the data dimension is high. We then propose a bias-corrected estimator using random matrix theory and establish its asymptotic normality. The superiority of the new (bias-corrected) estimator over existing alternatives is first checked by Monte-Carlo experiments with various combinations of $(p, n)$ (dimension and sample size). In order to demonstrate further potential benefits from the results of the paper to general probability PCA analysis, we provide evidence of net improvements in two popular procedures (Ulfarsson and Solo, 2008; Bai and Ng, 2002) for determining the number of principal components when the respective variance estimator proposed by these authors is replaced by the bias-corrected estimator. The new estimator is also used to derive new asymptotics for the related goodness-of-fit statistic under the high-dimensional scheme.

math.ST