Searcharxiv⌕ Search

arXiv subjects

Long Feng

Publications and source records attributed to Long Feng.

At least 73 records · Page 4Linked to original sources

Robust Mutual Fund Selection with False Discovery Rate Control

In this article, we address the challenge of identifying skilled mutual funds among a large pool of candidates, utilizing the linear factor pricing model. Assuming observable factors with a weak correlation structure for the idiosyncratic error, we propose a spatial-sign based multiple testing procedure (SS-BH). When latent factors are present, we first extract them using the elliptical principle component method (He et al. 2022) and then propose a factor-adjusted spatial-sign based multiple testing procedure (FSS-BH). Simulation studies demonstrate that our proposed FSS-BH procedure performs exceptionally well across various applications and exhibits robustness to variations in the covariance structure and the distribution of the error term. Additionally, real data application further highlights the superiority of the FSS-BH procedure.

stat.ME↗

Adaptive Sphericity Tests for High Dimensional Data

In this paper, we investigate sphericity testing in high-dimensional settings, where existing methods primarily rely on sum-type test procedures that often underperform under sparse alternatives. To address this limitation, we propose two max-type test procedures utilizing the sample covariance matrix and the sample spatial-sign covariance matrix, respectively. Furthermore, we introduce two Cauchy combination test procedures that integrate both sum-type and max-type tests, demonstrating their superiority across a wide range of sparsity levels in the alternative hypothesis. Our simulation studies corroborate these findings, highlighting the enhanced performance of our proposed methodologies in high-dimensional sphericity testi

stat.ME↗

Double Robust high dimensional alpha test for linear factor pricing model

In this paper, we investigate alpha testing for high-dimensional linear factor pricing models. We propose a spatial sign-based max-type test to handle sparse alternative cases. Additionally, we prove that this test is asymptotically independent of the spatial-sign-based sum-type test proposed by Liu et al. (2023). Based on this result, we introduce a Cauchy Combination test procedure that combines both the max-type and sum-type tests. Simulation studies and real data applications demonstrate that the new proposed test procedure is robust not only for heavy-tailed distributions but also for the sparsity of the alternative hypothesis.

stat.ME↗

Time Series Generative Learning with Application to Brain Imaging Analysis

This paper focuses on the analysis of sequential image data, particularly brain imaging data such as MRI, fMRI, CT, with the motivation of understanding the brain aging process and neurodegenerative diseases. To achieve this goal, we investigate image generation in a time series context. Specifically, we formulate a min-max problem derived from the $f$-divergence between neighboring pairs to learn a time series generator in a nonparametric manner. The generator enables us to generate future images by transforming prior lag-k observations and a random vector from a reference distribution. With a deep neural network learned generator, we prove that the joint distribution of the generated sequence converges to the latent truth under a Markov and a conditional invariance condition. Furthermore, we extend our generation mechanism to a panel data scenario to accommodate multiple samples. The effectiveness of our mechanism is evaluated by generating real brain MRI sequences from the Alzheimer's Disease Neuroimaging Initiative. These generated image sequences can be used as data augmentation to enhance the performance of further downstream tasks, such as Alzheimer's disease detection.

stat.ML↗

Adaptive Strategy of Testing Alphas in High Dimensional Linear Factor Pricing Models

In recent years, there has been considerable research on testing alphas in high-dimensional linear factor pricing models. In our study, we introduce a novel max-type test procedure that performs well under sparse alternatives. Furthermore, we demonstrate that this new max-type test procedure is asymptotically independent from the sum-type test procedure proposed by Pesaran and Yamagata (2017). Building on this, we propose a Fisher combination test procedure that exhibits good performance for both dense and sparse alternatives.

stat.ME↗

Testing Independence Between High-Dimensional Random Vectors Using Rank-Based Max-Sum Tests

In this paper, we address the problem of testing independence between two high-dimensional random vectors. Our approach involves a series of max-sum tests based on three well-known classes of rank-based correlations. These correlation classes encompass several popular rank measures, including Spearman's $ρ$, Kendall's $τ$, Hoeffding's D, Blum-Kiefer-Rosenblatt's R and Bergsma-Dassios-Yanagimoto's $τ^*$.The key advantages of our proposed tests are threefold: (1) they do not rely on specific assumptions about the distribution of random vectors, which flexibility makes them available across various scenarios; (2) they can proficiently manage non-linear dependencies between random vectors, a critical aspect in high-dimensional contexts; (3) they have robust performance, regardless of whether the alternative hypothesis is sparse or dense.Notably, our proposed tests demonstrate significant advantages in various scenarios, which is suggested by extensive numerical results and an empirical application in RNA microarray analysis.

stat.ME↗

Testing Alpha in High Dimensional Linear Factor Pricing Models with Dependent Observations

In this study, we introduce three distinct testing methods for testing alpha in high dimensional linear factor pricing model that deals with dependent data. The first method is a sum-type test procedure, which exhibits high performance when dealing with dense alternatives. The second method is a max-type test procedure, which is particularly effective for sparse alternatives. For a broader range of alternatives, we suggest a Cauchy combination test procedure. This is predicated on the asymptotic independence of the sum-type and max-type test statistics. Both simulation studies and practical data application demonstrate the effectiveness of our proposed methods when handling dependent observations.

stat.ME↗

Adaptive Rank-based Tests for High Dimensional Mean Problems

The Wilcoxon signed-rank test and the Wilcoxon-Mann-Whitney test are commonly employed in one sample and two sample mean tests for one-dimensional hypothesis problems. For high-dimensional mean test problems, we calculate the asymptotic distribution of the maximum of rank statistics for each variable and suggest a max-type test. This max-type test is then merged with a sum-type test, based on their asymptotic independence offered by stationary and strong mixing assumptions. Our numerical studies reveal that this combined test demonstrates robustness and superiority over other methods, especially for heavy-tailed distributions.

stat.ME↗

Limit Law for the Maximum Interpoint Distance of High Dimensional Dependent Variables

In this paper, we considier the limiting distribution of the maximum interpoint Euclidean distance $M_n=\max _{1 \leq i<j \leq n}\left\|\boldsymbol{X}_i-\boldsymbol{X}_j\right\|$, where $\boldsymbol{X}_1, \boldsymbol{X}_2, \ldots, \boldsymbol{X}_n$ be a random sample coming from a $p$-dimensional population with dependent sub-gaussian components. When the dimension tends to infinity with the sample size, we proves that $M_n^2$ under a suitable normalization asymptotically obeys a Gumbel type distribution. The proofs mainly depend on the Stein-Chen Poisson approximation method and high dimensional Gaussian approximation.

math.PR↗

Fisher's combined probability test for cross-sectional independence in panel data models with serial correlation

Testing cross-sectional independence in panel data models is of fundamental importance in econometric analysis with high-dimensional panels. Recently, econometricians began to turn their attention to the problem in the presence of serial dependence. The existing procedure for testing cross-sectional independence with serial correlation is based on the sum of the sample cross-sectional correlations, which generally performs well when the alternative has dense cross-sectional correlations, but suffers from low power against sparse alternatives. To deal with sparse alternatives, we propose a test based on the maximum of the squared sample cross-sectional correlations. Furthermore, we propose a combined test to combine the p-values of the max based and sum based tests, which performs well under both dense and sparse alternatives. The combined test relies on the asymptotic independence of the max based and sum based test statistics, which we show rigorously. We show that the proposed max based and combined tests have attractive theoretical properties and demonstrate the superior performance via extensive simulation results. We apply the two new tests to analyze the weekly returns on the securities in the S\&P 500 index under the Fama-French three-factor model, and confirm the usefulness of the proposed combined test in detecting cross-sectional independence.

stat.ME↗

Asymptotic Independence of the Quadratic form and Maximum of Independent Random Variables with Applications to High-Dimensional Tests

This paper establishes the asymptotic independence between the quadratic form and maximum of a sequence of independent random variables. Based on this theoretical result, we find the asymptotic joint distribution for the quadratic form and maximum, which can be applied into the high-dimensional testing problems. By combining the sum-type test and the max-type test, we propose the Fisher's combination tests for the one-sample mean test and two-sample mean test. Under this novel general framework, several strong assumptions in existing literature have been relaxed. Monte Carlo simulation has been done which shows that our proposed tests are strongly robust to both sparse and dense data.

stat.ME↗

Rank Based Tests for High Dimensional White Noise

The development of high-dimensional white noise test is important in both statistical theories and applications, where the dimension of the time series can be comparable to or exceed the length of the time series. This paper proposes several distribution-free tests using the rank based statistics for testing the high-dimensional white noise, which are robust to the heavy tails and do not quire the finite-order moment assumptions for the sample distributions. Three families of rank based tests are analyzed in this paper, including the simple linear rank statistics, non-degenerate U-statistics and degenerate U-statistics. The asymptotic null distributions and rate optimality are established for each family of these tests. Among these tests, the test based on degenerate U-statistics can also detect the non-linear and non-monotone relationships in the autocorrelations. Moreover, this is the first result on the asymptotic distributions of rank correlation statistics which allowing for the cross-sectional dependence in high dimensional data.

math.ST↗

Adaptive Testing for Alphas in Conditional Factor Models with High Dimensional Assets

This paper focuses on testing for the presence of alpha in time-varying factor pricing models, specifically when the number of securities N is larger than the time dimension of the return series T. We introduce a maximum-type test that performs well in scenarios where the alternative hypothesis is sparse. We establish the limit null distribution of the proposed maximum-type test statistic and demonstrate its asymptotic independence from the sum-type test statistics proposed by Ma et al.(2020).Additionally, we propose an adaptive test by combining the maximum-type test and sum-type test, and we show its advantages under various alternative hypotheses through simulation studies and two real data applications.

stat.ME↗

Testing for high-dimensional white noise

Testing for multi-dimensional white noise is an important subject in statistical inference. Such test in the high-dimensional case becomes an open problem waiting to be solved, especially when the dimension of a time series is comparable to or even greater than the sample size. To detect an arbitrary form of departure from high-dimensional white noise, a few tests have been developed. Some of these tests are based on max-type statistics, while others are based on sum-type ones. Despite the progress, an urgent issue awaits to be resolved: none of these tests is robust to the sparsity of the serial correlation structure. Motivated by this, we propose a Fisher's combination test by combining the max-type and the sum-type statistics, based on the established asymptotically independence between them. This combination test can achieve robustness to the sparsity of the serial correlation structure,and combine the advantages of the two types of tests. We demonstrate the advantages of the proposed test over some existing tests through extensive numerical results and an empirical analysis.

stat.ME↗

Sparse Kronecker Product Decomposition: A General Framework of Signal Region Detection in Image Regression

This paper aims to present the first Frequentist framework on signal region detection in high-resolution and high-order image regression problems. Image data and scalar-on-image regression are intensively studied in recent years. However, most existing studies on such topics focused on outcome prediction, while the research on image region detection is rather limited, even though the latter is often more important. In this paper, we develop a general framework named Sparse Kronecker Product Decomposition (SKPD) to tackle this issue. The SKPD framework is general in the sense that it works for both matrices (e.g., 2D grayscale images) and (high-order) tensors (e.g., 2D colored images, brain MRI/fMRI data) represented image data. Moreover, unlike many Bayesian approaches, our framework is computationally scalable for high-resolution image problems. Specifically, our framework includes: 1) the one-term SKPD; 2) the multi-term SKPD; and 3) the nonlinear SKPD. We propose nonconvex optimization problems to estimate the one-term and multi-term SKPDs and develop path-following algorithms for the nonconvex optimization. The computed solutions of the path-following algorithm are guaranteed to converge to the truth with a particularly chosen initialization even though the optimization is nonconvex. Moreover, the region detection consistency could also be guaranteed by the one-term and multi-term SKPD. The nonlinear SKPD is highly connected to shallow convolutional neural networks (CNN), particular to CNN with one convolutional layer and one fully connected layer. Effectiveness of SKPDs is validated by real brain imaging data in the UK Biobank database.

cs.CV↗

Asymptotic Independence of the Sum and Maximum of Dependent Random Variables with Applications to High-Dimensional Tests

For a set of dependent random variables, without stationary or the strong mixing assumptions, we derive the asymptotic independence between their sums and maxima. Then we apply this result to high-dimensional testing problems, where we combine the sum-type and max-type tests and propose a novel test procedure for the one-sample mean test, the two-sample mean test and the regression coefficient test in high-dimensional setting. Based on the asymptotic independence between sums and maxima, the asymptotic distributions of test statistics are established. Simulation studies show that our proposed tests have good performance regardless of data being sparse or not. Examples on real data are also presented to demonstrate the advantages of our proposed methods.

stat.ME↗

Computationally efficient and data-adaptive changepoint inference in high dimension

High-dimensional changepoint inference that adapts to various change patterns has received much attention recently. We propose a simple, fast yet effective approach for adaptive changepoint testing. The key observation is that two statistics based on aggregating cumulative sum statistics over all dimensions and possible changepoints by taking their maximum and summation, respectively, are asymptotically independent under some mild conditions. Hence we are able to form a new test by combining the p-values of the maximum- and summation-type statistics according to their limit null distributions. To this end, we develop new tools and techniques to establish asymptotic distribution of the maximum-type statistic under a more relaxed condition on componentwise correlations among all variables than that in existing literature. The proposed method is simple to use and computationally efficient. It is adaptive to different sparsity levels of change signals, and is comparable to or even outperforms existing approaches as revealed by our numerical studies.

stat.ME↗

Projected Robust PCA with Application to Smooth Image Recovery

Most high-dimensional matrix recovery problems are studied under the assumption that the target matrix has certain intrinsic structures. For image data related matrix recovery problems, approximate low-rankness and smoothness are the two most commonly imposed structures. For approximately low-rank matrix recovery, the robust principal component analysis (PCA) is well-studied and proved to be effective. For smooth matrix problem, 2d fused Lasso and other total variation based approaches have played a fundamental role. Although both low-rankness and smoothness are key assumptions for image data analysis, the two lines of research, however, have very limited interaction. Motivated by taking advantage of both features, we in this paper develop a framework named projected robust PCA (PRPCA), under which the low-rank matrices are projected onto a space of smooth matrices. Consequently, a large class of image matrices can be decomposed as a low-rank and smooth component plus a sparse component. A key advantage of this decomposition is that the dimension of the core low-rank component can be significantly reduced. Consequently, our framework is able to address a problematic bottleneck of many low-rank matrix problems: singular value decomposition (SVD) on large matrices. Theoretically, we provide explicit statistical recovery guarantees of PRPCA and include classical robust PCA as a special case.

stat.ML↗