SearcharxivSearch

arXiv subjects

Yuke Shi

Publications and source records attributed to Yuke Shi.

3 recordsLinked to original sources

Gaussian Multiplier Bootstrap Procedure for the $k$th Largest Coordinate of High-Dimensional Statistics

We consider the problem of Gaussian multiplier bootstrap procedures for the $k$th largest statistics and functions of the top $k$ order statistics, which are commonly encountered in high-dimensional statistical inference. Such a problem has been studied previously for $k=1$ (i.e., maxima). However, in many applications, a general $k$ ($k\geq 1$) is of great interest. We provide the upper bounds for the errors between Gaussian approximations and Gaussian multiplier approximations. The dimension $p$ is allowed to be larger than the sample size $n$. The effectiveness of the proposed methods is demonstrated via the computer numerical results and a real-world data analysis.

math.ST

Gaussian Approximations for the $k$th coordinate of sums of random vectors

We consider the problem of Gaussian approximation for the $\kappa$th coordinate of a sum of high-dimensional random vectors. Such a problem has been studied previously for $\kappa=1$ (i.e., maxima). However, in many applications, a general $\kappa\geq1$ is of great interest, which is addressed in this paper. We make four contributions: 1) we first show that the distribution of the $\kappa$th coordinate of a sum of random vectors, $\boldsymbol{X}= (X_{1},\cdots,X_{p})^{\sf T}= n^{-1/2}\sum_{i=1}^n \boldsymbol{x}_{i}$, can be approximated by that of Gaussian random vectors and derive their Kolmogorov's distributional difference bound; 2) we provide the theoretical justification for estimating the distribution of the $\kappa$th coordinate of a sum of random vectors using a Gaussian multiplier procedure, which multiplies the original vectors with i.i.d. standard Gaussian random variables; 3) we extend the Gaussian approximation result and Gaussian multiplier bootstrap procedure to a more general case where $\kappa$ diverges; 4) we further consider the Gaussian approximation for a square sum of the first $d$ largest coordinates of $\boldsymbol{X}$. All these results allow the dimension $p$ of random vectors to be as large as or much larger than the sample size $n$.

math.ST

Distance-based regression analysis for measuring associations

Distance-based regression model, as a nonparametric multivariate method, has been widely used to detect the association between variations in a distance or dissimilarity matrix for outcomes and predictor variables of interest in genetic association studies, genomic analyses, and many other research areas. Based on it, a pseudo-$F$ statistic which partitions the variation in distance matrices is often constructed to achieve the aim. To the best of our knowledge, the statistical properties of the pseudo-$F$ statistic has not yet been well established in the literature. To fill this gap, we study the asymptotic null distribution of the pseudo-$F$ statistic and show that it is asymptotically equivalent to a mixture of chi-squared random variables. Given that the pseudo-$F$ test statistic has unsatisfactory power when the correlations of the response variables are large, we propose a square-root $F$-type test statistic which replaces the similarity matrix with its square root. The asymptotic null distribution of the new test statistic and power of both tests are also investigated. Simulation studies are conducted to validate the asymptotic distributions of the tests and demonstrate that the proposed test has more robust power than the pseudo-$F$ test. Both test statistics are exemplified with a gene expression dataset for a prostate cancer pathway. Keywords: Asymptotic distribution, Chi-squared-type mixture, Nonparametric test, Pseudo-$F$ test, Similarity matrix.

math.ST