SearcharxivSearch

arXiv subjects

Hengjian Cui

Publications and source records attributed to Hengjian Cui.

4 recordsLinked to original sources

The Bag-and-Whisker Plot: A New Bagplot for Bivariate Data

The bagplot, also known as the "bag-and-bolster plot", is a notable extension of the boxplot from univariate to bivariate data. Although widely used, its practical application is hindered by two key limitations: the fixed inflation factor for outlier detection that does not adapt to the sample size, and the unstable convex hull used to visualize its fence. In this paper, we propose a new bagplot, namely the "bag-and-whisker plot", as an improvement method to address these limitations. Our framework recasts outlier detection as a multiple testing problem, yielding a data-adaptive fence that controls statistical error rates and enhances the reliability of outlier identification. To further resolve graphical instability, we introduce a refined visualization that abandons the convex hull (the bolster) with a direct rendering of the statistical fence, complemented by granular whiskers that effectively illustrate the data's spread. Extensive simulations and real-world data analyses demonstrate that our new bagplot exhibits superior adaptivity and robustness compared to the existing standard, and thus can be highly recommended for practical use. To increase the visibility of the work, a user-friendly R package named BagWhiskerPlot has been made publicly available on CRAN.

stat.ME

A Distribution-Free Test of Independence and Its Application to Variable Selection

Motivated by the importance of measuring the association between the response and predictors in high dimensional data, In this article, we propose a new mean variance test of independence between a categorical random variable and a continuous one based on mean variance index. The mean variance index is zero if and only if two variables are independent. Under the independence, we derive an explicit form of its asymptotic null distribution, which provides us with an efficient and fast way to compute the empirical p-value in practice. The number of classes of the categorical variable is allowed to diverge slowly to the infinity. It is essentially a rank test and thus distribution-free. No assumption on the distributions of two random variables is required and the test statistic is invariant under one-to-one transformations. It is resistent to heavy-tailed distributions and extreme values. We assess its performance by Monte Carlo simulations and demonstrate that the proposed test achieves a higher power in comparison with the existing tests. We apply the proposed MV test to a high dimensional colon cancer gene expression data to detect the significant genes associated with the tissue syndrome.

stat.ME

Partial Penalized Likelihood Ratio Test under Sparse Case

This work is concern with testing the low-dimensional parameters of interest with divergent dimensional data and variable selection for the rest under the sparse case. A consistent test via the partial penalized likelihood approach, called the partial penalized likelihood ratio test statistic is derived, and its asymptotic distributions under the null hypothesis and the local alternatives of order $n^{-1/2}$ are obtained under some regularity conditions. Meanwhile, the oracle property of the partial penalized likelihood estimator also holds. The proposed partial penalized likelihood ratio test statistic outperforms the full penalized likelihood ratio test statistic in term of size and power, and performs as well as the classical likelihood ratio test statistic. Moreover, the proposed method obtains the variable selection results as well as the p-values of testing. Numerical simulations and an analysis of Prostate Cancer data confirm our theoretical findings and demonstrate the promising performance of the proposed partial penalized likelihood in hypothesis testing and variable selection.

stat.ME

Depth weighted scatter estimators

General depth weighted scatter estimators are introduced and investigated. For general depth functions, we find out that these affine equivariant scatter estimators are Fisher consistent and unbiased for a wide range of multivariate distributions, and show that the sample scatter estimators are strong and \sqrtn-consistent and asymptotically normal, and the influence functions of the estimators exist and are bounded in general. We then concentrate on a specific case of the general depth weighted scatter estimators, the projection depth weighted scatter estimators, which include as a special case the well-known Stahel-Donoho scatter estimator whose limiting distribution has long been open until this paper. Large sample behavior, including consistency and asymptotic normality, and efficiency and finite sample behavior, including breakdown point and relative efficiency of the sample projection depth weighted scatter estimators, are thoroughly investigated. The influence function and the maximum bias of the projection depth weighted scatter estimators are derived and examined. Unlike typical high-breakdown competitors, the projection depth weighted scatter estimators can integrate high breakdown point and high efficiency while enjoying a bounded-influence function and a moderate maximum bias curve. Comparisons with leading estimators on asymptotic relative efficiency and gross error sensitivity reveal that the projection depth weighted scatter estimators behave very well overall and, consequently, represent very favorable choices of affine equivariant multivariate scatter estimators.

math.ST