Searcharxiv⌕ Search

arXiv subjects

Huanchao Zhou

Publications and source records attributed to Huanchao Zhou.

6 recordsLinked to original sources

Favourable Missingness in Semi-Supervised Classification for Exponential Mixture Models

Semi-supervised classifiers are commonly trained from samples in which all features are observed but some class labels are missing. When label missingness is independent of the observed data, unavailable class memberships reduce Fisher information relative to a completely classified sample. We study a different regime in which the probability of label missingness depends on posterior classification uncertainty, so that the observed missing-label indicators can themselves carry information about the Bayes decision boundary. Building on the conditionally weighted information decomposition of Ahfock and McLachlan, we develop this phenomenon for a two-component exponential mixture. Although the exponential model is non-Gaussian, asymmetric, and supported on the positive half-line, its log-posterior odds remain linear in the feature. We derive Bayes' rule and its exact error rate, formulate entropy-logistic and squared-discriminant missingness mechanisms, and obtain the full partially classified likelihood. We then derive a decomposition of the Fisher information into the complete-data information, the conditionally weighted loss due to missing labels, and the information contributed by the missing labels. Numerical quadrature identifies regions in which the full likelihood classifier has asymptotic relative efficiency above or below one. Monte Carlo experiments with finite training samples broadly support the population calculations, with the largest departures from the asymptotic predictions occurring near the transition at which the relative efficiency crosses one.

stat.CO↗

Spectrally-Corrected and Regularized Linear Discriminant Analysis for Spiked Covariance Model

This paper proposes an improved linear discriminant analysis called spectrally-corrected and regularized LDA (SRLDA). This method integrates the design ideas of the sample spectrally-corrected covariance matrix and the regularized discriminant analysis. With the support of a large-dimensional random matrix analysis framework, it is proved that SRLDA has a linear classification global optimal solution under the spiked model assumption. According to simulation data analysis, the SRLDA classifier performs better than RLDA and ILDA and is closer to the theoretical classifier. Experiments on different data sets show that the SRLDA algorithm performs better in classification and dimensionality reduction than currently used tools.

stat.ML↗

Exact Separation of Eigenvalues of Large Dimensional Noncentral Sample Covariance Matrices

Let $ \bbB_n =\frac{1}{n}(\bbR_n + \bbT^{1/2}_n \bbX_n)(\bbR_n + \bbT^{1/2}_n \bbX_n)^* $ where $ \bbX_n $ is a $ p \times n $ matrix with independent standardized random variables, $ \bbR_n $ is a $ p \times n $ non-random matrix, representing the information, and $ \bbT_{n} $ is a $ p \times p $ non-random nonnegative definite Hermitian matrix. Under some conditions on $ \bbR_n \bbR_n^* $ and $ \bbT_n $, it has been proved that for any closed interval outside the support of the limit spectral distribution, with probability one there will be no eigenvalues falling in this interval for all $ p $ sufficiently large. The purpose of this paper is to carry on with the study of the support of the limit spectral distribution, and we show that there is an exact separation phenomenon: with probability one, the proper number of eigenvalues lie on either side of these intervals.

math.PR↗

Analysis of the limiting spectral distribution of large dimensional General information-plus-noise type matrices

In this paper, we derive the analytical behavior of the limiting spectral distribution of non-central covariance matrices of the "general information-plus-noise" type, as studied in [14]. Through the equation defining its Stieltjes transform, it is shown that the limiting distribution has a continuous derivative away from zero, the derivative being analytic wherever it is positive, and we show the determination criterion for its support. We also extend the result in [14] to allow for all possible ratios of row to column of the underlying random matrix.

math.ST↗

No Eigenvalues Outside the Support of the Limiting Spectral Distribution of Large Dimensional noncentral Sample Covariance Matrices

Let $ \bbB_n =\frac{1}{n}(\bbR_n + \bbT^{1/2}_n \bbX_n)(\bbR_n + \bbT^{1/2}_n \bbX_n)^* $, where $ \bbX_n $ is a $ p \times n $ matrix with independent standardized random variables, $ \bbR_n $ is a $ p \times n $ non-random matrix and $ \bbT_{n} $ is a $ p \times p $ non-random, nonnegative definite Hermitian matrix. The matrix $\bbB_n$ is referred to as the information-plus-noise type matrix, where $\bbR_n$ contains the information and $\bbT^{1/2}_n \bbX_n$ is the noise matrix with the covariance matrix $\bbT_{n} $. It is known that, as $ n \to \infty $, if $ p/n $ converges to a positive number, the empirical spectral distribution of $ \bbB_n $ converges almost surely to a nonrandom limit, under some mild conditions. In this paper, we prove that, under certain conditions on the eigenvalues of $ \bbR_n $ and $ \bbT_n $, for any closed interval outside the support of the limit spectral distribution, with probability one there will be no eigenvalues falling in this interval for all $ n $ sufficiently large.

math.PR↗

The limiting spectral distribution of large dimensional general information-plus-noise type matrices

Let $ X_{n} $ be $ n\times N $ random complex matrices, $R_{n}$ and $T_{n}$ be non-random complex matrices with dimensions $n\times N$ and $n\times n$, respectively. We assume that the entries of $ X_{n} $ are independent and identically distributed, $ T_{n} $ are nonnegative definite Hermitian matrices and $T_{n}R_{n}R_{n}^{*}= R_{n}R_{n}^{*}T_{n} $. The general information-plus-noise type matrices are defined by $C_{n}=\frac{1}{N}T_{n}^{\frac{1}{2}} \left( R_{n} +X_{n}\right) \left(R_{n}+X_{n}\right)^{*}T_{n}^{\frac{1}{2}} $. In this paper, we establish the limiting spectral distribution of the large dimensional general information-plus-noise type matrices $C_{n}$. Specifically, we show that as $n$ and $N$ tend to infinity proportionally, the empirical distribution of the eigenvalues of $C_{n}$ converges weakly to a non-random probability distribution, which is characterized in terms of a system of equations of its Stieltjes transform.

math.ST↗