SearcharxivSearch

arXiv subjects

Yonghan Zhang

Publications and source records attributed to Yonghan Zhang.

3 recordsLinked to original sources

Distributed Online Estimation of Spiked Eigenvalues with Adaptive Weighting under Persistent Aspect Ratio Heterogeneity

We study online estimation of spiked covariance eigenvalues from observations distributed across $L$ nodes with heterogeneous and persistent effective sample sizes. In the proportional high-dimensional regime, local Rayleigh statistics are deterministically distorted by node-specific aspect ratios $c_{\ell,t}=p/N^{\mathrm{eff}}_{\ell,t}$, and direct aggregation of uncorrected statistics converges to the wrong limit. We propose a correct-then-aggregate framework in which each node removes its deterministic bias via an inverse Rayleigh transfer map, and the server fuses corrected estimates using adaptive soft-max weights based on predictable fluctuation metrics, transmitting only $O(k)$ scalars per active node per round. We establish consistency and asymptotic normality of the global estimator, enabling valid online inference, and derive non-asymptotic bounds quantifying how accuracy improves with the number of nodes and their effective sample sizes. The adaptive weights achieve variance reduction comparable to oracle inverse-variance weighting, confirming the data-driven construction is nearly efficient. Simulation studies validate these properties. An application to cross-venue monitoring of a dominant market factor shows the method tracks systemic risk in real time while substantially reducing communication cost relative to a centralized pooled approach.

math.ST

Transfer Learning for Linear Discriminant Analysis with a Shared Classification Signal

This paper studies transfer learning for linear discriminant analysis in high-dimensional two-class classification. We consider one target domain and several source domains, where the mean difference in each domain is decomposed into a deterministic common component and a domain-specific random deviation. The common component represents a shared classification signal across domains, while the random deviation captures domain-specific heterogeneity. Under spiked covariance models, we derive deterministic limits for the target-domain Gaussian-calibrated error of weighted transfer classifiers under both homogeneous and heterogeneous covariance settings. These limits quantify the effects of the shared signal, domain-specific variation, dimension-to-sample-size ratios, and spike structures on transfer performance. They further lead to oracle transfer weights and consistent data-driven plug-in estimators. We also characterize the intercept bias induced by unbalanced target-domain class sample sizes and provide an asymptotically optimal correction.

stat.ME

Structural Effect and Spectral Enhancement of High-Dimensional Regularized Linear Discriminant Analysis

Regularized linear discriminant analysis (RLDA) is a widely used tool for classification and dimensionality reduction, but its performance in high-dimensional scenarios is inconsistent. Existing theoretical analyses of RLDA often lack clear insight into how data structure affects classification performance. To address this issue, we derive a non-asymptotic approximation of the misclassification rate and thus analyze the structural effect and structural adjustment strategies of RLDA. Based on this, we propose the Spectral Enhanced Discriminant Analysis (SEDA) algorithm, which optimizes the data structure by adjusting the spiked eigenvalues of the population covariance matrix. By developing a new theoretical result on eigenvectors in random matrix theory, we derive an asymptotic approximation on the misclassification rate of SEDA. The bias correction algorithm and parameter selection strategy are then obtained. Experiments on synthetic and real datasets show that SEDA achieves higher classification accuracy and dimensionality reduction compared to existing LDA methods.

stat.ML