SearcharxivSearch

arXiv subjects

Yuanxing Chen

Publications and source records attributed to Yuanxing Chen.

6 recordsLinked to original sources

Penalized Network Cross-Validation for Nested Models by Edge-Sampling

In the network literature, a wide range of statistical models has been proposed to exploit structural patterns in the data. Therefore, model selection between different models is a fundamental problem. However, systematic theoretical understanding remains limited when comparisons involve different model classes. To address this challenging issue, we propose a penalized edge-sampling cross-validation framework for nested network model selection. By incorporating a model complexity penalty into the evaluation process, our method effectively mitigates the overfitting tendency of cross-validation and adapts to varying model structures. This framework supports comparisons among widely used models, including stochastic block models (SBMs), degree-corrected SBMs (DCBMs), and graphon models, providing the first consistency guarantees for model selection across these settings to the best of our knowledge. Empirical evaluations, including both simulated data and the ``Political Books'' network, demonstrate that our method yields stable and accurate performance across various scenarios.

stat.ME

Local spectral clustering for heterogeneous clustering structures

Classical clustering methods typically assume that all informative features support a single latent partition of the observations. This assumption can be overly restrictive for modern high-dimensional data, where different subsets of features may encode distinct notions of similarity and induce heterogeneous sample partitions, while some features may contain no meaningful clustering information. We develop a frequentist framework for local clustering that simultaneously identifies feature groups and estimates the sample clustering structure associated with each group. Our approach represents each sample partition by a label-invariant clustering matrix and groups features according to their shared clustering structures, thereby reformulating local clustering as a feature-grouping, or clustering-of-clusterings, problem. Under a heterogeneous sub-Gaussian mixture model, we construct feature-specific Gaussian-kernel similarity matrices and propose a local spectral clustering procedure based on a clustering-matrix optimization criterion. The proposed method avoids explicit likelihood specification and Bayesian posterior computation, accommodates heterogeneous feature distributions, and permits the presence of non-informative features. Extensive simulations and applications further demonstrate the practical utility and superiority of the proposed approach.

stat.ME

Fused Spatial Latent Block Models for Co-Clustering

Spatial transcriptomics is a rapidly growing technique that captures gene expression together with spatial coordinates in intact tissue sections, enabling in situ mapping of transcriptional activity. This technology offers unprecedented opportunities to study tissue heterogeneity and spatial gene expression patterns. Uncovering the associations between spatially variable gene modules and spot types can advance our understanding of pathological mechanisms. However, rigorous statistical methods that exploit spatial information to achieve spatially coherent co-clustering of spots and genes are still lacking, and theoretical investigations in this direction remain limited. We propose a fused spatial latent block model (F-SpLBM). Our model uses the LBM to uncover co-expression patterns between spots and genes, penalized fusion to automatically determine the number of co-clusters, and the Potts model to incorporate spatial information. We establish that the fusion-based procedure recovers the true block structure with the misclassification rate converging at a super-polynomial rate. We also prove asymptotic normality of the parameter estimators and quantify the accuracy gain from spatial smoothing. Simulations and real-data analyses demonstrate that F-SpLBM yields spatially coherent and biologically interpretable clustering results.

stat.ME

Cross-Validation in Bipartite Networks

Bipartite networks, which encode interactions between two distinct types of entities, arise widely in applications and exhibit inherent asymmetry across node sets. Despite a growing literature on bipartite community detection, estimating community numbers $(K_1, K_2)$, a critical issue for bipartite network analysis, remains theoretically underdeveloped without any model selection consistency established, to our knowledge. Indeed, the inherent asymmetry and the two-dimensional parameter space with possibly drastically different $K_1$ and $K_2$ pose unique challenges that differ from unipartite cases. In particular, the candidate models may simultaneously overfit one node set while underfitting the other. To address these challenges, we propose Bipartite Cross-Validation (BCV), a penalized cross-validation framework that jointly selects $(K_1,K_2)$ in a fully data-driven manner. We establish the first model selection consistency for bipartite networks, notably accommodating the regime where the numbers of communities scale with the network size, revealing the intricate interplay between sparsity and model complexity. Simulations and real-data applications demonstrate strong finite-sample performance of BCV.

stat.ME

Federated Online Learning for Heterogeneous Multisource Streaming Data

Federated learning has emerged as an essential paradigm for distributed multi-source data analysis under privacy concerns. Most existing federated learning methods focus on the ``static" datasets. However, in many real-world applications, data arrive continuously over time, forming streaming datasets. This introduces additional challenges for data storage and algorithm design, particularly under high-dimensional settings. In this paper, we propose a federated online learning (FOL) method for distributed multi-source streaming data analysis. To account for heterogeneity, a personalized model is constructed for each data source, and a novel ``subgroup" assumption is employed to capture potential similarities, thereby enhancing model performance. We adopt the penalized renewable estimation method and the efficient proximal gradient descent for model training. The proposed method aligns with both federated and online learning frameworks: raw data are not exchanged among sources, ensuring data privacy, and only summary statistics of previous data batches are required for model updates, significantly reducing storage demands. Theoretically, we establish the consistency properties for model estimation, variable selection, and subgroup structure recovery, demonstrating optimal statistical efficiency. Simulations illustrate the effectiveness of the proposed method. Furthermore, when applied to the financial lending data and the web log data, the proposed method also exhibits advantageous prediction performance. Results of the analysis also provide some practical insights.

stat.ML

Heterogeneity-aware Clustered Distributed Learning for Multi-source Data Analysis

In diverse fields ranging from finance to omics, it is increasingly common that data is distributed and with multiple individual sources (referred to as ``clients'' in some studies). Integrating raw data, although powerful, is often not feasible, for example, when there are considerations on privacy protection. Distributed learning techniques have been developed to integrate summary statistics as opposed to raw data. In many of the existing distributed learning studies, it is stringently assumed that all the clients have the same model. To accommodate data heterogeneity, some federated learning methods allow for client-specific models. In this article, we consider the scenario that clients form clusters, those in the same cluster have the same model, and different clusters have different models. Further considering the clustering structure can lead to a better understanding of the ``interconnections'' among clients and reduce the number of parameters. To this end, we develop a novel penalization approach. Specifically, group penalization is imposed for regularized estimation and selection of important variables, and fusion penalization is imposed to automatically cluster clients. An effective ADMM algorithm is developed, and the estimation, selection, and clustering consistency properties are established under mild conditions. Simulation and data analysis further demonstrate the practical utility and superiority of the proposed approach.

stat.ME