arXiv · 2606.16460
Module-structured mixture factor models for molecular subtype discovery in transcriptomic data
Abstract
High-throughput gene expression data exhibit high dimensionality, complex intergene dependence, and pronounced biological heterogeneity across samples, presenting major challenges for unsupervised clustering and disease subtype discovery. We introduce a module-structured mixture factor model that combines finite mixture modeling with low-rank latent factor representations defined at the gene-module level. By explicitly modeling gene modules in both the mean and covariance structure, the proposed framework decomposes expression variability into global gene-specific effects, cluster-specific module-level shifts, latent dependence within modules, and gene-specific residual noise. An Expectation--Conditional Maximization algorithm is applied for parameter estimation, allowing stable and scalable inference in high-dimensional transcriptomic settings. This framework enables interpretable unsupervised identification of disease-associated molecular subtypes and phenotypic heterogeneity across two autoimmune diseases using a large clinical transcriptomic dataset.
Explore related subjects
Keep this discovery
Jinran Wu, Geoffrey J. McLachlan, Saumyadipta Pyne. 2026-06-15. Module-structured mixture factor models for molecular subtype discovery in transcriptomic data. https://arxiv.org/abs/2606.16460
Cite the original work for its findings. Save a collection to share your selection of sources.