SearcharxivSearch

arXiv subjects

Xiaoke Qin

Publications and source records attributed to Xiaoke Qin.

2 recordsLinked to original sources

Extending Cluster-Weighted Factor Analyzers for multivariate prediction and high-dimensional interpretability

Cluster-weighted factor analyzers (CWFA) are a versatile class of mixture models designed to estimate the joint distribution of a random vector that includes a response variable along with a set of explanatory variables. They are particularly valuable in situations involving high dimensionality. This paper enhances CWFA models in two notable ways. First, it enables the prediction of multiple response variables while considering their potential interactions. Second, it identifies factors associated with disjoint groups of explanatory variables, thereby improving interpretability. This development leads to the introduction of the multivariate cluster-weighted disjoint factor analyzers (MCWDFA) model. An alternating expectation-conditional maximization algorithm is employed for parameter estimation. The effectiveness of the proposed model is assessed through an extensive simulation study that examines various scenarios. The proposal is applied to crime data from the United States, sourced from the UCI Machine Learning Repository, with the aim of capturing potential latent heterogeneity within communities and identifying groups of socio-economic features that are similarly associated with factors predicting crime rates. Results provide valuable insights into the underlying structures influencing crime rates which may potentially be helpful for effective cluster-specific policymaking and social interventions.

stat.ME

Finite mixtures of matrix-variate Poisson-log normal distributions for three-way count data

Three-way data structures, characterized by three entities, the units, the variables and the occasions, are frequent in biological studies. In RNA sequencing, three-way data structures are obtained when high-throughput transcriptome sequencing data are collected for $n$ genes across $p$ conditions at $r$ occasions. Matrix variate distributions offer a natural way to model three-way data and mixtures of matrix variate distributions can be used to cluster three-way data. Clustering of gene expression data is carried out as means of discovering gene co-expression networks. In this work, a mixture of matrix variate Poisson-log normal distributions is proposed for clustering read counts from RNA sequencing. By considering the matrix variate structure, full information on the conditions and occasions of the RNA sequencing dataset is simultaneously considered, and the number of covariance parameters to be estimated is reduced. We propose three different frameworks for parameter estimation: a Markov chain Monte Carlo based approach, a variational Gaussian approximation based approach, and a hybrid approach. Various information criteria are used for model selection. The models are applied to both real and simulated data, and we demonstrate that the proposed approaches can recover the underlying cluster structure in both cases. In simulation studies where the true model parameters are known, our proposed approach shows good parameter recovery.

stat.ME