SearcharxivSearch

arXiv subjects

Kristoffer Hellton

Publications and source records attributed to Kristoffer Hellton.

3 recordsLinked to original sources

When and why are principal component scores a good tool for visualizing high-dimensional data?

Principal component analysis (PCA) is a popular dimension reduction technique often used to visualize high-dimensional data structures. In genomics, this can involve millions of variables, but only tens to hundreds of observations. Theoretically, such extreme high-dimensionality will cause biased or inconsistent eigenvector estimates, but in practice the principal component scores are used for visualization with great success. In this paper, we explore when and why the classical principal component scores can be used to visualize structures in high-dimensional data, even when there are few observations compared to the number of variables. Our argument is two-fold: First, we argue that eigenvectors related to pervasive signals will have eigenvalues scaling linearly with the number of variables. Second, we prove that for linearly increasing eigenvalues, the sample component scores will be scaled and rotated versions of the population scores, asymptotically. Thus the visual information of the sample scores will be unchanged, even though the sample eigenvectors are biased. In the case of pervasive signals, the principal component scores can be used to visualize the population structures, even in extreme high-dimensional situations.

math.ST

Multiple Model-Free Knockoffs

Model-free knockoffs is a recently proposed technique for identifying covariates that is likely to have an effect on a response variable. The method is an efficient method to control the false discovery rate in hypothesis tests for separate covariates. This paper presents a generalisation of the technique using multiple sets of model-free knockoffs. This is formulated as an open question in Candes et al. [4]. With multiple knockoffs, we are able to reduce the randomness in the knockoffs, making the result stronger. Since we use the same structure for generating all the knockoffs, the computational resources is far smaller than proportional with the number of knockoffs. We prove a bound on the asymptotic false discovery rate when the number of sets increases that is better then the published bounds for one set.

stat.ME

Integrative clustering of high-dimensional data with joint and individual clusters, with an application to the Metabric study

When measuring a range of different genomic, epigenomic, transcriptomic and other variables, an integrative approach to analysis can strengthen inference and give new insights. This is also the case when clustering patient samples, and several integrative cluster procedures have been proposed. Common for these methodologies is the restriction of a joint cluster structure, which is equal for all data layers. We instead present Joint and Individual Clustering (JIC), which estimates both joint and data type-specific clusters simultaneously, as an extension of the JIVE algorithm (Lock et. al, 2013). The method is compared to iCluster, another integrative clustering method, and simulations show that JIC is clearly advantageous when both individual and joint clusters are present. The method is used to cluster patients in the Metabric study, integrating gene expression data and copy number aberrations (CNA). The analysis suggests a division into three joint clusters common for both data types and seven independent clusters specific for CNA. Both the joint and CNA-specific clusters are significantly different with respect to survival, also when adjusting for age and treatment.

stat.ME