SearcharxivSearch

arXiv subjects

Jan O. Bauer

Publications and source records attributed to Jan O. Bauer.

6 recordsLinked to original sources

Beyond Regularization: Inherently Sparse Principal Component Analysis

Sparse principal component analysis (sparse PCA) is a widely used technique for dimensionality reduction in multivariate analysis, addressing two key limitations of standard PCA. First, sparse PCA can be implemented in high-dimensional low sample size settings, such as genetic microarrays. Second, it improves interpretability as components are regularized to zero. However, over-regularization of sparse singular vectors can cause them to deviate greatly from the population singular vectors, potentially misrepresenting the data structure. Additionally, sparse singular vectors are often not orthogonal, resulting in shared information between components, which complicates the calculation of variance explained. To address these challenges, we propose a methodology for sparse PCA that reflects the inherent structure of the data matrix. Specifically, we identify uncorrelated submatrices of the data matrix, meaning that the covariance matrix exhibits a sparse block diagonal structure. Such sparse matrices commonly occur in high-dimensional settings. The singular vectors of such a data matrix are inherently sparse, which improves interpretability while capturing the underlying data structure. Furthermore, these singular vectors are orthogonal by construction, ensuring that they do not share information. We demonstrate the effectiveness of our method through simulations and provide real data applications. Supplementary materials for this article are available online.

stat.ME

Localized Functional Principal Component Analysis Based on Covariance Structure

Functional principal component analysis (FPCA) is a widely used technique in functional data analysis for identifying the primary sources of variation in a sample of random curves. The eigenfunctions obtained from standard FPCA typically have non-zero support across the entire domain. In applications, however, it is often desirable to analyze eigenfunctions that are non-zero only on specific portions of the original domain-and exhibit zero regions when little is contributed to a specific direction of variability-allowing for easier interpretability. Our method identifies sparse characteristics of the underlying stochastic process and derives localized eigenfunctions by mirroring these characteristics without explicitly enforcing sparsity. Specifically, we decompose the stochastic process into uncorrelated sub-processes, each supported on disjoint intervals. Applying FPCA to these sub-processes yields localized eigenfunctions that are naturally orthogonal. In contrast, approaches that enforce localization through penalization must additionally impose orthogonality. Moreover, these approaches can suffer from over-regularization, resulting in eigenfunctions and eigenvalues that deviate from the inherent structure of their population counterparts, potentially misrepresenting data characteristics. Our approach avoids these issues by preserving the inherent structure of the data. Moreover, since the sub-processes have disjoint supports, the eigenvalues associated to the localized eigenfunctions allow for assessing the importance of each sub-processes in terms of its contribution to the total explained variance. We illustrate the effectiveness of our method through simulations and real data applications. Supplementary material for this article is available online.

stat.ME

High-Dimensional Block Diagonal Covariance Structure Detection Using Singular Vectors

The assumption of independent subvectors arises in many aspects of multivariate analysis. In most real-world applications, however, we lack prior knowledge about the number of subvectors and the specific variables within each subvector. Yet, testing all these combinations is not feasible. For example, for a data matrix containing 15 variables, there are already 1 382 958 545 possible combinations. Given that zero correlation is a necessary condition for independence, independent subvectors exhibit a block diagonal covariance matrix. This paper focuses on the detection of such block diagonal covariance structures in high-dimensional data and therefore also identifies uncorrelated subvectors. Our nonparametric approach exploits the fact that the structure of the covariance matrix is mirrored by the structure of its eigenvectors. However, the true block diagonal structure is masked by noise in the sample case. To address this problem, we propose to use sparse approximations of the sample eigenvectors to reveal the sparse structure of the population eigenvectors. Notably, the right singular vectors of a data matrix with an overall mean of zero are identical to the sample eigenvectors of its covariance matrix. Using sparse approximations of these singular vectors instead of the eigenvectors makes the estimation of the covariance matrix obsolete. We demonstrate the performance of our method through simulations and provide real data examples. Supplementary materials for this article are available online.

stat.ME

Divisive Hierarchical Clustering of Variables Identified by Singular Vectors

In this work, we introduce a novel methodology for divisive hierarchical clustering. Our divisive (``top-down'') approach is motivated by the fact that agglomerative hierarchical clustering (``bottom-up''), which is commonly used for hierarchical clustering, is not the best choice for all settings. The proposed methodology approximates the similarity matrix by a block diagonal matrix to identify clusters. While divisively clustering $p$ elements involves evaluating $2^{p-1}-1$ possible splits, which makes the task computationally costly, this approximation effectively reduces this number to at most $p(p-1)$ candidates, ensuring computational feasibility. We elaborate on the methodology and describe the incorporation of linkage functions to assess distances between clusters. We further show that these distances are ultrametric, ensuring that the resulting hierarchical cluster structure can be uniquely represented by a dendrogram, with interpretable heights. Additionally, the proposed methodology exhibits the flexibility to also optimize objectives of other clustering methods, and it can outperform these. The methodology is also applicable for constructing balanced clusters. To validate the efficiency of our approach, we conduct simulation studies and analyze real-world data. Supplementary materials for this article can be accessed online.

stat.ME

Principal Loading Analysis

This paper proposes a tool for dimension reduction where the dimension of the original space is reduced: a Principal Loading Analysis (PLA). PLA is a tool to reduce dimensions by discarding variables. The intuition is that variables are dropped which distort the covariance matrix only by a little. Our method is introduced and an algorithm for conducting PLA is provided. Further, we give bounds for the noise arising in the sample case.

math.ST

Correlation Based Principal Loading Analysis

Principal loading analysis is a dimension reduction method that discards variables which have only a small distorting effect on the covariance matrix. We complement principal loading analysis and propose to rather use a mix of both, the correlation and covariance matrix instead. Further, we suggest to use rescaled eigenvectors and provide updated algorithms for all proposed changes.

stat.ME