SearcharxivSearch

arXiv subjects

Dale S. Kim

Publications and source records attributed to Dale S. Kim.

3 recordsLinked to original sources

The Correlation Thresholding Algorithm for Exploratory Factor Analysis

Exploratory factor analysis is often used in the social sciences to estimate potential measurement models. To do this, several important issues need to be addressed: (1) determining the number of factors, (2) learning constraints in the factor loadings, and (3) selecting a solution amongst rotationally equivalent choices. Traditionally, these issues are treated separately. This work examines the Correlation Thresholding (CT) algorithm, which uses a graph-theoretic perspective to solve all three simultaneously, from a unified framework. Despite this advantage, it relies on several assumptions that may not hold in practice. We discuss the implications of these assumptions and assess the sensitivity of the CT algorithm to them for practical use in exploratory factor analysis. This is examined over a series of simulation studies, as well as a real data example. The CT algorithm shows reasonable robustness against violating these assumptions and very competitive performance in comparison to other methods.

stat.ME

Structure Learning of Latent Factors via Clique Search on Correlation Thresholded Graphs

Despite the widespread application of latent factor analysis, existing methods suffer from the following weaknesses: requiring the number of factors to be known, lack of theoretical guarantees for learning the model structure, and nonidentifiability of the parameters due to rotation invariance properties of the likelihood. We address these concerns by proposing a fast correlation thresholding (CT) algorithm that simultaneously learns the number of latent factors and a rotationally identifiable model structure. Our novel approach translates this structure learning problem into the search for so-called independent maximal cliques in a thresholded correlation graph that can be easily constructed from the observed data. Our clique analysis technique scales well up to thousands of variables, while competing methods are not applicable in a reasonable amount of running time. We establish a finite-sample error bound and high-dimensional consistency for the structure learning of our method. Through a series of simulation studies and a real data example, we show that the CT algorithm is an accurate method for learning the structure of factor analysis models and is robust to violations of its assumptions.

stat.ME

A Hybrid EM Algorithm for Linear Two-Way Interactions with Missing Data

We study an EM algorithm for estimating product-term regression models with missing data. The study of such problems in the likelihood tradition has thus far been restricted to an EM algorithm method using full numerical integration. However, under most missing data patterns, we show that this problem can be solved analytically, and numerical approximations are only needed under specific conditions. Thus we propose a hybrid EM algorithm, which uses analytic solutions when available and approximate solutions only when needed. The theoretical framework of our algorithm is described herein, along with two numerical experiments using both simulated and real data. We show that our algorithm confers higher accuracy to the estimation process, relative to the existing full numerical integration method. We conclude with a discussion of applications, extensions, and topics of further research.

stat.ME