SearcharxivSearch

arXiv subjects

Motonori Oka

Publications and source records attributed to Motonori Oka.

5 recordsLinked to original sources

Sparse Bayesian joint modal estimation for exploratory item factor analysis

This study presents a scalable Bayesian estimation algorithm for sparse estimation in exploratory item factor analysis based on a classical Bayesian estimation method, namely Bayesian joint modal estimation (BJME). BJME estimates the model parameters and factor scores that maximize the complete-data joint posterior density. The algorithm's scalability is achieved through an alternating optimization scheme that iteratively updates model parameters and latent variables. Simulation studies show that the proposed algorithm has high computational efficiency and accuracy in variable selection over latent factors and the recovery of the model parameters. Moreover, we conducted a real data analysis using large-scale data from a psychological assessment that targeted the Big Five personality traits. This result indicates that the proposed algorithm achieves computationally efficient parameter estimation and extracts the interpretable factor loading structure.

stat.ME

Boosting Stochastic Optimisation for High-dimensional Latent Variable Models

Latent variable models are widely used in social and behavioural sciences, including education, psychology, and political science. With the increasing availability of large and complex datasets, high-dimensional latent variable models have become more common. However, estimating such models via marginal maximum likelihood is computationally challenging because it requires evaluating a large number of high-dimensional integrals. Stochastic optimisation, which combines stochastic approximation and sampling techniques, has been shown to be effective. It iterates between sampling latent variables from their posterior distribution under current parameter estimates and updating the model parameters using an approximate stochastic gradient constructed from the latent variable samples. In this paper, we investigate strategies to improve the performance of stochastic optimisation for high-dimensional latent variable models. The improvement is achieved through two strategies: a Metropolis-adjusted Langevin sampler that uses the gradient of the negative complete-data log-likelihood to sample latent variables efficiently, and a minibatch gradient technique that uses only a subset of observations when sampling latent variables and constructing stochastic gradients. Our simulation studies show that combining these strategies yields the best overall performance among competitors. An application to a personality test with 30 latent dimensions further demonstrates that the proposed algorithm scales effectively to high-dimensional settings.

stat.CO

Scalable Bayesian Approach for the DINA Q-matrix Estimation Combining Stochastic Optimization and Variational Inference

Diagnostic classification models (DCMs) offer statistical tools to inspect the fined-grained attribute of respondents' strengths and weaknesses. However, the diagnosis accuracy deteriorates when misspecification occurs in the predefined item-attribute relationship, which is encoded into a Q-matrix. To prevent such misspecification, methodologists have recently developed several Bayesian Q-matrix estimation methods for greater estimation flexibility. However, these methods become infeasible in the case of large-scale assessments with a large number of attributes and items. In this study, we focused on the deterministic inputs, noisy "and" gate (DINA) model and proposed a new framework for the Q-matrix estimation to find the Q-matrix with the maximum marginal likelihood. Based on this framework, we developed a scalable estimation algorithm for the DINA Q-matrix by constructing an iteration algorithm that utilizes stochastic optimization and variational inference. The simulation and empirical studies reveal that the proposed method achieves high-speed computation, good accuracy, and robustness to potential misspecifications, such as initial value's choices and hyperparameter settings. Thus, the proposed method can be a useful tool for estimating a Q-matrix in large-scale settings.

stat.CO

Assessing the Performance of Diagnostic Classification Models in Small Sample Contexts with Different Estimation Methods

Fueled by the call for formative assessments, diagnostic classification models (DCMs) have recently gained popularity in psychometrics. Despite their potential for providing diagnostic information that aids in classroom instruction and students' learning, empirical applications of DCMs to classroom assessments have been highly limited. This is partly because how DCMs with different estimation methods perform in small sample contexts is not yet well-explored. Hence, this study aims to investigate the performance of respondent classification and item parameter estimation with a comprehensive simulation design that resembles classroom assessments using different estimation methods. The key findings are the following: (1) although the marked difference in respondent classification accuracy was not observed among the maximum likelihood (ML), Bayesian, and nonparametric methods, the Bayesian method provided slightly more accurate respondent classification in parsimonious DCMs than the ML method, and in complex DCMs, the ML method yielded the slightly better result than the Bayesian method; (2) while item parameter recovery was poor in both Bayesian and ML methods, the Bayesian method exhibited unstable slip values owing to the multimodality of their posteriors under complex DCMs, and the ML method produced irregular estimates that appear to be well-estimated due to a boundary problem under parsimonious DCMs.

stat.CO

Variational Bayesian Inference for a Polytomous-Attribute Saturated Diagnostic Classification Model with Parallel Computing

As a statistical tool to assist formative assessments in educational settings, diagnostic classification models (DCMs) have been increasingly used to provide diagnostic information regarding examinees' attributes. DCMs often adopt a dichotomous division such as the mastery and non-mastery of attributes to express the mastery states of attributes. However, many practical settings involve different levels of mastery states rather than a simple dichotomy in a single attribute. Although this practical demand can be addressed by polytomous-attribute DCMs, their computational cost in a Markov chain Monte Carlo estimation impedes their large-scale application due to the larger number of polytomous-attribute mastery patterns than that of binary-attribute ones. This study considers a scalable Bayesian estimation method for polytomous-attribute DCMs and developed a variational Bayesian (VB) algorithm for a polytomous-attribute saturated DCM -- a generalization of polytomous-attribute DCMs -- by building on the existing literature on polytomous-attribute DCMs and VB for binary-attribute DCMs. Furthermore, we proposed the configuration of parallel computing for the proposed VB algorithm to achieve better computational efficiency. Monte Carlo simulations revealed that our method exhibited the high performance in parameter recovery under a wide range of conditions. An empirical example is used to demonstrate the utility of our method.

stat.CO