SearcharxivSearch

arXiv subjects

André Beauducel

Publications and source records attributed to André Beauducel.

16 recordsLinked to original sources

How to improve the regression factor score predictor when individuals have different factor loadings

Previous research has shown that ignoring individual differences of factor loadings in conventional factor models may reduce the determinacy of factor score predictors. Therefore, the aim of the present study is to propose a heterogeneous regression factor score with larger determinacy than the conventional regression factor score when individuals have different factor loadings. First, a method for the estimation of individual loadings is proposed. The individual loading estimates are used to compute the heterogeneity-based regression factor score predictor. Then, a binomial test for loading heterogeneity of a factor is recommended to compute the heterogeneity-based regression factor score predictor only when the test is significant. Otherwise, the conventional regression factor score predictor should be used. A simulation study reveals that the heterogeneity-based regression factor score predictor has larger determinacy than the conventional regression factor score predictor in populations with substantial loading heterogeneity. An empirical example based on subsamples drawn randomly from a large sample of Big Five Markers indicates that the determinacy can be improved for the factor emotional stability when the heterogeneity-based regression factor score is computed.

stat.ME

How does the elimination of group mean-differences affect factor score determinacy?

The present study investigates to what degree the common variance of the factor score predictor with the original factor, i.e., the determinacy coefficient or the validity of the factor score predictor, depends on the mean-difference between groups. When mean-differences between groups in the factor score predictor are eliminated by means of covariance analysis, regression, or group specific norms, this may reduce the covariance of the factor score predictor with the common factor. It is shown that in a one-factor model with the same group mean-difference on all observed variables, the common factor cannot be distinguished from a common factor representing the group mean-difference. It is also shown that for common factor loadings equal or larger than .60, the elimination of a d = .50 mean-difference between two groups in the factor score predictor leads to only small decreases of the determinacy coefficient. A compensation-factor k is proposed allowing for the estimation of the number of additional observed variables necessary to recover the size of the determinacy coefficient before elimination of a group mean-difference. It turns out that for factor loadings equal or larger than .60 only a few additional items are needed in order to recover the initial determinacy coefficient after the elimination of moderate or large group mean-differences.

stat.AP

Robust oblique Target-rotation for small samples

Introduction: Oblique Target-rotation in the context of exploratory factor analysis is a relevant method for the investigation of the oblique independent clusters model. It was argued that minimizing single cross-loadings by means of target rotation may lead to large effects of sampling error on the target rotated factor solutions. Method: In order to minimize effects of sampling error on results of Target-rotation we propose to compute the mean cross-loadings for each block of salient loadings of the independent clusters model and to perform target rotation for the block-wise mean cross-loadings. The resulting transformation-matrix is than applied to the complete unrotated loading matrix in order to produce mean Target-rotated factors. Results: A simulation study based on correlated independent factor models revealed that mean oblique Target-rotation resulted in smaller negative bias of factor inter-correlations than conventional Target-rotation based on single loadings, especially when sample size was small and when the number of factors was large. An empirical example revealed that the similarity of Target-rotated factors computed for small subsamples with Target-rotated factors of the total sample was more pronounced for mean Target-rotation than for conventional Target-rotation. Discussion: Mean Target-rotation can be recommended in the context of oblique independent factor models, especially for small samples. An R-script and an SPSS-script for this form of Target-rotation are provided in the Appendix.

stat.ME

Bias of determinacy coefficients in confirmatory factor analysis based on categorical variables

The relevance of determinacy coefficients as indicators for the validity of factor score predictors has regularly been emphasized. Previous simulation studies revealed biased determinacy coefficients for factor score predictors based on categorical variables. Therefore, and because there are different possibilities to compute determinacy coefficients, the present study compared bias of determinacy coefficients for the best linear factor score predictor and for a correlation-preserving factor score predictor based on confirmatory factor models with observed variables with 2, 4, 6, and 8 categories and maximum likelihood estimation, diagonally weighted least squares estimation, and Bayesian estimation. Positive bias was found when data were based on variables with two categories, population factors were correlated, and when there were unmodeled cross-loadings. Based on the results, the correction for sampling error, the use of maximum likelihood or Bayesian parameters, and data with at least four categories are recommended to avoid overestimation of parameter-based determinacy coefficients.

stat.AP

The trade-off between factor score determinacy and the preservation of inter-factor correlations

Regression factor score predictors have the maximum factor score determinacy, i.e., the maximum correlation with the corresponding factor, but they do not have the same inter-correlations as the factors. As it might be useful to compute factor score predictors that have the same inter-correlations as the factors, correlation-preserving factor score predictors have been proposed. However, correlation-preserving factor score predictors have smaller correlations with the corresponding factors (factor score determinacy) than regression factor score predictors. Thus, higher factor score determinacy goes along with bias of the inter-correlations and unbiased inter-correlations go along with lower factor score determinacy. The aim of the present study was therefore to investigate the size of the trade-off between factor score determinacy and bias of inter-correlations by means of a simulation study. It turns out that under several conditions very small gains of factor score determinacy of the regression factor score predictor go along with a large bias of inter-correlations. Instead of using the regression factor score predictor by default, it is proposed to check whether substantial bias of inter-correlations can be avoided without substantial loss of factor score determinacy by using a correlation-preserving factor score predictor. A syntax that allows to compute correlation-preserving factor score predictors from regression factor score predictors and to compare factor score determinacy and inter-correlations of the factor score predictors is given in the Appendix.

stat.AP

Correlation-preserving mean plausible values as a basis for prediction in the context of Bayesian structural equation modeling

Mean plausible values can be computed when Bayesian structural equation modeling (BSEM) is performed. As mean plausible values do not preserve the inter-factor correlations, they yield path coefficients that are different from the estimated path coefficients of the model. As it might be of interest to perform exactly the same prediction on the level of plausible values that has been estimated by BSEM, correlation-preserving mean plausible values were proposed. An example for the computation of the correlation preserving mean plausible values is given and the corresponding syntax is given in the Appendix.

stat.AP

R-factor analysis of data generated by a combination of R- and Q-factors leads to biased loading estimates

Effects of performing R-factor analysis of observed variables based on population models comprising R- and Q-factors were investigated. It was noted that estimating a model comprising R- and Q-factors has to face loading indeterminacy beyond rotational indeterminacy. Although R-factor analysis of data based on a population model comprising R- and Q-factors is nevertheless possible, this may lead to model error. Accordingly, even in the population, the resulting R-factor loadings are not necessarily close estimates of the original population R-factor loadings. It was shown in a simulation study that large Q-factor variance induces an increase of the variation of R-factor loading estimates beyond chance level. The results indicate that performing R-factor analysis with data based on a population model comprising R- and Q-factors may result in substantial loading bias. Tests of the multivariate kurtosis of observed variables are proposed as an indicator of possible Q-factor variance in observed variables as a prerequisite for R-factor analysis.

stat.AP

Coefficients of factor score determinacy for mean plausible values of Bayesian factor analysis

In the context of Bayesian factor analysis, it is possible to compute mean plausible values, which might be used as covariates or predictors or in order to provide individual scores for the Bayesian latent variables. Previous simulation studies ascertained the validity of the plausible values by the mean squared difference of the plausible values and the generating factor scores. However, the generating factor scores are unknown in empirical studies so that an indicator that is solely based on model parameters is needed in order to evaluate the validity of factor score estimates in empirical studies. The coefficient of determinacy is based on model parameters and can be computed whenever Bayesian factor analysis is performed in empirical settings. Therefore, the central aim of the present simulation study was to compare the coefficient of determinacy based on model parameters with the correlation of mean plausible values with the generating factors. It was found that the coefficient of determinacy yields an acceptable estimate for the validity of mean plausible values. As for small sample sizes and a small salient loading size the coefficient of determinacy overestimates the validity, it is recommended to report the coefficient of determinacy together with a bias-correction in order to estimate the validity of mean plausible values in empirical settings.

stat.AP

Heterogeneous item populations across individuals: Consequences for the factor model, item inter-correlations, and scale validity

The paper is devoted to the consequences of blind random selection of items from different item populations that might be based on completely uncorrelated factors for item inter-correlations and corresponding factor loadings. Based on the model of essentially parallel measurements, we explore the consequences of presenting items from different populations across individuals and items from identical populations within each individual for the factor model and item inter-correlations in the total population of individuals. Moreover, we explore the consequences of presenting items from different as well as identical item populations across and within individuals. We show that correlations can be substantial in the total population of individuals even when -- in subpopulations of individuals -- items are drawn from populations with uncorrelated factors. In order to address this challenge for the validity of a scale, we propose a method that helps to detect whether item inter-correlations result from different item populations in different subpopulations of individuals and evaluate the method by means of a simulation study. Based on the analytical results and on results from a simulation study, we provide recommendations for the detection of subpopulations of individuals responding to items from different item populations.

stat.AP

Score Predictor Factor Analysis as model for the identification of single-item indicators

Score Predictor Factor Analysis (SPFA) was introduced as a method that allows to compute factor score predictors that are -- under some conditions -- more highly correlated with the common factors resulting from factor analysis than the factor score predictors computed from the common factor model. In the present study, we investigate SPFA as a model in its own rights. In order to provide a basis for this, the properties and the utility of SPFA factor score predictors and the possibility to identify single-item indicators in SPFA loading matrices were investigated. Regarding the factor score predictors, the main result is that the best linear predictor of the score predictor factor analysis has not only perfect determinacy but is also correlation preserving. Regarding the SPFA loadings it was found in a simulation study that five or more population factors that are represented by only one variable with a rather substantial loading can more accurately be identified by means of SPFA than with conventional factor analysis. Moreover, the percentage of correctly identified single-item indicators was substantially larger for SPFA than for the common factor model. It is therefore argued that SPFA is a tool that can be especially helpful when very short scales or single-item indicators are to be identified.

stat.AP

Score predictor factor analysis: Reproducing observed covariances by means of factor score predictors

The model implied by factor score predictors does not reproduce the non-diagonal elements of the observed covariance matrix as well as the factor loadings. It is therefore investigated whether it is possible to estimate factor loadings for which the model implied by the factor score predictors optimally reproduces the non-diagonal elements of the observed covariance matrix. Accordingly, loading estimates are proposed for which the model implied by the factor score predictors allows for a least-squares approximation of the non-diagonal elements of the observed covariance matrix. This estimation method is termed Score predictor factor analysis and algebraically compared with Minres factor analysis as well as principal component analysis. A population based and a sample based simulation study was performed in order to compare Score predictor factor analysis, Minres factor analysis, and principal component analysis. It turns out that the non-diagonal elements of the observed covariance matrix can more exactly be reproduced from the factor score predictors computed from Score predictor factor analysis than from the factor score predictors computed from Minres factor analysis and from principal components. Moreover, Score predictor factor analysis can be helpful to identify factors when the factor model does not perfectly fit to the data because of model error.

stat.AP

Differences of Type I error rates for ANOVA and Multilevel-Linear-Models using SAS and SPSS for repeated measures designs

To derive recommendations on how to analyze longitudinal data, we examined Type I error rates of Multilevel Linear Models (MLM) and repeated measures Analysis of Variance (rANOVA) using SAS and SPSS.We performed a simulation with the following specifications: To explore the effects of high numbers of measurement occasions and small sample sizes on Type I error, measurement occasions of m = 9 and 12 were investigated as well as sample sizes of n = 15, 20, 25 and 30. Effects of non-sphericity in the population on Type I error were also inspected: 5,000 random samples were drawn from two populations containing neither a within-subject nor a between-group effect. They were analyzed including the most common options to correct rANOVA and MLM-results: The Huynh-Feldt-correction for rANOVA (rANOVA-HF) and the Kenward-Roger-correction for MLM (MLM-KR), which could help to correct progressive bias of MLM with an unstructured covariance matrix (MLM-UN). Moreover, uncorrected rANOVA and MLM assuming a compound symmetry covariance structure (MLM-CS) were also taken into account. The results showed a progressive bias for MLM-UN for small samples which was stronger in SPSS than in SAS. Moreover, an appropriate bias correction for Type I error via rANOVA-HF and an insufficient correction by MLM-UN-KR for n < 30 were found. These findings suggest MLM-CS or rANOVA if sphericity holds and a correction of a violation via rANOVA-HF. If an analysis requires MLM, SPSS yields more accurate Type I error rates for MLM-CS and SAS yields more accurate Type I error rates for MLM-UN.

stat.AP

Varimax rotation based on gradient projection needs between 10 and more than 500 random start loading matrices for optimal performance

Gradient projection rotation (GPR) is a promising method to rotate factor or component loadings by different criteria. Since the conditions for optimal performance of GPR-Varimax are widely unknown, this simulation study investigates GPR towards the Varimax criterion in principal component analysis. The conditions of the simulation study comprise two sample sizes (n = 100, n = 300), with orthogonal simple structure population models based on four numbers of components (3, 6, 9, 12), with- and without Kaiser-normalization, and six numbers of random start loading matrices for GPR-Varimax rotation (1, 10, 50, 100, 500, 1,000). GPR-Varimax rotation always performed better when at least 10 random matrices were used for start loadings instead of the identity matrix. GPR-Varimax worked better for a small number of components, larger (n = 300) as compared to smaller (n = 100) samples, and when loadings were Kaiser-normalized before rotation. To ensure optimal (stationary) performance of GPR-Varimax in recovering orthogonal simple structure, we recommend using at least 10 iterations of start loading matrices for the rotation of up to three components and 50 iterations for up to six components. For up to nine components, rotation should be based on a sample size of at least 300 cases, Kaiser-normalization, and more than 50 different start loading matrices. For more than nine components, GPR-Varimax rotation should be based on at least 300 cases, Kaiser-normalization, and at least 500 different start loading matrices.

stat.CO

On optimal allocation of treatment/condition variance in principal component analysis

The allocation of a (treatment) condition-effect on the wrong principal component (misallocation of variance) in principal component analysis (PCA) has been addressed in research on event-related potentials of the electroencephalogram. However, the correct allocation of condition-effects on PCA components might be relevant in several domains of research. The present paper investigates whether different loading patterns at each condition-level are a basis for an optimal allocation of between-condition variance on principal components. It turns out that a similar loading shape at each condition-level is a necessary condition for an optimal allocation of between-condition variance, whereas a similar loading magnitude is not necessary.

stat.AP

Treating reflective indicators as causal-formative indicators in order to compute factor score estimates or unit-weighted scales

Individual scores on common factors are required in some applied settings (e.g., business and marketing settings). Common factors are based on reflective indicators, but their scores cannot unambiguously be determined. Therefore, factor score estimates and unit-weighted scales are used in order to provide individual scores. It is shown that these scores are based on treating the reflective indicators as if they were causal-formative indicators. This modification of the causal status of the indicators should be justified. Therefore, the fit of the models implied by factor score estimates and unit-weighted scales should be investigated in order to ascertain the validity of the scores.

stat.AP

A Schmid-Leiman based transformation resulting in perfect inter-correlations of three types of factor score predictors

Factor score predictors are to be computed when the individual scores on the factors are of interest. Conditions for a perfect inter-correlation of the regression/best linear factor score predictor, the best linear conditionally unbiased predictor, and the determinant best linear correlation-preserving predictor are presented. When these three types of factor score predictors are perfectly correlated for corresponding factors, the factor score predictors computed from one method will have the virtues of the factor score predictors computed from the other methods. A Schmid-Leiman based transformation for which the three types of factor score predictors are perfectly correlated for corresponding orthogonal factors is proposed.

stat.AP