Searcharxiv⌕ Search

arXiv subjects

Dandan Chen Kaptur

Publications and source records attributed to Dandan Chen Kaptur.

5 recordsLinked to original sources

Dependencies in Item-Adaptive CAT Data and Differential Item Functioning Detection: A Multilevel Framework

Differential item functioning (DIF) detection is an important yet understudied problem in computerized adaptive testing (CAT). In this article, we proposed a two-level logistic model to improve DIF detection in CAT by explicitly accounting for nuisance effects arising from CAT-induced structural dependency. First, we conceptualized that adaptive item selection induces systematic dependencies among examinees and items through provisional ability estimates, whereas traditional single-level DIF methods assume independent observations and may yield misleading results in CAT settings. Then, using a numeric example and Monte Carlo simulations, we compared our proposed two-level model with competing single-level models under various CAT conditions, manipulating test length, exposure control, ability estimator, DIF type, and DIF prevalence. Item-level Type-I error and statistical power conditional on joint model convergence were reported for each model. We showed that the proposed two-level model has improved control of spurious DIF and competitive power relative to single-level models, particularly with shorter tests and smaller exposure rates. However, we observed that the model convergence varied systematically across simulated conditions, highlighting that inferential accuracy and convergence reliability are intertwined in complex CAT DIF settings. Through this study, we underscored both the promise of multilevel DIF modeling in CAT and the need for future research to jointly evaluate convergence and inferential performance when assessing DIF models.

stat.AP↗

Enhancing Systematic Reviews with Large Language Models: Using GPT-4 and Kimi

This research delved into GPT-4 and Kimi, two Large Language Models (LLMs), for systematic reviews. We evaluated their performance by comparing LLM-generated codes with human-generated codes from a peer-reviewed systematic review on assessment. Our findings suggested that the performance of LLMs fluctuates by data volume and question complexity for systematic reviews.

cs.CL↗

Examining Differential Item Functioning (DIF) in Self-Reported Health Survey Data: Via Multilevel Modeling

Few health-related constructs or measures have received a critical evaluation in terms of measurement equivalence, such as self-reported health survey data. Differential item functioning (DIF) analysis is crucial for evaluating measurement equivalence in self-reported health surveys, which are often hierarchical in structure. Traditional single-level DIF methods in this case fall short, making multilevel models a better alternative. We highlight the benefits of multilevel modeling for DIF analysis, when applying a health survey data set to multilevel binary logistic regression (for analyzing binary response data) and multilevel multinominal logistic regression (for analyzing polytomous response data), and comparing them with their single-level counterparts. Our findings show that multilevel models fit better and explain more variance than single-level models. This article is expected to raise awareness of multilevel modeling and help healthcare researchers and practitioners understand the use of multilevel modeling for DIF analysis.

stat.AP↗

Examinees' Rapid-Guessing Patterns in Computerized Adaptive Testing for Interim Assessment: From Hierarchical Clustering

Interim assessment is frequently administered via computerized adaptive testing (CAT), offering direct support to teaching and learning. This study attempted to fill a vital knowledge gap about the nuanced landscape of examinees' rapid-guessing patterns in CAT in the interim assessment context. We analyzed a sample of 146,519 examinees in Grades 1-8 who participated in a widely used CAT, using hierarchical clustering, a robust data science methodology for uncovering insights in data. We found that examinees' rapid-guessing patterns varied across item positions, content domains, chronological grades, examinee clusters, and examinees' overall rapid-guessing level on the test, suggesting a nuanced interplay between testing features and examinees' behavior. Our study contributes to the literature on rapid guessing in CATs for interim assessment, offering a comprehensive and nuanced pattern analysis and demonstrating the application of hierarchical clustering to process data analysis in testing.

stat.AP↗

Evaluating Four Methods for Detecting Differential Item Functioning in Large-Scale Assessments with More Than Two Groups

This study evaluated four multi-group differential item functioning (DIF) methods (the root mean square deviation approach, Wald-1, generalized logistic regression procedure, and generalized Mantel-Haenszel method) via Monte Carlo simulation of controlled testing conditions. These conditions varied in the number of groups, the ability and sample size of the DIF-contaminated group, the parameter associated with DIF, and the proportion of DIF items. When comparing Type-I error rates and powers of the methods, we showed that the RMSD approach yielded the best Type-I error rates when it was used with model-predicted cutoff values. Also, this approach was found to be overly conservative when used with the commonly used cutoff value of 0.1. Implications for future research for educational researchers and practitioners were discussed.

stat.AP↗