Searcharxiv⌕ Search

arXiv subjects

Hua-Hua Chang

Publications and source records attributed to Hua-Hua Chang.

3 recordsLinked to original sources

A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments

The rapid expansion of large-scale assessments and the growing adoption of automatic item generation have intensified concerns about incidental content redundancy, where construct-irrelevant elements such as wording or contextual framing become unintentionally repetitive across items. Traditional similarity metrics like BLEU or cosine similarity, often fail to capture the nuanced structural and semantic layers that drive perceived redundancy simultaneously. This study proposes a dual-dimensional framework for Automated Item Similarity Analysis (AISA) powered by Large Language Models (LLMs), operationalizing similarity through Structured Decomposition and Semantic Relatedness. Psychometric validation indicates that LLM-derived metrics align more closely with indicators of construct-irrelevant local dependence and yield more coherent item parameter groupings than traditional text-based measures. The framework is further evaluated through its application in Computerized Adaptive Testing (CAT). Simulations reveal that incorporating LLM-based similarity constraints into item selection improves estimation stability and reduces bias with minimal efficiency trade-offs, outperforming constraints based on conventional metrics. These findings highlight the potential of LLM-powered AISA to support scalable bank curation, content-aware test assembly, and experience-sensitive adaptive testing across diverse assessment contexts.

cs.AI↗

Sequential Design for Computerized Adaptive Testing that Allows for Response Revision

In computerized adaptive testing (CAT), items (questions) are selected in real time based on the already observed responses, so that the ability of the examinee can be estimated as accurately as possible. This is typically formulated as a non-linear, sequential, experimental design problem with binary observations that correspond to the true or false responses. However, most items in practice are multiple-choice and dichotomous models do not make full use of the available data. Moreover, CAT has been heavily criticized for not allowing test-takers to review and revise their answers. In this work, we propose a novel CAT design that is based on the polytomous nominal response model and in which test-takers are allowed to revise their responses at any time during the test. We show that as the number of administered items goes to infinity, the proposed estimator is (i) strongly consistent for any item selection and revision strategy and (ii) asymptotically normal when the items are selected to maximize the Fisher information at the current ability estimate and the number of revisions is smaller than the number of items. We also present the findings of a simulation study that supports our asymptotic results.

math.ST↗

Nonlinear sequential designs for logistic item response theory models with applications to computerized adaptive tests

Computerized adaptive testing is becoming increasingly popular due to advancement of modern computer technology. It differs from the conventional standardized testing in that the selection of test items is tailored to individual examinee's ability level. Arising from this selection strategy is a nonlinear sequential design problem. We study, in this paper, the sequential design problem in the context of the logistic item response theory models. We show that the adaptive design obtained by maximizing the item information leads to a consistent and asymptotically normal ability estimator in the case of the Rasch model. Modifications to the maximum information approach are proposed for the two- and three-parameter logistic models. Similar asymptotic properties are established for the modified designs and the resulting estimator. Examples are also given in the case of the two-parameter logistic model to show that without such modifications, the maximum likelihood estimator of the ability parameter may not be consistent.

math.ST↗