SearcharxivSearch

arXiv subjects

Steven W. Nydick

Publications and source records attributed to Steven W. Nydick.

3 recordsLinked to original sources

Bayesian Consensus Calibration of Continuously Evolving IRT Item Banks

AI-based item generation and NLP-based prediction of item parameters are producing item banks that are substantially larger, sparser, and more frequently updated than conventional banks. Hierarchical Bayesian item response theory (IRT) is a natural calibration framework for such banks, but the common practice of refitting the entire accumulated response history at each update is costly and can exceed available memory. We describe \emph{consensus calibration}, a divide-and-conquer procedure that calibrates each time period independently and reconstructs the pooled posterior in two layers. First, the posterior draws of each period are mapped to a common metric by a robust characteristic-curve linking (Haebara) that is solved separately for each draw, which propagates the uncertainty of the linking transformation into the linked posteriors. Second, the linked item posteriors are combined as a product of Gaussian densities from which the population prior contributed by each period is removed and a single prior---obtained by consensus across the per-period population posteriors---is reinstated. The correction targets the posterior dispersion, not only its location. As evidence for consensus calibration, we compare it to a pooled single-run analysis on a large operational assessment in terms of item-parameter recovery, an uncertainty-by-exposure diagnostic, and the ability distributions.

stat.ME

Analytically Corrected Bayesian Modularization for Local Item Calibration

The continuous calibration of pilot items embedded in operational assessments is challenging when pilot samples are small and adaptively routed. We formalize a Bayesian modularization framework for local item calibration that blocks feedback from pilot responses to the operational latent scale: Plausible Values are drawn from the operational posterior, and each pilot item is calibrated by its own local logistic regression. This construction is computationally scalable, protects operational trait estimates from malfunctioning pilot items, and admits Firth's penalized likelihood for sparse routed samples. Because treating imputed traits as fixed predictors induces attenuation, we derive closed-form disattenuation mappings that recover the generating item parameters under normal-ogive, posterior-normality, homoscedasticity, and joint-normality approximations. The resulting Modular Local Calibration (MLC) estimator reaches the target of marginal-likelihood calibration without per-item numerical integration. A multivariate delta-method covariance combined with Rubin's-rules pooling propagates operational item-parameter uncertainty into the focal item standard errors. Monte Carlo simulations for the unidimensional 2PL show that corrected MLC substantially reduces attenuation bias and yields nominal-to-conservative interval coverage in the studied conditions, including under restricted-range MAR routing.

stat.ME

A Scalable Parametric Item Calibration Engine (SPICE) for Explanatory IRT with Sparse Data

We describe a Bayesian multidimensional explanatory IRT model, and an associated Markov Chain Monte Carlo (MCMC) estimation procedure and the corresponding development of calibration software, designed for psychometric analyses of large numbers of sparsely-linked persons and items. Such data structures can arise, for example, from adaptive assessments using large banks of automatically generated items with individual test takers receiving a very small proportion of the entire bank. We discuss how our choices for model specification, data structures, and algorithm implementation combine to create a scalable method for explanatory IRT that can support a variety of psychometric operations with sparse data.

stat.ME