SearcharxivSearch

arXiv subjects

Kirk P. Vanacore

Publications and source records attributed to Kirk P. Vanacore.

2 recordsLinked to original sources

GEEPERs: Principal Stratification using Principal Scores and Stacked Estimating Equations

Principal stratification is a framework for making sense of causal effects conditioned on variables that may themselves have been affected by the treatment. For instance, in an evaluation of an educational intervention, some subjects in the treatment group may not fully utilize the intervention, and researchers may be interested in how this subgroup is affected. Most principal stratification estimators rely on strong structural or modeling assumptions and often require advanced statistical training to fit and evaluate, making them inaccessible to many applied researchers. In this paper, we introduce a new principal effect estimator for one-way noncompliance based on a binary indicator. Estimates may be computed using conventional regression methods (though the standard errors require a specialized sandwich estimator) and do not rely on distributional assumptions. We present a simulation study that demonstrates the novel method's greater robustness compared to popular alternatives and illustrate the method through a real-data analysis.

stat.ME

Modernizing Ground Truth: Four Shifts Toward Improving Reliability and Validity in AI in Education

Generative Artificial Intelligence (GenAI) is now widespread in education, yet the efficacy of GenAI systems remains constrained by the quality and interpretation of the labeled data used to train and evaluate them. Studies commonly report inter-rater reliability (IRR), often summarized by a single coefficient such as Cohen's kappa (k), as a gatekeeper to ``ground truth.'' We argue that many educational assessment and practice support settings include challenges, such as high-inference constructs, skewed label distributions, and temporally segmented multimodal data, which yield potential misapplication or misinterpretation of threshold-based heuristics for IRR. The growing use of large language models as annotators and judges introduces risks such as automation bias and circular validation. We propose four practical shifts for establishing ground truth: (1) treat IRR as a diagnostic signal to localize disagreement and refine constructs rather than a mechanical acceptance threshold (e.g., k > 0.8); (2) require transparent reporting of rater expertise, codebook development, reconciliation procedures, and segmentation rules; (3) mitigate risks in LLM annotation through bias audits and verification workflows; and (4) complement agreement statistics with validity and effectiveness evidence for the intended use, including uncertainty-aware labeling (e.g., assigning different labels to the same item to capture nuance), criterion-related checks (e.g., predictive tests to check if labels forecast the intended outcome), and close-the-loop evaluations of whether systems trained on these labels improve learning beyond a reasonable control. We illustrate these shifts through case studies of multimodal tutoring data and provide actionable recommendations toward strengthening the evidence base of labeled AIED datasets.

cs.CY