SearcharxivSearch

arXiv subjects

Jeff Eicher

Publications and source records attributed to Jeff Eicher.

3 recordsLinked to original sources

Measuring the Cross-Lingual Comprehension Gap: How the language of the evidence shapes what language models understand

Language models are often evaluated as though capabilities demonstrated in English remain equally available when the same content is presented in other languages. Traditional multilingual benchmarks rarely isolate language while holding content, question, reference answer, model, and evaluation unit constant. We define the Cross-Lingual Comprehension Gap (CLCG) as the reduction in response quality when the same content and question are presented in a target language rather than in English. Using ParallelQA-18, a professionally human-translated parallel corpus, we evaluate five models from five laboratories on a stratified sample of 150 articles across 18 languages (English reference; Portuguese high-resource baseline; 16 targets spanning Joshi et al. 2020 classes 0-4). A within-item design varies only passage language. The primary estimator contrasts English versus pooled target-language Token-F1 micro-means on higher-complexity open-ended questions, with article-cluster bootstrap intervals. The primary pooled CLCG is 0.078 (95% CI 0.072-0.084), about a 17% reduction relative to the English score; the equal-language macro summary is 0.077. Net of Portuguese, the macro gap is 0.016 (95% CI 0.013-0.020). Language-level CLCG is negatively associated with Joshi resource class (rho = -0.594, p = 0.015, n = 16). In blinded paired human evaluations, higher-resource responses are preferred in 61.6% of decisive judgments (estimated preference probability 0.655, 95% CI 0.558-0.741). Capabilities shown in English should not be assumed to transfer equally to other languages; English-centered evaluations may overestimate quality for users of low-resource languages.

cs.CL

A Unified Survival Benchmark for Temporal Dropout Risk Prediction in Learning Analytics

Student dropout is a persistent concern in Learning Analytics, yet comparative studies frequently evaluate predictive models under heterogeneous protocols, prioritizing discrimination over temporal interpretability and calibration. This study introduces a survival-oriented benchmark for temporal dropout risk modelling using the Open University Learning Analytics Dataset (OULAD). Two arms are compared: Family A: Dynamic Weekly, with models in person-period representation, and Family B: Static Early-Window, with an expanded roster of families: tree-based survival, parametric, and neural models. The evaluation protocol integrates four analytical layers: predictive performance, ablation, explainability, and calibration. Results are reported within each family separately, because a single numerical cross-family ranking would conflate genuine model differences with artifacts of temporal representation, to which survival metrics are known to be sensitive. Within Family B, Random Survival Forest showed the highest point estimates for time-dependent concordance and the lowest Brier scores across all three horizons; within Family A, Poisson Piecewise-Exponential showed the lowest point estimate for integrated Brier score within a tight five-model cluster. No-refit bootstrap resampling qualifies these positions as directional signals, not claims of strict superiority. Ablation and explainability analyses converged, across all models, on a shared finding: the dominant predictive signal was not primarily demographic or structural, but temporal and behavioral. Calibration corroborated this pattern in the better-discriminating models, except for XGBoost AFT, the sole outlier (analyzed in the Discussion). These results support unified, multi-dimensional benchmarking in Learning Analytics and situate dropout risk as a temporal-behavioral process rather than a function of static background attributes.

cs.LG

An Auditable Policy-Simulation Framework for Student Dropout in Intervention-Free Data

This study proposes a temporal modeling framework with a counterfactual policy-simulation layer for student dropout in higher education, using LMS engagement data and administrative withdrawal records. Dropout is operationalized as a time-to-event outcome at the enrollment level; weekly risk is modeled in discrete time via penalized, class-balanced logistic regression over person--period rows. Under a late-event temporal holdout, the model attains row-level AUCs of 0.8350 (train) and 0.8405 (test), with aggregate calibration acceptable but sparsely supported in the highest-risk bins. Ablation analyses indicate performance is sensitive to feature set composition, underscoring the role of temporal engagement signals. A scenario-indexed policy layer produces survival contrasts $\Delta S(T)$ under an explicit trigger/schedule contract: positive contrasts are confined to the shock branch ($T_{\rm policy}=18$: 0.0102, 0.0260, 0.0819), while the mechanism-aware branch is negative ($\Delta S_{\rm mech}(18)=-0.0078$, $\Delta S_{\rm mech}(38)=-0.0134$). A subgroup analysis by gender quantifies scenario-induced survival gaps via bootstrap; contrasts are directionally stable but small. Results are not causally identified; they demonstrate the framework's capacity for internal structural scenario comparison under observational data constraints.

cs.LG