SearcharxivSearch

arXiv subjects

Philipp Doebler

Publications and source records attributed to Philipp Doebler.

6 recordsLinked to original sources

Cross-Course Generalizability of SRL-Aligned Predictive Models Using Digital Learning Traces

STEM dropout rates remain high at universities, particularly in computer science programs with theory-intensive courses. Digital learning environments now capture rich behavioral data that could help identify struggling students early, yet the generalizability of data-driven prediction models across courses and institutions remains uncertain. Guided by self-regulated learning (SRL) theory, this study analyzed multimodal digital-trace data from three undergraduate theoretical computer science courses (N1 = 137, N2 = 104, N3 = 148) at two universities. Weekly SRL-aligned digital-trace indicators were modeled using Elastic Net, Random Forest, and XGBoost to evaluate predictive performance over time and across settings, and model calibration both within and across courses. Early prediction of at-risk students was feasible, with SRL-related behaviors such as time management, effort regulation, and sustained engagement emerging as key predictors. While Random Forest achieved the highest in-sample accuracy, Elastic Net generalized more robustly across contexts. Out-of-sample accuracy and calibration declined between institutions with different base rates, underscoring the contextual nature of predictive analytics in higher education. These findings suggest that digital learning traces enable early identification of at-risk students within courses, but generalizing predictive models beyond their original context requires caution, particularly if the at-risk rates differ between contexts.

cs.CY

Interpretable Prediction Rule Ensembles in the Presence of Missing Data

Prediction Rule Ensembles (PREs) are robust and interpretable statistical learning techniques with potential for predictive analytics, yet their efficacy in the presence of missing data is untested. This study uses multiple imputation to fill in missing values, but uses a data stacking approach instead of a traditional model pooling approach to combine the results. We perform a simulation study to compare imputation methods under realistic conditions, focusing on sample sizes of $N=200$ and $N=400$ across 1,000 replications. Evaluated techniques include multiple imputation by chained equations with predictive mean matching (MICE PMM), MICE with Random Forest (MICE RF), Random Forest imputation with the ranger algorithm (missRanger), and imputation using extreme gradient boosting (MIXGBoost), with results compared to listwise deletion. Because stacking multiple imputed datasets can overly complicate models, we additionally explore different coarsening levels to simplify and enhance the interpretability and performance of PRE models. Our findings highlight a trade-off between predictive performance and model complexity in selecting imputation methods. While MIXGBoost and MICE PMM yield high rule recovery rates, they also increase false positives in rule selection. In contrast, MICE RF and missRanger promote rule sparsity. MIXGBoost achieved the greatest MSE reduction, followed by MICE PMM, MICE RF, and missRanger. Avoiding too-course rounding of variables helps to reduce model size with marginal loss in performance. Listwise deletion has an adverse impact on model validity. Our results emphasize the importance of choosing suitable imputation techniques based on research goals and of advancing methods for handling missing data in statistical learning.

stat.AP

Adapting tree-based multiple imputation methods for multi-level data? A simulation study

When data have a hierarchical structure, such as students nested within classrooms, ignoring dependencies between observations can compromise the validity of imputation procedures. Standard tree-based imputation methods implicitly assume independence between observations, limiting their applicability in multilevel data settings. Although Multivariate Imputation by Chained Equations (MICE) is widely used for hierarchical data, it has limitations, including sensitivity to model specification and computational complexity. Alternative tree-based approaches have shown promise for individual-level data, but remain largely unexplored for hierarchical contexts. In this simulation study, we systematically evaluate the performance of novel tree-based methods--Chained Random Forests and Extreme Gradient Boosting (mixgb)--explicitly adapted for multi-level data by incorporating dummy variables indicating cluster membership. We compare these tree-based methods and their adapted versions with traditional MICE imputation in terms of coefficient estimation bias, type I error rates and statistical power, under different cluster sizes, missingness mechanisms and missingness rates, using both random intercept and random slope data generation models. The results show that MICE provides robust and accurate inference for level 2 variables, especially at low missingness rates. However, the adapted boosting approach (mixgb with cluster dummies) consistently outperforms other methods for Level-1 variables at higher missingness rates (30%, 50%). For level 2 variables, while MICE retains better power at moderate missingness (30%), adapted boosting becomes superior at high missingness (50%), regardless of the missingness mechanism or cluster size. These findings highlight the potential of appropriately adapted tree-based imputation methods as effective alternatives to conventional MICE in multilevel data analyses.

stat.AP

Testing for Publication Bias in Diagnostic Meta-Analysis: A Simulation Study

The present study investigates the performance of several statistical tests to detect publication bias in diagnostic meta-analysis by means of simulation. While bivariate models should be used to pool data from primary studies in diagnostic meta-analysis, univariate measures of diagnostic accuracy are preferable for the purpose of detecting publication bias. In contrast to earlier research, which focused solely on the diagnostic odds ratio or its logarithm ($\ln\omega$), the tests are combined with four different univariate measures of diagnostic accuracy. For each combination of test and univariate measure, both type I error rate and statistical power are examined under diverse conditions. The results indicate that tests based on linear regression or rank correlation cannot be recommended in diagnostic meta-analysis, because type I error rates are either inflated or power is too low, irrespective of the applied univariate measure. In contrast, the combination of trim and fill and $\ln\omega$ has non-inflated or only slightly inflated type I error rates and medium to high power, even under extreme circumstances (at least when the number of studies per meta-analysis is large enough). Therefore, we recommend the application of trim and fill combined with $\ln\omega$ to detect funnel plot asymmetry in diagnostic meta-analysis. Please cite this paper as published in Statistics in Medicine (https://doi.org/10.1002/sim.6177).

stat.ME

Optimal design of the Wilcoxon-Mann-Whitney-test

In scientific research, many hypotheses relate to the comparison of two independent groups. Usually, it is of interest to use a design (i.e., the allocation of sample sizes $m$ and $n$ for fixed $N = m + n$) that maximizes the power of the applied statistical test. It is known that the two-sample t-tests for homogeneous and heterogeneous variances may lose substantial power when variances are unequal but equally large samples are used. We demonstrate that this is not the case for the non-parametric Wilcoxon-Mann-Whitney-test, whose application in biometrical research fields is motivated by two examples from cancer research. We prove the optimality of the design $m = n$ in case of symmetric and identically shaped distributions using normal approximations and show that this design generally offers power only negligibly lower than the optimal design for a wide range of distributions. Please cite this paper as published in the Biometrical Journal (https://doi.org/10.1002/bimj.201600022).

stat.ME

Fisher transformation based Confidence Intervals of Correlations in Fixed- and Random-Effects Meta-Analysis

Meta-analyses of correlation coefficients are an important technique to integrate results from many cross-sectional and longitudinal research designs. Uncertainty in pooled estimates is typically assessed with the help of confidence intervals, which can double as hypothesis tests for two-sided hypotheses about the underlying correlation. A standard approach to construct confidence intervals for the main effect is the Hedges-Olkin-Vevea Fisher-z (HOVz) approach, which is based on the Fisher-z transformation. Results from previous studies (Field, 2005; Hafdahl and Williams, 2009), however, indicate that in random-effects models the performance of the HOVz confidence interval can be unsatisfactory. To this end, we propose improvements of the HOVz approach, which are based on enhanced variance estimators for the main effect estimate. In order to study the coverage of the new confidence intervals in both fixed- and random-effects meta-analysis models, we perform an extensive simulation study, comparing them to established approaches. Data were generated via a truncated normal and beta distribution model. The results show that our newly proposed confidence intervals based on a Knapp-Hartung-type variance estimator or robust heteroscedasticity consistent sandwich estimators in combination with the integral z-to-r transformation (Hafdahl, 2009) provide more accurate coverage than existing approaches in most scenarios, especially in the more appropriate beta distribution simulation model.

stat.ME