Searcharxiv⌕ Search

arXiv subjects

Sunday A. Adetunji

Publications and source records attributed to Sunday A. Adetunji.

3 recordsLinked to original sources

A Leakage-Proof Benchmark and Conformal Selective Triage for Electrohysterogram-Based Preterm Birth Prediction

Preterm birth remains a major cause of neonatal morbidity and mortality worldwide. Electrohysterography (EHG), a noninvasive measure of uterine myoelectrical activity, has been studied for preterm-birth prediction, but performance estimates may be biased when segments from the same maternal record are split across training and validation data. We formalized the distinction between segment-level and patient-independent record-grouped validation, established a patient-independent benchmark on the Term-Preterm Electrohysterogram Database, and evaluated class-conditional conformal selective prediction. All 300 records (38 preterm) were analyzed across three prespecified regimes. A 92-feature elastic-net logistic model was evaluated with record-grouped nested cross-validation; preprocessing, model fitting, Platt calibration, and conformal estimation were confined to training data, and performance was calculated from strictly out-of-fold record-level predictions with 1,000 bootstrap resamples. Under patient-independent evaluation, AUROC was 0.493 (95% CI 0.467-0.520), AUPRC 0.122 (0.095-0.152), and Brier score 0.115 (0.094-0.137). AUROC was 0.514 at or before 26 weeks and 0.469 thereafter. At miscoverage alpha=0.10, marginal coverage was 0.897, abstention 72.7%, and singleton-prediction accuracy 0.624. These findings establish a record-separated reference benchmark for EHG prediction and a broader validation principle for segmented physiological data: the unit of resampling should correspond to the unit at which predictive performance is intended to generalize. Conformal prediction further quantifies when the available information supports a singleton classification and when uncertainty warrants deferral.

q-bio.QM↗

AI-Derived Reproductive Phenotypes and Explainable ML for Concurrent Early Multimorbidity in U.S. Women: NHANES 2017-March 2020

Background:Adverse reproductive history is a multisystemic risk factor, but evidence is constrained by isolated outcome studies, limited adjustment, and non-interpretable algorithmic models. We re-frame the estimand from prediction to concurrent risk classification and emphasize calibration, interpretability, and systematic error. Methods:We analyzed 1,602 U.S. women aged 20-44 years from NHANES 2017-March 2020 with reproductive-history variables, chronic-condition indicators, and PHQ-9 data. Restricted multimorbidity was defined as at least two of hypertension, hypercholesterolemia, cardiovascular disease, kidney disease, and kidney stones. Features were summarized using principal components analysis and k-means clustering. We compared multivariable logistic regression with XGBoost and used SHAP values to quantify contributions. Results:Early multimorbidity occurred in 6.6% (106/1,602); 71.0% had no chronic condition and 22.4% had one. Adverse reproductive burden was common: 58% had at least one adverse reproductive factor and 12.6% had three or more. Four latent phenotypes emerged (n=398, 508, 102, 594), including a fragile subgroup in which 77.5% met the multimorbidity definition. In holdout evaluation, XGBoost improved discrimination relative to logistic regression (ROC-AUC 0.766 vs 0.667), but showed worse probability accuracy and calibration (Brier 0.069 vs 0.059; expected calibration error 0.113 vs 0.037). Dominant drivers were age, PHQ-9 score, income-to-poverty ratio, race/ethnicity, education, and the adverse reproductive index. Conclusions: Principal components analysis and k-means phenotyping revealed that adverse reproductive life-course structure is strongly clustered with concurrent early multimorbidity in U.S. women aged 20-44 years. Although XGBoost improved discrimination, calibration and feature attribution remained essential for reliable translation into practice

q-bio.OT↗

Teaching Longitudinal Linear Mixed Models End-to-End: A Reproducible Case Study in Mouse Body-Weight Growth

Background: Linear mixed-effects models are central for analyzing longitudinal continuous data, yet many learners meet them as scattered formulas or software output rather than as a coherent workflow. There is a need for a single, reproducible case study that links questions, model building, diagnostics, and interpretation. Methods: We reanalyze a published mouse body-weight experiment with 31 mice in three groups weighed weekly for 12 weeks. After reshaping the data to long format and using profile plots to motivate linear time trends, we fit three random-intercept linear mixed models: a common-slope model, a fully interacted group-by-time model, and a parsimonious model with group-specific intercepts, a shared slope for two groups, and an extra slope for the third. Models are compared using maximum likelihood, AIC, BIC, and likelihood ratio tests, and linear contrasts are used to estimate group differences in weekly means and 12 week gains. Results: The parsimonious model fits as well as the fully interacted model and clearly outperforms the common-slope model, revealing small and similar gains in two groups and much steeper growth in the third, with highly significant contrasts for excess weight gain. Interpretation: This case study gives a complete, executable workflow for longitudinal linear mixed modeling, from raw data and exploratory plots through model selection, diagnostics, and targeted contrasts. By making explicit the mapping from scientific questions to model terms and estimable contrasts, and by providing R code and a stepwise checklist, it serves as a practical template for teaching and applied work in biostatistics, epidemiology, and related fields

stat.ME↗