Do stable performance metrics guarantee stable model predictions?: An empirical investigation using the GUSTO-1 dataset
Model validation is routinely used to quantify and correct for overfitting in performance estimates, but it does not fully address the stability of models developed from slightly different training samples. We investigated whether stable discrimination implies stable individual predictions and compared logistic regression (LR) with artificial neural networks (ANNs). Using data from 40,830 participants in the GUSTO-I trial, with 30-day mortality as the outcome, we created six sample-size scenarios with varying events per variable (EPV). In each scenario, LR and ANN models were developed using eight predictors, and 200 bootstrap resamples were used to assess stability in discrimination, mean absolute prediction error (MAPE), and the classification instability index (CII). For LR, discrimination stabilized at an EPV of 44.50, whereas MAPE and CII did not stabilize until an EPV of 89.12. ANN discrimination was more volatile, with optimism stabilizing only at an EPV of 178.25, and ANN predictions were severely miscalibrated at an EPV of 4.50. Overall, discrimination stabilized earlier than individual predictions for both approaches, with greater instability and larger sample-size requirements for ANN. Model developers should therefore assess and report prediction-level stability, because discrimination metrics alone may not reflect the reliability of individual predictions.