SearcharxivSearch

arXiv subjects

Stephen H Friend

Publications and source records attributed to Stephen H Friend.

2 recordsLinked to original sources

Learning Disease vs Participant Signatures: a permutation test approach to detect identity confounding in machine learning diagnostic applications

Recently, Saeb et al (2017) showed that, in diagnostic machine learning applications, having data of each subject randomly assigned to both training and test sets (record-wise data split) can lead to massive underestimation of the cross-validation prediction error, due to the presence of "subject identity confounding" caused by the classifier's ability to identify subjects, instead of recognizing disease. To solve this problem, the authors recommended the random assignment of the data of each subject to either the training or the test set (subject-wise data split). The adoption of subject-wise split has been criticized in Little et al (2017), on the basis that it can violate assumptions required by cross-validation to consistently estimate generalization error. In particular, adopting subject-wise splitting in heterogeneous data-sets might lead to model under-fitting and larger classification errors. Hence, Little et al argue that perhaps the overestimation of prediction errors with subject-wise cross-validation, rather than underestimation with record-wise cross-validation, is the reason for the discrepancies between prediction error estimates generated by the two splitting strategies. In order to shed light on this controversy, we focus on simpler classification performance metrics and develop permutation tests that can detect identity confounding. By focusing on permutation tests, we are able to evaluate the merits of record-wise and subject-wise data splits under more general statistical dependencies and distributional structures of the data, including situations where cross-validation breaks down. We illustrate the application of our tests using synthetic and real data from a Parkinson's disease study.

stat.AP

Towards personalized causal inference of medication response in mobile health: an instrumental variable approach for randomized trials with imperfect compliance

Mobile health studies can leverage longitudinal sensor data from smartphones to guide the application of personalized medical interventions. In this paper, we propose that adoption of an instrumental variable approach for randomized trials with imperfect compliance provides a natural framework for personalized causal inference of medication response in mobile health studies. Randomized treatment suggestions can be easily delivered to the study participants via electronic messages popping up on the smart-phone screen. Under quite general assumptions we can identify the causal effect of the actual treatment on the response in the presence of unobserved confounders. We implement a personalized randomization test of the null hypothesis of no causal effect of treatment on response, and evaluate its performance in a large scale simulation study encompassing data generated from linear and non-linear time series models under several simulation conditions. In particular, we evaluate the empirical power of the proposed test under varying degrees of compliance between the suggested and actual treatment adopted by the participant. Our investigations provide encouraging results in terms of power and control of type I error rates. Finally, we compare the proposed instrumental variable approach to a simple intent-to-treat strategy, and develop randomization confidence intervals for the causal effects.

stat.AP