Searcharxiv⌕ Search

arXiv subjects

Damjan Krstajic

Publications and source records attributed to Damjan Krstajic.

5 recordsLinked to original sources

Why comparing survival curves between two subgroups may be misleading

We analyse an issue when comparing survival curves between two subgroups. We show that there is a direct relationship between estimates of subgroups' survival at a time point and positive and negative predictive values in the binary classification settings. Our findings present a case where current methods of comparing survival curves between subgroups may be misleading. We think that this ought to be taken into account during the validation of prognostic diagnostic tests that predict two prognostic subgroups for a given disease or treatment, when the validation data set consists of censored data.

stat.ME↗

A critical assessment of conformal prediction methods applied in binary classification settings

In recent years there has been an increase in the number of scientific papers that suggest using conformal predictions in drug discovery. We consider that some versions of conformal predictions applied in binary settings are embroiled in pitfalls, not obvious at first sight, and that it is important to inform the scientific community about them. In the paper we first introduce the general theory of conformal predictions and follow with an explanation of the versions currently dominant in drug discovery research today. Finally, we provide cases for their critical assessment in binary classification settings.

stat.AP↗

Missed opportunities in large scale comparison of QSAR and conformal prediction methods and their applications in drug discovery

Recently Bosc et al. (J Cheminform 11(1): 4, 2019), published an article describing a case study that directly compares conformal predictions with traditional QSAR methods for large-scale predictions of target-ligand binding. We consider this study to be very important. Unfortunately, we have found several issues in the authors' approach as well as in the presentation of their findings.

stat.AP↗

Binary classification models with "Uncertain" predictions

Binary classification models which can assign probabilities to categories such as "the tissue is 75% likely to be tumorous" or "the chemical is 25% likely to be toxic" are well understood statistically, but their utility as an input to decision making is less well explored. We argue that users need to know which is the most probable outcome, how likely that is to be true and, in addition, whether the model is capable enough to provide an answer. It is the last case, where the potential outcomes of the model explicitly include "don't know" that is addressed in this paper. Including this outcome would better separate those predictions that can lead directly to a decision from those where more data is needed. Where models produce an "Uncertain" answer similar to a human reply of "don't know" or "50:50" in the examples we refer to earlier, this would translate to actions such as "operate on tumour" or "remove compound from use" where the models give a "more true than not" answer. Where the models judge the result "Uncertain" the practical decision might be "carry out more detailed laboratory testing of compound" or "commission new tissue analyses". The paper presents several examples where we first analyse the effect of its introduction, then present a methodology for separating "Uncertain" from binary predictions and finally, we provide arguments for its use in practice.

stat.AP↗

How real is the random censorship model in medical studies?

In survival analysis the random censorship model refers to censoring and survival times being independent of each other. It is one of the fundamental assumptions in the theory of survival analysis. We explain the reason for it being so ubiquitous, and we investigate its presence in medical studies. We differentiate two types of censoring in medical studies (dropout and administrative), and we explain their importance in examining the existence of the random censorship model. We show that in order to presume the random censorship model it is not enough to have a design study which conforms to it, but that one needs to provide evidence for its presence in the results. Blindly presuming the random censorship model might lead to the Kaplan-Meier estimator producing biased results, which might have serious consequences when estimating survival in medical studies.

stat.AP↗