SearcharxivSearch

arXiv subjects

Christian Stock

Publications and source records attributed to Christian Stock.

3 recordsLinked to original sources

Assessing covariate-adjusted risk differences in small-sample clinical trials

Binary endpoints are common in clinical trials and conditional odds ratios have traditionally been used to assess treatment effects. However, the interpretation of odds ratios is difficult, they are non-collapsible, and conditional odds-ratios obtained from regression models additionally rely on modeling assumptions in order to be a relevant overall summary measure for the trial. As an alternative, risk differences have gained increasing prominence as a more interpretable, clinically meaningful and assumption-lean measure of treatment effects. This shift has also been motivated by new regulatory guidance, which emphasizes the relevance of marginal estimands and encourages covariate adjustment. Yet, covariate-adjusted inference for risk differences, particularly in smaller samples, has methodological subtleties and lacks well-established best practices. We conduct a simulation study comparing methods for estimating and testing risk differences in small-sample (N$\,\leq\,$150) randomized clinical trials with prognostic categorical baseline covariates, focusing on exact unconditional tests, Mantel-Haenszel methods, and $g$-computation (standardization) approaches. We find that several $g$-computation approaches exhibit inflated Type I error in very small samples when standard Wald-type inference is applied, whereas robust or penalized variants improve error control at the expense of power. Classical methods such as the Mantel-Haenszel and Suissa-Shuster tests remain robust but may forgo efficiency gains from covariate adjustment. Overall, our results suggest that misalignment between estimand and variance estimation may contribute to the Type I error inflation, beyond the impact of small sample size alone. Based on these results, we provide practical recommendations to guide method selection that align the estimand, variance estimation, and inferential target.

stat.ME

Why rankings of biomedical image analysis competitions should be interpreted with care

International challenges have become the standard for validation of biomedical image analysis methods. Given their scientific impact, it is surprising that a critical analysis of common practices related to the organization of challenges has not yet been performed. In this paper, we present a comprehensive analysis of biomedical image analysis challenges conducted up to now. We demonstrate the importance of challenges and show that the lack of quality control has critical consequences. First, reproducibility and interpretation of the results is often hampered as only a fraction of relevant information is typically provided. Second, the rank of an algorithm is generally not robust to a number of variables such as the test data used for validation, the ranking scheme applied and the observers that make the reference annotations. To overcome these problems, we recommend best practice guidelines and define open research questions to be addressed in the future.

cs.CV

Clickstream analysis for crowd-based object segmentation with confidence

With the rapidly increasing interest in machine learning based solutions for automatic image annotation, the availability of reference annotations for algorithm training is one of the major bottlenecks in the field. Crowdsourcing has evolved as a valuable option for low-cost and large-scale data annotation; however, quality control remains a major issue which needs to be addressed. To our knowledge, we are the first to analyze the annotation process to improve crowd-sourced image segmentation. Our method involves training a regressor to estimate the quality of a segmentation from the annotator's clickstream data. The quality estimation can be used to identify spam and weight individual annotations by their (estimated) quality when merging multiple segmentations of one image. Using a total of 29,000 crowd annotations performed on publicly available data of different object classes, we show that (1) our method is highly accurate in estimating the segmentation quality based on clickstream data, (2) outperforms state-of-the-art methods for merging multiple annotations. As the regressor does not need to be trained on the object class that it is applied to it can be regarded as a low-cost option for quality control and confidence analysis in the context of crowd-based image annotation.

cs.CV