SearcharxivSearch

arXiv subjects

Alexander Bartel

Publications and source records attributed to Alexander Bartel.

5 recordsLinked to original sources

Performance evaluation of deep learning models for image analysis: considerations for visual control and statistical metrics

Deep learning-based automated image analysis (DL-AIA) has been shown to outperform trained pathologists in tasks related to feature quantification. Related to these capacities the use of DL-AIA tools is currently extending from proof-of-principle studies to routine applications such as patient samples (diagnostic pathology), regulatory safety assessment (toxicologic pathology), and recurrent research tasks. To ensure that DL-AIA applications are safe and reliable, it is critical to conduct a thorough and objective generalization performance assessment (i.e., the ability of the algorithm to accurately predict patterns of interest) and possibly evaluate model robustness (i.e., the algorithm's capacity to maintain predictive accuracy on images from different sources). In this article, we review the practices for performance assessment in veterinary pathology publications by which two approaches were identified: 1) Exclusive visual performance control (i.e. eyeballing of algorithmic predictions) plus validation of the models application utilizing secondary performance indices, and 2) Statistical performance control (alongside the other methods), which requires a dataset creation and separation of an hold-out test set prior to model training. This article compares the strengths and weaknesses of statistical and visual performance control methods. Furthermore, we discuss relevant considerations for rigorous statistical performance evaluation including metric selection, test dataset image composition, ground truth label quality, resampling methods such as bootstrapping, statistical comparison of multiple models, and evaluation of model stability. It is our conclusion that visual and statistical evaluation have complementary strength and a combination of both provides the greatest insight into the DL model's performance and sources of error.

cs.CV

Nuclear Pleomorphism in Canine Cutaneous Mast Cell Tumors: Comparison of Reproducibility and Prognostic Relevance between Estimates, Manual Morphometry and Algorithmic Morphometry

Variation in nuclear size and shape is an important criterion of malignancy for many tumor types; however, categorical estimates by pathologists have poor reproducibility. Measurements of nuclear characteristics (morphometry) can improve reproducibility, but manual methods are time consuming. The aim of this study was to explore the limitations of estimates and develop alternative morphometric solutions for canine cutaneous mast cell tumors (ccMCT). We assessed the following nuclear evaluation methods for measurement accuracy, reproducibility, and prognostic utility: 1) anisokaryosis (karyomegaly) estimates by 11 pathologists; 2) gold standard manual morphometry of at least 100 nuclei; 3) practicable manual morphometry with stratified sampling of 12 nuclei by 9 pathologists; and 4) automated morphometry using a deep learning-based segmentation algorithm. The study dataset comprised 96 ccMCT with available outcome information. The study dataset comprised 96 ccMCT with available outcome information. Inter-rater reproducibility of karyomegaly estimates was low ($\kappa$ = 0.226), while it was good (ICC = 0.654) for practicable morphometry of the standard deviation (SD) of nuclear size. As compared to gold standard manual morphometry (AUC = 0.839, 95% CI: 0.701 - 0.977), the prognostic value (tumor-specific survival) of SDs of nuclear area for practicable manual morphometry (12 nuclei) and automated morphometry were high with an area under the ROC curve (AUC) of 0.868 (95% CI: 0.737 - 0.991) and 0.943 (95% CI: 0.889 - 0.996), respectively. This study supports the use of manual morphometry with stratified sampling of 12 nuclei and algorithmic morphometry to overcome the poor reproducibility of estimates.

cs.CV

Systematic Review of Methods and Prognostic Value of Mitotic Activity. Part 2: Canine Tumors

One of the most relevant prognostication tests for tumors is cellular proliferation, which is most commonly measured by the mitotic activity in routine tumor sections. The goal of this systematic review is to scholarly analyze the methods and prognostic relevance of histologically measuring mitotic activity in canine tumors. A total of 137 articles that correlated the mitotic activity in canine tumors with patient outcome were identified through a systematic (PubMed and Scopus) and manual (Google Scholar) literature search and eligibility screening process. These studies determined the mitotic count (MC, number of mitotic figures per tumor area) in 126 instances, presumably the mitotic count (method not specified) in 6 instances and the mitotic index (MI, proportion of mitotic figures per tumor cells) in 5 instances. A particularly high risk of bias was identified in the available details of the MC methods and statistical analysis, which often did not quantify the prognostic discriminative ability of the MC and only reported p-values. A significant association of the MC with survival was found in 72/109 (66%) studies. However, survival was evaluated by at least three studies in only 7 tumor types/groups, of which a prognostic relevance is apparent for mast cell tumors of the skin, cutaneous melanoma and soft tissue sarcoma of the skin. None of the studies on the MI found a prognostic relevance. Further studies on the MC and MI with standardized methods are needed to prove the prognostic benefit of this test for further tumor types.

q-bio.SC

Systematic Review of Methods and Prognostic Value of Mitotic Activity. Part 1: Feline Tumors

Increased proliferation is a key driver of tumorigenesis, and quantification of mitotic activity is a standard task for prognostication. The goal of this systematic review is scholarly analysis of all available references on mitotic activity in feline tumors, and to provide an overview of the measuring methods and prognostic value. A systematic literature search in PubMed and Scopus and a manual search in Google Scholar was conducted. All articles on feline tumors that correlated mitotic activity with patient outcome were identified. Data analysis revealed that of the eligible 42 articles, the mitotic count (MC, mitotic figures per tumor area) was evaluated in 39 instances and the mitotic index (MI, mitotic figures per tumor cells) in three instances. The risk of bias was considered high for most studies (26/42, 62%) based on small study populations, insufficient details of the MC/MI methods, and lack of statistical measures for diagnostic accuracy or effect on outcome. The MC/MI methods varied markedly between studies. A significant association of the MC with survival was determined in 21/29 (72%) studies, while one study found an inverse effect. There were three tumor types with at least four studies and a prognostic association was found in 5/6 studies on mast cell tumors, 5/5 on mammary tumors and 3/4 on soft tissue sarcomas. The MI was shown to correlate with survival by two research groups, however a comparison to the MC was not conducted. An updated systematic review will be needed with of new literature for different tumor types.

q-bio.SC

How Many Annotators Do We Need? -- A Study on the Influence of Inter-Observer Variability on the Reliability of Automatic Mitotic Figure Assessment

Density of mitotic figures in histologic sections is a prognostically relevant characteristic for many tumours. Due to high inter-pathologist variability, deep learning-based algorithms are a promising solution to improve tumour prognostication. Pathologists are the gold standard for database development, however, labelling errors may hamper development of accurate algorithms. In the present work we evaluated the benefit of multi-expert consensus (n = 3, 5, 7, 9, 11) on algorithmic performance. While training with individual databases resulted in highly variable F$_1$ scores, performance was notably increased and more consistent when using the consensus of three annotators. Adding more annotators only resulted in minor improvements. We conclude that databases by few pathologists and high label accuracy may be the best compromise between high algorithmic performance and time investment.

cs.CV