SearcharxivSearch

arXiv subjects

Ingo Roeder

Publications and source records attributed to Ingo Roeder.

7 recordsLinked to original sources

Fair Distributed Machine Learning with Imbalanced Data as a Stackelberg Evolutionary Game

Decentralised learning enables the training of deep learning algorithms without centralising data sets, resulting in benefits such as improved data privacy, operational efficiency and the fostering of data ownership policies. However, significant data imbalances pose a challenge in this framework. Participants with smaller datasets in distributed learning environments often achieve poorer results than participants with larger datasets. Data imbalances are particularly pronounced in medical fields and are caused by different patient populations, technological inequalities and divergent data collection practices. In this paper, we consider distributed learning as an Stackelberg evolutionary game. We present two algorithms for setting the weights of each node's contribution to the global model in each training round: the Deterministic Stackelberg Weighting Model (DSWM) and the Adaptive Stackelberg Weighting Model (ASWM). We use three medical datasets to highlight the impact of dynamic weighting on underrepresented nodes in distributed learning. Our results show that the ASWM significantly favours underrepresented nodes by improving their performance by 2.713% in AUC. Meanwhile, nodes with larger datasets experience only a modest average performance decrease of 0.441%.

cs.LG

Automatic quality control framework for more reliable integration of machine learning-based image segmentation into medical workflows

Machine learning algorithms underpin modern diagnostic-aiding software, which has proved valuable in clinical practice, particularly in radiology. However, inaccuracies, mainly due to the limited availability of clinical samples for training these algorithms, hamper their wider applicability, acceptance, and recognition amongst clinicians. We present an analysis of state-of-the-art automatic quality control (QC) approaches that can be implemented within these algorithms to estimate the certainty of their outputs. We validated the most promising approaches on a brain image segmentation task identifying white matter hyperintensities (WMH) in magnetic resonance imaging data. WMH are a correlate of small vessel disease common in mid-to-late adulthood and are particularly challenging to segment due to their varied size, and distributional patterns. Our results show that the aggregation of uncertainty and Dice prediction were most effective in failure detection for this task. Both methods independently improved mean Dice from 0.82 to 0.84. Our work reveals how QC methods can help to detect failed segmentation cases and therefore make automatic segmentation more reliable and suitable for clinical practice.

eess.IV

A comparative study of semi- and self-supervised semantic segmentation of biomedical microscopy data

In recent years, Convolutional Neural Networks (CNNs) have become the state-of-the-art method for biomedical image analysis. However, these networks are usually trained in a supervised manner, requiring large amounts of labelled training data. These labelled data sets are often difficult to acquire in the biomedical domain. In this work, we validate alternative ways to train CNNs with fewer labels for biomedical image segmentation using. We adapt two semi- and self-supervised image classification methods and analyse their performance for semantic segmentation of biomedical microscopy images.

cs.CV

An interpretable automated detection system for FISH-based HER2 oncogene amplification testing in histo-pathological routine images of breast and gastric cancer diagnostics

Histo-pathological diagnostics are an inherent part of the everyday work but are particularly laborious and associated with time-consuming manual analysis of image data. In order to cope with the increasing diagnostic case numbers due to the current growth and demographic change of the global population and the progress in personalized medicine, pathologists ask for assistance. Profiting from digital pathology and the use of artificial intelligence, individual solutions can be offered (e.g. detect labeled cancer tissue sections). The testing of the human epidermal growth factor receptor 2 (HER2) oncogene amplification status via fluorescence in situ hybridization (FISH) is recommended for breast and gastric cancer diagnostics and is regularly performed at clinics. Here, we develop an interpretable, deep learning (DL)-based pipeline which automates the evaluation of FISH images with respect to HER2 gene amplification testing. It mimics the pathological assessment and relies on the detection and localization of interphase nuclei based on instance segmentation networks. Furthermore, it localizes and classifies fluorescence signals within each nucleus with the help of image classification and object detection convolutional neural networks (CNNs). Finally, the pipeline classifies the whole image regarding its HER2 amplification status. The visualization of pixels on which the networks' decision occurs, complements an essential part to enable interpretability by pathologists.

eess.IV

Domain specific cues improve robustness of deep learning based segmentation of ct volumes

Machine Learning has considerably improved medical image analysis in the past years. Although data-driven approaches are intrinsically adaptive and thus, generic, they often do not perform the same way on data from different imaging modalities. In particular Computed tomography (CT) data poses many challenges to medical image segmentation based on convolutional neural networks (CNNs), mostly due to the broad dynamic range of intensities and the varying number of recorded slices of CT volumes. In this paper, we address these issues with a framework that combines domain-specific data preprocessing and augmentation with state-of-the-art CNN architectures. The focus is not limited to optimise the score, but also to stabilise the prediction performance since this is a mandatory requirement for use in automated and semi-automated workflows in the clinical environment. The framework is validated with an architecture comparison to show CNN architecture-independent effects of our framework functionality. We compare a modified U-Net and a modified Mixed-Scale Dense Network (MS-D Net) to compare dilated convolutions for parallel multi-scale processing to the U-Net approach based on traditional scaling operations. Finally, we propose an ensemble model combining the strengths of different individual methods. The framework performs well on a range of tasks such as liver and kidney segmentation, without significant differences in prediction performance on strongly differing volume sizes and varying slice thickness. Thus our framework is an essential step towards performing robust segmentation of unknown real-world samples.

eess.IV

Statistical and mathematical modeling of spatiotemporal dynamics of stem cells

Statistical and mathematical modeling are crucial to describe, interpret, compare and predict the behavior of complex biological systems including the organization of hematopoietic stem and progenitor cells in the bone marrow environment. The current prominence of high-resolution and live-cell imaging data provides an unprecedented opportunity to study the spatiotemporal dynamics of these cells within their stem cell niche and learn more about aberrant, but also unperturbed, normal hematopoiesis. However, this requires careful quantitative statistical analysis of the spatial and temporal behavior of cells and the interaction with their microenvironment. Moreover, such quantification is a prerequisite for the construction of hypothesis-driven mathematical models that can provide mechanistic explanations by generating spatiotemporal dynamics that can be directly compared to experimental observations. Here, we provide a brief overview of statistical methods in analyzing spatial distribution of cells, cell motility, cell shapes and cellular genealogies. We also describe cell- based modeling formalisms that allow researchers to simulate emergent behavior in a multicellular system based on a set of hypothesized mechanisms. Together, these methods provide a quantitative workflow for the analytic and synthetic study of the spatiotemporal behavior of hematopoietic stem and progenitor cells.

q-bio.QM

Towards an understanding of lineage specification in hematopoietic stem cells: A mathematical model for the interaction of transcription factors GATA-1 and PU.1

In addition to their self-renewal capabilities, hematopoietic stem cells guarantee the continuous supply of fully differentiated, functional cells of various types in the peripheral blood. The process which controls differentiation into the different lineages of the hematopoietic system (erythroid, myeloid, lymphoid) is referred to as lineage specification. It requires a potentially multi-step decision sequence which determines the fate of the cells and their successors. It is generally accepted that lineage specification is regulated by a complex system of interacting transcription factors. However, the underlying principles controlling this regulation are currently unknown. Here, we propose a simple quantitative model describing the interaction of two transcription factors. This model is motivated by experimental observations on the transcription factors GATA-1 and PU.1, both known to act as key regulators and potential antagonists in the erythroid vs. myeloid differentiation processes of hematopoietic progenitor cells. We demonstrate the ability of the model to account for the observed switching behavior of a transition from a state of low expression of both factors (undifferentiated state) to the dominance of one factor (differentiated state). Depending on the parameter choice, the model predicts two different possibilities to explain the experimentally suggested, stem cell characterizing priming state of low level co-expression. Whereas increasing transcription rates are sufficient to induce differentiation in one scenario, an additional system perturbation (by stochastic fluctuations or directed impulses) of transcription factor levels is required in the other case.

q-bio.MN