SearcharxivSearch

arXiv subjects

Alexander Seitel

Publications and source records attributed to Alexander Seitel.

12 recordsLinked to original sources

Medical Imaging AI Competitions Lack Fairness

Benchmarking competitions are central to the development of artificial intelligence (AI) in medical imaging, defining performance standards and shaping methodological progress. However, it remains unclear whether these benchmarks provide data that are sufficiently representative, accessible, and reusable to support clinically meaningful AI. In this work, we assess fairness along two complementary dimensions: (1) whether challenge datasets capture the diversity of real-world clinical data, and (2) whether they are accessible and legally reusable in line with the FAIR principles. To address this question, we conducted a large-scale systematic study of 249 biomedical image analysis challenges comprising 458 tasks across 19 imaging modalities. Our findings reveal limited representation across challenge datasets with respect to geographic location, imaging modalities, and problem types, raising concerns about how well current benchmarks reflect real-world clinical diversity. Despite their widespread influence, challenge datasets were frequently constrained by restrictive or ambiguous access conditions, inconsistent or non-compliant licensing practices, and incomplete documentation, limiting reproducibility and long-term reuse. Together, these shortcomings expose foundational fairness limitations in our benchmarking ecosystem and highlight a disconnect between leaderboard success and clinical relevance.

cs.CV

Current validation practice undermines surgical AI development

Surgical data science (SDS) is rapidly advancing, yet clinical adoption of artificial intelligence (AI) in surgery remains limited, with inadequate validation as an important contributing factor. Existing validation practices often neglect the temporal and hierarchical structure of intraoperative videos, yielding misleading or clinically irrelevant results. We introduce a comprehensive catalogue of validation pitfalls in AI-based surgical video analysis, derived from a multi-stage Delphi process with 92 international experts. Pitfalls span three categories: (1) data, (2) metric selection/configuration, and (3) aggregation and reporting. A systematic review of surgical AI papers reveals that these pitfalls are widespread. Experiments on surgical video datasets show that ignoring temporal and hierarchical data structures can understate uncertainty, obscure critical failure modes, and alter algorithm rankings. To address these shortcomings, we provide consensus-based best practices compiled. Together, this work provides an evidence-based framework for rigorous validation of surgical video analysis algorithms, guiding benchmarking, reporting, regulatory review, and clinical translation.

q-bio.OT

Anthropomorphic tissue-mimicking phantoms for oximetry validation in multispectral optical imaging

Significance: Optical imaging of blood oxygenation (sO$_2$) can be achieved based on the differential absorption spectra of oxy- and deoxy-haemoglobin. A key challenge in realising clinical validation of the sO$_2$ biomarkers is the absence of reliable sO$_2$ reference standards, including test objects. Aim: To enable quantitative testing of multispectral imaging methods for assessment of sO$_2$ by introducing anthropomorphic phantoms with appropriate tissue-mimicking optical properties. Approach: We used the stable copolymer-in-oil base material to create physical anthropomorphic structures and optimised dyes to mimic the optical absorption of blood across a wide spectral range. Using 3D-printed phantom moulds generated from a magnetic resonance image of a human forearm, we moulded the material into an anthropomorphic shape. Using both reflectance hyperspectral imaging (HSI) and photoacoustic tomography (PAT), we acquired images of the forearm phantoms and evaluated the performance of linear spectral unmixing (LSU). Results: Based on 10 fabricated forearm phantoms with vessel-like structures featuring five distinct sO$_2$ levels (between 0 and 100%), we showed that the measured absorption spectra of the material correlated well with HSI and PAT data with a Pearson correlation coefficient consistently above 0.8. Further, the application of LSU enabled a quantification of the mean absolute error in sO$_2$ assessment with HSI and PAT. Conclusion: Our anthropomorphic tissue-mimicking phantoms hold potential to provide a robust tool for developing, standardising, and validating optical imaging of sO$_2$.

physics.med-ph

Unsupervised Domain Transfer with Conditional Invertible Neural Networks

Synthetic medical image generation has evolved as a key technique for neural network training and validation. A core challenge, however, remains in the domain gap between simulations and real data. While deep learning-based domain transfer using Cycle Generative Adversarial Networks and similar architectures has led to substantial progress in the field, there are use cases in which state-of-the-art approaches still fail to generate training images that produce convincing results on relevant downstream tasks. Here, we address this issue with a domain transfer approach based on conditional invertible neural networks (cINNs). As a particular advantage, our method inherently guarantees cycle consistency through its invertible architecture, and network training can efficiently be conducted with maximum likelihood training. To showcase our method's generic applicability, we apply it to two spectral imaging modalities at different scales, namely hyperspectral imaging (pixel-level) and photoacoustic tomography (image-level). According to comprehensive experiments, our method enables the generation of realistic spectral data and outperforms the state of the art on two downstream classification tasks (binary and multi-class). cINN-based domain transfer could thus evolve as an important method for realistic synthetic data generation in the field of spectral imaging and beyond.

eess.IV

Photoacoustic image synthesis with generative adversarial networks

Photoacoustic tomography (PAT) has the potential to recover morphological and functional tissue properties with high spatial resolution. However, previous attempts to solve the optical inverse problem with supervised machine learning were hampered by the absence of labeled reference data. While this bottleneck has been tackled by simulating training data, the domain gap between real and simulated images remains an unsolved challenge. We propose a novel approach to PAT image synthesis that involves subdividing the challenge of generating plausible simulations into two disjoint problems: (1) Probabilistic generation of realistic tissue morphology, and (2) pixel-wise assignment of corresponding optical and acoustic properties. The former is achieved with Generative Adversarial Networks (GANs) trained on semantically annotated medical imaging data. According to a validation study on a downstream task our approach yields more realistic synthetic images than the traditional model-based approach and could therefore become a fundamental step for deep learning-based quantitative PAT (qPAT).

eess.IV

Standard error estimation in meta-analysis of studies reporting medians

We consider the setting of an aggregate data meta-analysis of a continuous outcome of interest. When the distribution of the outcome is skewed, it is often the case that some primary studies report the sample mean and standard deviation of the outcome and other studies report the sample median along with the first and third quartiles and/or minimum and maximum values. To perform meta-analysis in this context, a number of approaches have recently been developed to impute the sample mean and standard deviation from studies reporting medians. Then, standard meta-analytic approaches with inverse-variance weighting are applied based on the (imputed) study-specific sample means and standard deviations. In this paper, we illustrate how this common practice can severely underestimate the within-study standard errors, which results in overestimation of between-study heterogeneity in random effects meta-analyses. We propose a straightforward bootstrap approach to estimate the standard errors of the imputed sample means. Our simulation study illustrates how the proposed approach can improve estimation of the within-study standard errors and between-study heterogeneity. Moreover, we apply the proposed approach in a meta-analysis to identify risk factors of a severe course of COVID-19.

stat.ME

Semantic segmentation of multispectral photoacoustic images using deep learning

Photoacoustic (PA) imaging has the potential to revolutionize functional medical imaging in healthcare due to the valuable information on tissue physiology contained in multispectral photoacoustic measurements. Clinical translation of the technology requires conversion of the high-dimensional acquired data into clinically relevant and interpretable information. In this work, we present a deep learning-based approach to semantic segmentation of multispectral photoacoustic images to facilitate image interpretability. Manually annotated photoacoustic {and ultrasound} imaging data are used as reference and enable the training of a deep learning-based segmentation algorithm in a supervised manner. Based on a validation study with experimentally acquired data from 16 healthy human volunteers, we show that automatic tissue segmentation can be used to create powerful analyses and visualizations of multispectral photoacoustic images. Due to the intuitive representation of high-dimensional information, such a preprocessing algorithm could be a valuable means to facilitate the clinical translation of photoacoustic imaging.

eess.IV

Surgical Data Science -- from Concepts toward Clinical Translation

Recent developments in data science in general and machine learning in particular have transformed the way experts envision the future of surgery. Surgical Data Science (SDS) is a new research field that aims to improve the quality of interventional healthcare through the capture, organization, analysis and modeling of data. While an increasing number of data-driven approaches and clinical applications have been studied in the fields of radiological and clinical data science, translational success stories are still lacking in surgery. In this publication, we shed light on the underlying reasons and provide a roadmap for future advances in the field. Based on an international workshop involving leading researchers in the field of SDS, we review current practice, key achievements and initiatives as well as available standards and tools for a number of topics relevant to the field, namely (1) infrastructure for data acquisition, storage and access in the presence of regulatory constraints, (2) data annotation and sharing and (3) data analytics. We further complement this technical perspective with (4) a review of currently available SDS products and the translational progress from academia and (5) a roadmap for faster clinical translation and exploitation of the full potential of SDS, based on an international multi-round Delphi process.

cs.CY

Tattoo tomography: Freehand 3D photoacoustic image reconstruction with an optical pattern

Purpose: Photoacoustic tomography (PAT) is a novel imaging technique that can spatially resolve both morphological and functional tissue properties, such as the vessel topology and tissue oxygenation. While this capacity makes PAT a promising modality for the diagnosis, treatment and follow-up of various diseases, a current drawback is the limited field-of-view (FoV) provided by the conventionally applied 2D probes. Methods: In this paper, we present a novel approach to 3D reconstruction of PAT data (Tattoo tomography) that does not require an external tracking system and can smoothly be integrated into clinical workflows. It is based on an optical pattern placed on the region of interest prior to image acquisition. This pattern is designed in a way that a tomographic image of it enables the recovery of the probe pose relative to the coordinate system of the pattern. This allows the transformation of a sequence of acquired PA images into one common global coordinate system and thus the consistent 3D reconstruction of PAT imaging data. Results: An initial feasibility study conducted with experimental phantom data and in vivo forearm data indicates that the Tattoo approach is well-suited for 3D reconstruction of PAT data with high accuracy and precision. Conclusion: In contrast to previous approaches to 3D ultrasound (US) or PAT reconstruction, the Tattoo approach neither requires complex external hardware nor training data acquired for a specific application. It could thus become a valuable tool for clinical freehand PAT.

physics.med-ph

Navigated interventions in the head and neck area: standardized assessment of a new handy field generator

Electromagnetic (EM) tracking enables localization of surgical instruments within the magnetic field emitted by an EM field generator (FG). Usually, the larger a FG is, the larger its tracking volume is. However, the company NDI (Northern Digital Inc., Waterloo, ON, Canada) recently introduced the Planar 10-11 FG, which combines a compact construction (97mm x 112mm x 31mm) with a relatively large, cylindrical tracking volume (diameter: 340mm, height: 340mm). Using the standardized assessment protocol of Hummel et al., the FG was tested with regard to its tracking accuracy and to its robustness with respect to external sources of disturbance. The mean positional error (5cm distance metric according to Hummel protocol) was 0.59mm, with a mean jitter of 0.26mm in the standard setup. The mean orientational error was found to be 0.10°. The highest positional error (4.82mm) due to metallic sources of disturbance was caused by the steel SST 303. In contrast, steel SST 416 caused the lowest positional error (0.10mm). Overall, the Planar 10-11 FG tends to achieve better tracking accuracy results compared to other NDI FGs. Due to its compact construction and portability, the FG could contribute to increased clinical use of EM tracking systems.

physics.med-ph

Open-source Tracked Ultrasound with Anser EMT

Image-guided Interventions (IGT) have shown a huge potential to improve medical procedures or even allow for new treatment options. Most ultrasound(US)-based IGT systems use electromagnetic (EM) tracking for localizing US probes and instruments. However, EM tracking is not always reliable in clinical settings because the EM field can be disturbed by medical equipment. So far, most researchers used and studied commercial EM trackers with their IGT systems which in turn limited the possibilities to customize the trackers in order minimize distortions and make the systems robust for clinical use. In light of current good scientific practice initiatives that increasingly request research to publish the source code corresponding to a paper, the aim of this work was to test the feasibility of using the open-source EM tracker Anser EMT for localizing US probes in a clinical US suite for the first time. The standardized protocol of Hummel et al. yielded a jitter of 0.1$\pm$0.1 mm and a position error of 1.1$\pm$0.7 mm, which is comparable to 0.1 mm and 1.0 mm of a commercial NDI Aurora system. The rotation error of Anser EMT was 0.15$\pm$0.16°, which is lower than at least 0.4° for the commercial tracker. We consider tracked US as feasible with Anser EMT if an accuracy of 1-2 mm is sufficient for a specific application.

physics.med-ph

Clickstream analysis for crowd-based object segmentation with confidence

With the rapidly increasing interest in machine learning based solutions for automatic image annotation, the availability of reference annotations for algorithm training is one of the major bottlenecks in the field. Crowdsourcing has evolved as a valuable option for low-cost and large-scale data annotation; however, quality control remains a major issue which needs to be addressed. To our knowledge, we are the first to analyze the annotation process to improve crowd-sourced image segmentation. Our method involves training a regressor to estimate the quality of a segmentation from the annotator's clickstream data. The quality estimation can be used to identify spam and weight individual annotations by their (estimated) quality when merging multiple segmentations of one image. Using a total of 29,000 crowd annotations performed on publicly available data of different object classes, we show that (1) our method is highly accurate in estimating the segmentation quality based on clickstream data, (2) outperforms state-of-the-art methods for merging multiple annotations. As the regressor does not need to be trained on the object class that it is applied to it can be regarded as a low-cost option for quality control and confidence analysis in the context of crowd-based image annotation.

cs.CV