SearcharxivSearch

arXiv subjects

Thomas Keil

Publications and source records attributed to Thomas Keil.

3 recordsLinked to original sources

Can synthetic data reproduce real-world findings in epidemiology? A replication study using adversarial random forests

Synthetic data holds substantial potential to address practical challenges in epidemiology due to restricted data access and privacy concerns. However, many current methods suffer from limited quality, high computational demands, and complexity for non-experts. Furthermore, common evaluation strategies for synthetic data often fail to directly reflect statistical utility and measure privacy risks sufficiently. Against this background, a critical underexplored question is whether synthetic data can reliably reproduce key findings from epidemiological research while preserving privacy. We propose adversarial random forests (ARF) as an efficient and convenient method for synthesizing tabular epidemiological data. To evaluate its performance, we replicated statistical analyses from six epidemiological publications covering blood pressure, anthropometry, myocardial infarction, accelerometry, loneliness, and diabetes, from the German National Cohort (NAKO Gesundheitsstudie), the Bremen STEMI Registry U45 Study, and the Guelph Family Health Study. We further assessed how dataset dimensionality and variable complexity affect the quality of synthetic data, and contextualized ARF's performance by comparison with commonly used tabular data synthesizers in terms of utility, privacy, generalisation, and runtime. Across all replicated studies, results on ARF-generated synthetic data consistently aligned with original findings. Even for datasets with relatively low sample size-to-dimensionality ratios, replication outcomes closely matched the original results across descriptive and inferential analyses. Reduced dimensionality and variable complexity further enhanced synthesis quality. ARF demonstrated favourable performance regarding utility, privacy preservation, and generalisation relative to other synthesizers and superior computational efficiency.

q-bio.QM

Subdomains of Post-COVID-Syndrome (PCS) -- A Population-Based Study

Post-COVID Syndrome (PCS), encompassing the multifaceted sequelae of COVID-19, can be severity-graded using a score comprising 12 different long-term symptom complexes. Acute COVID-19 severity and individual resilience were previously identified as key predictors of this score. This study validated these predictors and examined their relationship to PCS symptom complexes, using an expanded dataset (n=3,372) from the COVIDOM cohort study. Classification and Regression Tree (CART) analysis resolved the detailed relationship between the predictors and the constituting symptom complexes of the PCS score. Among newly recruited COVIDOM participants (n=1,930), the PCS score was again found to be associated with both its putative predictors. Of the score-constituting symptom complexes, neurological symptoms, sleep disturbance, and fatigue were predicted by individual resilience, whereas acute disease severity predicted exercise intolerance, chemosensory deficits, joint or muscle pain, signs of infection, and fatigue. These associations inspired the definition of two novel PCS scores that included the above-mentioned subsets of symptom complexes only. Both novel scores were inversely correlated with quality of life, measured by the EQ-5D-5L index. The newly defined scores may enhance the assessment of PCS severity, both in a research context and to delineate distinct PCS subdomains with different therapeutic and interventional needs in clinical practise.

q-bio.QM

SPINEPS -- Automatic Whole Spine Segmentation of T2-weighted MR images using a Two-Phase Approach to Multi-class Semantic and Instance Segmentation

Purpose. To present SPINEPS, an open-source deep learning approach for semantic and instance segmentation of 14 spinal structures (ten vertebra substructures, intervertebral discs, spinal cord, spinal canal, and sacrum) in whole body T2w MRI. Methods. During this HIPPA-compliant, retrospective study, we utilized the public SPIDER dataset (218 subjects, 63% female) and a subset of the German National Cohort (1423 subjects, mean age 53, 49% female) for training and evaluation. We combined CT and T2w segmentations to train models that segment 14 spinal structures in T2w sagittal scans both semantically and instance-wise. Performance evaluation metrics included Dice similarity coefficient, average symmetrical surface distance, panoptic quality, segmentation quality, and recognition quality. Statistical significance was assessed using the Wilcoxon signed-rank test. An in-house dataset was used to qualitatively evaluate out-of-distribution samples. Results. On the public dataset, our approach outperformed the baseline (instance-wise vertebra dice score 0.929 vs. 0.907, p-value<0.001). Training on auto-generated annotations and evaluating on manually corrected test data from the GNC yielded global dice scores of 0.900 for vertebrae, 0.960 for intervertebral discs, and 0.947 for the spinal canal. Incorporating the SPIDER dataset during training increased these scores to 0.920, 0.967, 0.958, respectively. Conclusions. The proposed segmentation approach offers robust segmentation of 14 spinal structures in T2w sagittal images, including the spinal cord, spinal canal, intervertebral discs, endplate, sacrum, and vertebrae. The approach yields both a semantic and instance mask as output, thus being easy to utilize. This marks the first publicly available algorithm for whole spine segmentation in sagittal T2w MR imaging.

eess.IV