SearcharxivSearch

arXiv subjects

Nadine Binder

Publications and source records attributed to Nadine Binder.

9 recordsLinked to original sources

Typical Healthcare Pathways as a Basis for Admixture Modeling of Patient Trajectories

Background: Understanding whether patients follow similar or distinct patterns of care is important for characterizing clinical practice, identifying patient subgroups, and supporting quality improvement. However, routine healthcare trajectories are difficult to compare directly because patients may differ in their diagnostic workup, treatment sequencing, timing of clinical events, and documentation practices. Despite this variation, trajectories often contain recurring patterns at the cohort level. Methods: To address this challenge, we present a framework that explicitly separates cohort-level typical pathway identification from patient-level inference. At the cohort level, we derive an interpretable representation of care processes using a rule-based algorithm to identify typical healthcare pathways, resulting in a compact pathway graph. These pathways are then modeled as Markov chains and used as structured components in an admixture model, allowing each patient to be represented as a probabilistic mixture of typical pathways rather than being assigned to a single pathway component. The resulting admixture weights provide a compact representation of patient trajectories for subgroup characterization. We further assess the stability of the identified pathways and inferred admixture representations across multiple train-test splits. Results: Across train-test splits, the framework demonstrated consistent pathway structures and patient-level mixture patterns. Applied to routine care data from prostate cancer patients undergoing radical prostatectomy, the framework identified interpretable care patterns and supported the identification of patient subgroups with similar clinical event patterns. Conclusion: Overall, the proposed framework provides an interpretable and stable approach for summarizing treatment pathways and characterizing patient subgroups in real-world practice.

stat.ME

Overcoming Selection Bias in Statistical Studies With Amortized Bayesian Inference

Selection bias arises when the probability that an observation enters a dataset depends on variables related to the quantities of interest, leading to systematic distortions in estimation and uncertainty quantification. For example, in epidemiological or survey settings, individuals with certain outcomes may be more likely to be included, resulting in biased prevalence estimates with potentially substantial downstream impact. Classical corrections, such as inverse-probability weighting or explicit likelihood-based models of the selection process, rely on tractable likelihoods, which limits their applicability in complex stochastic models with latent dynamics or high-dimensional structure. Simulation-based inference enables Bayesian analysis without tractable likelihoods but typically assumes missingness at random and thus fails when selection depends on unobserved outcomes or covariates. Here, we develop a bias-aware simulation-based inference framework that explicitly incorporates selection into neural posterior estimation. By embedding the selection mechanism directly into the generative simulator, the approach enables amortized Bayesian inference without requiring tractable likelihoods. This recasting of selection bias as part of the simulation process allows us to both obtain debiased estimates and explicitly test for the presence of bias. The framework integrates diagnostics to detect discrepancies between simulated and observed data and to assess posterior calibration. The method recovers well-calibrated posterior distributions across three statistical applications with diverse selection mechanisms, including settings in which likelihood-based approaches yield biased estimates. These results recast the correction of selection bias as a simulation problem and establish simulation-based inference as a practical and testable strategy for parameter estimation under selection bias.

stat.ML

Temporal Trends in Incidence of Dementia in a Birth Cohorts Analysis of the Framingham Heart Study

Background: Dementia leads to a high burden of disability and the number of dementia patients worldwide doubled between 1990 and 2016. Nevertheless, some studies indicated a decrease in dementia risk which may be due to a bias caused by conventional analysis methods that do not adequately account for missing disease information due to death. Methods: This study re-examines potential trends in dementia incidence over four decades in the Framingham Heart Study. We apply a multistate modeling framework tailored to interval-censored illness-death data and define three non-overlapping birth cohorts (1915-1924, 1925-1934, and 1935-1944). Trends are evaluated based on both dementia prevalence and dementia risk, using age as the underlying timescale. Additionally, age-conditional dementia probabilities stratified by sex are estimated. Results: A total of 731 out of 3828 individuals were diagnosed with dementia. The multistate model analysis revealed no temporal decline in dementia risk across birth cohorts, irrespective of sex. When stratified by sex and adjusted for education, women consistently exhibited higher lifetime age-conditional risks (46%-50%) than men (30%-34%) over the study period. Conclusions: We recommend using a combination of multistate approach and separation into birth cohorts to adequately estimate trends of disease risk in cohort studies as well as to communicate patient-relevant outcomes such age-conditional disease risks.

stat.AP

Testing similarity of competing risks models by comparing transition probabilities

Assessing whether patient populations exhibit comparable event dynamics is important for evaluating treatment equivalence, pooling cohorts and comparing clinical pathways. Existing similarity tests for competing risks models measure distances between transition intensities, which describe instantaneous event rates. In biomedical applications, similarity may be more naturally formulated through transition probabilities, which quantify cumulative event risks over a clinically relevant horizon. Assuming constant cause-specific transition intensities, we develop a framework for testing similarity based on a maximum-type distance between vectors of transition-probability functions. We propose a constrained parametric bootstrap test and establish asymptotic level control and consistency under administrative and independent exponential random right censoring. The constant-intensity formulation is motivated by small-data settings in which few events are observed and nonparametric estimators may be unstable. Simulations across sample sizes, censoring mechanisms and degrees of dissimilarity show that the proposed test can attain larger finite-sample rejection probabilities than an intensity-based benchmark under comparable alternatives. An application to routine prostate cancer data illustrates how the procedure identifies the smallest examined margin for which similarity of 90-day readmission-probability functions can be established under the fitted model. The method provides an interpretable and practically implementable basis for similarity assessment in parametric competing risks models.

stat.ME

Can synthetic data reproduce real-world findings in epidemiology? A replication study using adversarial random forests

Synthetic data holds substantial potential to address practical challenges in epidemiology due to restricted data access and privacy concerns. However, many current methods suffer from limited quality, high computational demands, and complexity for non-experts. Furthermore, common evaluation strategies for synthetic data often fail to directly reflect statistical utility and measure privacy risks sufficiently. Against this background, a critical underexplored question is whether synthetic data can reliably reproduce key findings from epidemiological research while preserving privacy. We propose adversarial random forests (ARF) as an efficient and convenient method for synthesizing tabular epidemiological data. To evaluate its performance, we replicated statistical analyses from six epidemiological publications covering blood pressure, anthropometry, myocardial infarction, accelerometry, loneliness, and diabetes, from the German National Cohort (NAKO Gesundheitsstudie), the Bremen STEMI Registry U45 Study, and the Guelph Family Health Study. We further assessed how dataset dimensionality and variable complexity affect the quality of synthetic data, and contextualized ARF's performance by comparison with commonly used tabular data synthesizers in terms of utility, privacy, generalisation, and runtime. Across all replicated studies, results on ARF-generated synthetic data consistently aligned with original findings. Even for datasets with relatively low sample size-to-dimensionality ratios, replication outcomes closely matched the original results across descriptive and inferential analyses. Reduced dimensionality and variable complexity further enhanced synthesis quality. ARF demonstrated favourable performance regarding utility, privacy preservation, and generalisation relative to other synthesizers and superior computational efficiency.

q-bio.QM

Small Data Explainer -- The impact of small data methods in everyday life

The emergence of breakthrough artificial intelligence (AI) techniques has led to a renewed focus on how small data settings, i.e., settings with limited information, can benefit from such developments. This includes societal issues such as how best to include under-represented groups in data-driven policy and decision making, or the health benefits of assistive technologies. We provide a conceptual overview, clarify the relationship between small data and big data, and identify common themes from exemplary case studies and application areas. Potential solutions are described in a more detailed technical overview of current data analysis and modelling techniques, highlighting contributions from different disciplines, such as knowledge-driven modelling from statistics and data-driven modelling from computer science. By linking application settings, conceptual contributions and specific techniques, we highlight what is already feasible and suggest what an agenda for fully leveraging small data might look like.

cs.CY

Challenges and proposed solutions in modeling multimodal medical data: A systematic review

Multimodal data modeling has emerged as a powerful approach in clinical research, enabling the integration of diverse data types such as imaging, genomics, wearable sensors, and electronic health records. Despite its potential to improve diagnostic accuracy and support personalized care, modeling such heterogeneous data presents significant technical challenges. This systematic review synthesizes findings from 69 studies to identify common obstacles, including missing modalities, limited sample sizes, dimensionality imbalance, interpretability issues, and finding the optimal fusion techniques. We highlight recent methodological advances, such as transfer learning, generative models, attention mechanisms, and neural architecture search that offer promising solutions. By mapping current trends and innovations, this review provides a comprehensive overview of the field and offers practical insights to guide future research and development in multimodal modeling for medical applications.

cs.LG

Testing similarity of parametric competing risks models for identifying potentially similar pathways in healthcare

The identification of similar patient pathways is a crucial task in healthcare analytics. A flexible tool to address this issue are parametric competing risks models, where transition intensities may be specified by a variety of parametric distributions, thus in particular being possibly time-dependent. We assess the similarity between two such models by examining the transitions between different health states. This research introduces a method to measure the maximum differences in transition intensities over time, leading to the development of a test procedure for assessing similarity. We propose a parametric bootstrap approach for this purpose and provide a proof to confirm the validity of this procedure. The performance of our proposed method is evaluated through a simulation study, considering a range of sample sizes, differing amounts of censoring, and various thresholds for similarity. Finally, we demonstrate the practical application of our approach with a case study from urological clinical routine practice, which inspired this research.

stat.ME

Similarity of competing risks models with constant intensities in an application to clinical healthcare pathways involving prostate cancer surgery

The recent availability of routine medical data, especially in a university-clinical context, may enable the discovery of typical healthcare pathways, i.e., typical temporal sequences of clinical interventions or hospital readmissions. However, such pathways are heterogeneous in a large provider such as a university hospital, and it is important to identify similar care pathways that can still be considered typical pathways. We understand the pathway as a temporal process with possible transitions from a single initial treatment state to hospital readmission of different types, which constitutes a competing risk setting. In this paper, we propose a multi-state model-based approach to uncover pathway similarity between two groups of individuals. We describe a new bootstrap procedure for testing the similarity of transition intensities from two competing risk models with constant transition intensities. In a large simulation study, we investigate the performance of our similarity approach with respect to different sample sizes and different similarity thresholds. The studies are motivated by an application from urological clinical routine and we show how the results can be transferred to the application example.

stat.ME