Searcharxiv⌕ Search

arXiv subjects

Stefan Konigorski

Publications and source records attributed to Stefan Konigorski.

17 recordsLinked to original sources

Causal inference for N-of-1 trials

The aim of personalized medicine is to tailor treatment decisions to individuals' characteristics. N-of-1 trials are within-person crossover trials that hold the promise of targeting individual-specific effects. While the idea behind N-of-1 trials might seem simple, analyzing and interpreting N-of-1 trials is not straightforward. Here we ground N-of-1 trials in a formal causal inference framework and formalize intuitive claims from the N-of-1 trials literature. We focus on causal inference from a single N-of-1 trial and define a conditional average treatment effect (CATE) that represents a target in this setting, which we call the U-CATE. We discuss assumptions sufficient for identification and estimation of the U-CATE under different causal models where the treatment schedule is assigned at baseline. A simple mean difference is an unbiased, asymptotically normal estimator of the U-CATE in simple settings. We also consider settings where carryover effects, trends over time, time-varying common causes of the outcome, and outcome-outcome effects are present. In these more complex settings, we show that a time-varying g-formula identifies the U-CATE under explicit assumptions. Finally, we analyze data from N-of-1 trials about acne symptoms and show how different assumptions about the data generating process can lead to different analytical strategies.

stat.ME↗

Linear-LLM-SCM: Benchmarking LLMs for Coefficient Elicitation in Linear-Gaussian Causal Models

Large language models (LLMs) have shown potential in identifying qualitative causal relations, but their ability to perform quantitative causal reasoning---estimating effect sizes that parametrize functional relationships---remains underexplored in continuous domains. We introduce Linear-LLM-SCM, a plug-and-play framework for evaluating LLMs on Linear Gaussian structural causal model parametrization when a directed acyclic graph (DAG) is given. The framework decomposes a DAG into local parent-child sets and prompts an LLM to produce a regression-style structural equation per node, which is aggregated and compared against available ground-truth parameters. Our experiments with seven real-world DAGs effect ground truth illustrate limitations of LLMs as quantitative causal parameterizers. Across most models, we observe variability in coefficient estimates and sensitivity to structural perturbations. We open-sourced the framework to further encourage the community to work on studies toward the use of LLM for causal effect elicitation in safety-critical domain, e.g., healthcare.

cs.LG↗

Know You Before You Speak: User-State Modeling for LLM Personalization in Multi-Turn Conversation

Personalized dialogue requires more than recalling explicit user histories: systems also need to infer hidden user states that evolve through interaction and shape appropriate response strategies. Existing memory- and profile-based methods primarily reuse observable user information, offering limited support for modeling user-state dynamics or selecting actions based on how they shape future user states. We propose PUMA (Prospective User-state Modeling for Action selection), a framework grounded in the Free Energy Principle (FEP) that formulates personalization as decision-making under partial observability, centered on an explicit user state model that captures latent user states and their action-conditioned dynamics. At each turn, PUMA maintains a belief over the user's hidden state, refines the user state model for observation generation and action-conditioned state transition, and selects dialogue actions by minimizing expected free energy, balancing epistemic and pragmatic objectives under a unified criterion. This formulation shifts personalization from passive memory retrieval to model-based decision-making over user evolution. We instantiate PUMA on healthcare-oriented counseling and motivational interviewing benchmarks with latent state annotations for rigorous evaluation. Experiments show that PUMA improves long-horizon dialogue outcomes while maintaining strong response quality, and a cross-dataset study demonstrates more reliable user-state estimation and next-state prediction.

cs.CL↗

Digital N-of-1 Trials and their Application in Experimental Physiology

Traditionally, studies in experimental physiology have been conducted in small groups of human participants, animal models or cell lines. Identifying optimal study designs that achieve sufficient power for drawing proper statistical inferences to detect group level effects with small sample sizes has been challenging. Moreover, average effects derived from traditional group-level inference do not necessarily apply to individual participants. Here, we introduce N-of-1 trials as an innovative study design that can be used to draw valid statistical inference about the effects of interventions on individual participants and can be aggregated across multiple study participants to provide population-level inferences more efficiently than standard group randomized trials. N-of-1 trials have been used in healthcare settings since the late 1980s, but without large-scale adoption and with few applications in experimental physiology research settings. In this manuscript, we introduce the key components and design features of N-of-1 trials, describe statistical analysis and interpretations of the results, and describe some available digital tools to facilitate their use using examples from experimental physiology.

stat.AP↗

Personalization of Large Foundation Models for Health Interventions

Large foundation models (LFMs) transform healthcare AI in prevention, diagnostics, and treatment. However, whether LFMs can provide truly personalized treatment recommendations remains an open question. Recent research has revealed multiple challenges for personalization, including the fundamental generalizability paradox: models achieving high accuracy in one clinical study perform at chance level in others, demonstrating that personalization and external validity exist in tension. This exemplifies broader contradictions in AI-driven healthcare: the privacy-performance paradox, scale-specificity paradox, and the automation-empathy paradox. As another challenge, the degree of causal understanding required for personalized recommendations, as opposed to mere predictive capacities of LFMs, remains an open question. N-of-1 trials -- crossover self-experiments and the gold standard for individual causal inference in personalized medicine -- resolve these tensions by providing within-person causal evidence while preserving privacy through local experimentation. Despite their impressive capabilities, this paper argues that LFMs cannot replace N-of-1 trials. We argue that LFMs and N-of-1 trials are complementary: LFMs excel at rapid hypothesis generation from population patterns using multimodal data, while N-of-1 trials excel at causal validation for a given individual. We propose a hybrid framework that combines the strengths of both to enable personalization and navigate the identified paradoxes: LFMs generate ranked intervention candidates with uncertainty estimates, which trigger subsequent N-of-1 trials. Clarifying the boundary between prediction and causation and explicitly addressing the paradoxical tensions are essential for responsible AI integration in personalized medicine.

cs.AI↗

Co-Exploration and Co-Exploitation via Shared Structure in Multi-Task Bandits

We propose a novel Bayesian framework for efficient exploration in contextual multi-task multi-armed bandit settings, where the context is only observed partially and dependencies between reward distributions are induced by latent context variables. In order to exploit these structural dependencies, our approach integrates observations across all tasks and learns a global joint distribution, while still allowing personalised inference for new tasks. In this regard, we identify two key sources of epistemic uncertainty, namely structural uncertainty in the latent reward dependencies across arms and tasks, and user-specific uncertainty due to incomplete context and limited interaction history. To put our method into practice, we represent the joint distribution over tasks and rewards using a particle-based approximation of a log-density Gaussian process. This representation enables flexible, data-driven discovery of both inter-arm and inter-task dependencies without prior assumptions on the latent variables. Empirically, we demonstrate that our method outperforms baselines such as hierarchical model bandits, especially in settings with model misspecification or complex latent heterogeneity.

cs.LG↗

Combining Unsupervised Learning and Statistical Inference For Multimodal N-of-1 Trials

N-of-1 trials are within-person crossover trials allowing both personalized and population-level inference on the effect of health interventions. Using the full potential of modern technologies, multimodal N-of-1 trials can integrate multimedia data for measuring health outcomes. However, methodology required for automated applications in large multimodal trials is not available yet. Here, we present an unsupervised approach for modeling multimodal N-of-1 trials, bypassing the need for expensive outcome labeling by medical experts. First, an autoencoder is trained on the outcome medical images. Then, the dimensionality of embeddings is reduced by extracting the first principal component, which is finally tested for its association with the treatment. Results from imaging simulation studies show high power in detecting a treatment effect while controlling type I error rates. An application to imaging N-of-1 trials of acne severity identifies individual treatment effects and supports that our methodology can enable large clinical multimodal N-of-1 trials.

stat.AP↗

Personalized Oncology: Feasibility of Evaluating Treatment Effects for Individual Patients

The effectiveness of personalized oncology treatments ultimately depends on whether outcomes can be causally attributed to the treatment. Advances in precision oncology have improved molecular profiling of individuals, and tailored therapies have led to more effective treatments for select patient groups. However, treatment responses still vary among individuals. As cancer is a heterogeneous and dynamic disease with varying treatment outcomes across different molecular types and resistance mechanisms, it requires customized approaches to identify cause-and-effect relationships. N-of-1 trials, or single-subject clinical trials, are designed to evaluate individual treatment effects. Several works have described different causal frameworks to identify treatment effects in N-of-1 trials, yet whether these approaches can be extended to single-cancer patient settings remains unclear. To explore this possibility, a longitudinal dataset from a single metastatic cancer patient with adaptively chosen treatments was considered. The dataset consisted of a detailed treatment plan as well as biomarker and lesion measurements recorded over time. After data processing, a treatment period with sufficient data points to conduct causal inference was selected. Under this setting, a causal framework was applied to define an estimand, identify causal relationships and assumptions, and calculate an individual-specific treatment effect using a time-varying g-formula. Through this application, we illustrate explicitly when and how causal treatment effects can be estimated in single-patient oncology settings. Our findings not only demonstrate the feasibility of applying causal methods in a single-cancer patient setting but also offer a blueprint for using causal methods across a broader spectrum of cancer types in individualized settings.

stat.AP↗

Automated Demand Forecasting in small to medium-sized enterprises

In response to the growing demand for accurate demand forecasts, this research proposes a generalized automated sales forecasting pipeline tailored for small- to medium-sized enterprises (SMEs). Unlike large corporations with dedicated data scientists for sales forecasting, SMEs often lack such resources. To address this, we developed a comprehensive forecasting pipeline that automates time series sales forecasting, encompassing data preparation, model training, and selection based on validation results. The development included two main components: model preselection and the forecasting pipeline. In the first phase, state-of-the-art methods were evaluated on a showcase dataset, leading to the selection of ARIMA, SARIMAX, Holt-Winters Exponential Smoothing, Regression Tree, Dilated Convolutional Neural Networks, and Generalized Additive Models. An ensemble prediction of these models was also included. Long-Short-Term Memory (LSTM) networks were excluded due to suboptimal prediction accuracy, and Facebook Prophet was omitted for compatibility reasons. In the second phase, the proposed forecasting pipeline was tested with SMEs in the food and electric industries, revealing variable model performance across different companies. While one project-based company derived no benefit, others achieved superior forecasts compared to naive estimators. Our findings suggest that no single model is universally superior. Instead, a diverse set of models, when integrated within an automated validation framework, can significantly enhance forecasting accuracy for SMEs. These results emphasize the importance of model diversity and automated validation in addressing the unique needs of each business. This research contributes to the field by providing SMEs access to state-of-the-art sales forecasting tools, enabling data-driven decision-making and improving operational efficiency.

econ.EM↗

Analyzing Population-Level Trials as N-of-1 Trials: an Application to Gait

Studying individual causal effects of health interventions is of interest whenever intervention effects are heterogeneous between study participants. Conducting N-of-1 trials, which are single-person randomized controlled trials, is the gold standard for their analysis. In this study, we propose to re-analyze existing population-level studies as N-of-1 trials as an alternative, and we use gait as a use case for illustration. Gait data were collected from 16 young and healthy participants under fatigued and non-fatigued, as well as under single-task (only walking) and dual-task (walking while performing a cognitive task) conditions. We first computed standard population-level ANOVA models to evaluate differences in gait parameters (stride length and stride time) across conditions. Then, we estimated the effect of the interventions on gait parameters on the individual level through Bayesian linear mixed models, viewing each participant as their own trial, and compared the results. The results illustrated that while few overall population-level effects were visible, individual-level analyses showed nuanced differences between participants. Baseline values of the gait parameters varied largely among all participants, and the changes induced by fatigue and cognitive task performance were also highly heterogeneous, with some individuals showing effects in opposite direction. These differences between population-level and individual-level analyses were more pronounced for the fatigue intervention compared to the cognitive task intervention. Following our empirical analysis, we discuss re-analyzing population studies through the lens of N-of-1 trials more generally and highlight important considerations and requirements. Our work encourages future studies to investigate individual effects using population-level data.

stat.AP↗

Designing and evaluating an online reinforcement learning agent for physical exercise recommendations in N-of-1 trials

Personalized adaptive interventions offer the opportunity to increase patient benefits, however, there are challenges in their planning and implementation. Once implemented, it is an important question whether personalized adaptive interventions are indeed clinically more effective compared to a fixed gold standard intervention. In this paper, we present an innovative N-of-1 trial study design testing whether implementing a personalized intervention by an online reinforcement learning agent is feasible and effective. Throughout, we use a new study on physical exercise recommendations to reduce pain in endometriosis for illustration. We describe the design of a contextual bandit recommendation agent and evaluate the agent in simulation studies. The results show that, first, implementing a personalized intervention by an online reinforcement learning agent is feasible. Second, such adaptive interventions have the potential to improve patients' benefits even if only few observations are available. As one challenge, they add complexity to the design and implementation process. In order to quantify the expected benefit, data from previous interventional studies is required. We expect our approach to be transferable to other interventions and clinical interventions.

cs.LG↗

Anytime-valid inference in N-of-1 trials

App-based N-of-1 trials offer a scalable experimental design for assessing the effects of health interventions at an individual level. Their practical success depends on the strong motivation of participants, which, in turn, translates into high adherence and reduced loss to follow-up. One way to maintain participant engagement is by sharing their interim results. Continuously testing hypotheses during a trial, known as "peeking", can also lead to shorter, lower-risk trials by detecting strong effects early. Nevertheless, traditionally, results are only presented upon the trial's conclusion. In this work, we introduce a potential outcomes framework that permits interim peeking of the results and enables statistically valid inferences to be drawn at any point during N-of-1 trials. Our work builds on the growing literature on valid confidence sequences, which enables anytime-valid inference with uniform type-1 error guarantees over time. We propose several causal estimands for treatment effects applicable in an N-of-1 trial and demonstrate, through empirical evaluation, that the proposed approach results in valid confidence sequences over time. We anticipate that incorporating anytime-valid inference into clinical trials can significantly enhance trial participation and empower participants.

stat.ME↗

Multimodal N-of-1 trials: A Novel Personalized Healthcare Design

N-of-1 trials aim to estimate treatment effects on the individual level and can be applied to personalize a wide range of physical and digital interventions in mHealth. In this study, we propose and apply a framework for multimodal N-of-1 trials in order to allow the inclusion of health outcomes assessed through images, audio or videos. We illustrate the framework in a series of N-of-1 trials that investigate the effect of acne creams on acne severity assessed through pictures. For the analysis, we compare an expert-based manual labelling approach with different deep learning-based pipelines where in a first step, we train and fine-tune convolutional neural networks (CNN) on the images. Then, we use a linear mixed model on the scores obtained in the first step in order to test the effectiveness of the treatment. The results show that the CNN-based test on the images provides a similar conclusion as tests based on manual expert ratings of the images, and identifies a treatment effect in one individual. This illustrates that multimodal N-of-1 trials can provide a powerful way to identify individual treatment effects and can enable large-scale studies of a large variety of health outcomes that can be actively and passively assessed using technological advances in order to personalized health interventions.

stat.AP↗

StudyMe: A New Mobile App for User-Centric N-of-1 Trials

N-of-1 trials are multi-crossover self-experiments that allow individuals to systematically evaluate the effect of interventions on their personal health goals. Although several tools for N-of-1 trials exist, none support non-experts in conducting their own user-centric trials. In this study we present StudyMe, an open-source mobile application that is freely available from https://play.google.com/store/apps/details?id=health.studyu.me and offers users flexibility and guidance in configuring every component of their trials. We also present research that informed the development of StudyMe. Through an initial survey with 272 participants, we learned that individuals are interested in a variety of personal health aspects and have unique ideas on how to improve them. In an iterative, user-centered development process with intermediate user tests we developed StudyMe that also features an educational part to communicate N-of-1 trial concepts. A final empirical evaluation of StudyMe showed that all participants were able to create their own trials successfully using StudyMe and the app achieved a very good usability rating. Our findings suggest that StudyMe provides a significant step towards enabling individuals to apply a systematic science-oriented approach to personalize health-related interventions and behavior modifications in their everyday lives.

cs.HC↗

StudyU: a platform for designing and conducting innovative digital N-of-1 trials

N-of-1 trials are the gold standard study design to evaluate individual treatment effects and derive personalized treatment strategies. Digital tools have the potential to initiate a new era of N-of-1 trials in terms of scale and scope, but fully-functional platforms are not yet available. Here, we present the open source StudyU platform which includes the StudyU designer and StudyU app. With the StudyU designer, scientists are given a collaborative web application to digitally specify, publish, and conduct N-of-1 trials. The StudyU app is a smartphone application with innovative user-centric elements for participants to partake in the published trials and assess the effects of different interventions on their health. Thereby, the StudyU platform allows clinicians and researchers worldwide to easily design and conduct digital N-of-1 trials in a safe manner. We envision that StudyU can change the landscape of personalized treatments both for patients and healthy individuals, democratize and personalize evidence generation for self-optimization and medicine, and can be integrated in clinical practice.

cs.HC↗

Directed Acyclic Graphs and causal thinking in clinical risk prediction modeling

Background: In epidemiology, causal inference and prediction modeling methodologies have been historically distinct. Directed Acyclic Graphs (DAGs) are used to model a priori causal assumptions and inform variable selection strategies for causal questions. Although tools originally designed for prediction are finding applications in causal inference, the counterpart has remained largely unexplored. The aim of this theoretical and simulation-based study is to assess the potential benefit of using DAGs in clinical risk prediction modeling. Methods and Findings: We explore how incorporating knowledge about the underlying causal structure can provide insights about the transportability of diagnostic clinical risk prediction models to different settings. A single-predictor model in the causal direction is likely to have better transportability than one in the anticausal direction. We further probe whether causal knowledge can be used to improve predictor selection. We empirically show that the Markov Blanket, the set of variables including the parents, children, and parents of the children of the outcome node in a DAG, is the optimal set of predictors for that outcome. Conclusions: Our findings challenge the generally accepted notion that a change in the distribution of the predictors does not affect diagnostic clinical risk prediction model calibration if the predictors are properly included in the model. Furthermore, using DAGs to identify Markov Blanket variables may be a useful, efficient strategy to select predictors in clinical risk prediction models if strong knowledge of the underlying causal structure exists or can be learned.

stat.ME↗

Integrating omics and MRI data with kernel-based tests and CNNs to identify rare genetic markers for Alzheimer's disease

For precision medicine and personalized treatment, we need to identify predictive markers of disease. We focus on Alzheimer's disease (AD), where magnetic resonance imaging scans provide information about the disease status. By combining imaging with genome sequencing, we aim at identifying rare genetic markers associated with quantitative traits predicted from convolutional neural networks (CNNs), which traditionally have been derived manually by experts. Kernel-based tests are a powerful tool for associating sets of genetic variants, but how to optimally model rare genetic variants is still an open research question. We propose a generalized set of kernels that incorporate prior information from various annotations and multi-omics data. In the analysis of data from the Alzheimer's Disease Neuroimaging Initiative (ADNI), we evaluate whether (i) CNNs yield precise and reliable brain traits, and (ii) the novel kernel-based tests can help to identify loci associated with AD. The results indicate that CNNs provide a fast, scalable and precise tool to derive quantitative AD traits and that new kernels integrating domain knowledge can yield higher power in association tests of very rare variants.

stat.ML↗