SearcharxivSearch

arXiv subjects

Daniela Zöller

Publications and source records attributed to Daniela Zöller.

4 recordsLinked to original sources

Improving prediction models by incorporating external data with weights based on similarity

In clinical settings, we often face the challenge of building prediction models based on small observational data sets. For example, such a data set might be from a medical center in a multi-center study. Differences between centers might be large, thus requiring specific models based on the data set from the target center. Still, we want to borrow information from the external centers, to deal with small sample sizes. There are approaches that either assign weights to each external data set or each external observation. To incorporate information on differences between data sets and observations, we propose an approach that combines both into weights that can be incorporated into a likelihood for fitting regression models. Specifically, we suggest weights at the data set level that incorporate information on how well the models that provide the observation weights distinguish between data sets. Technically, this takes the form of inverse probability weighting. We explore different scenarios where covariates and outcomes differ among data sets, informing our simulation design for method evaluation. The concept of effective sample size is used for understanding the effectiveness of our subgroup modeling approach. We demonstrate our approach through a clinical application, predicting applied radiotherapy doses for cancer patients. Generally, the proposed approach provides improved prediction performance when external data sets are similar. We thus provide a method for quantifying similarity of external data sets to the target data set and use this similarity to include external observations for improving performance in a target data set prediction modeling task with small data.

stat.ME

Distributed Multivariate Regression Modeling For Selecting Biomarkers Under Data Protection Constraints

The discovery of clinical biomarkers requires large patient cohorts and is aided by a pooled data approach across institutions. In many countries, data protection constraints, especially in the clinical environment, forbid the exchange of individual-level data between different research institutes, impeding the conduct of a joint analyses. To circumvent this problem, only non-disclosive aggregated data is exchanged, which is often done manually and requires explicit permission before transfer, i.e., the number of data calls and the amount of data should be limited. This does not allow for more complex tasks such as variable selection, as only simple aggregated summary statistics are typically transferred. Other methods have been proposed that require more complex aggregated data or use input data perturbation, but these methods can either not deal with a high number of biomarkers or lose information. Here, we propose a multivariable regression approach for identifying biomarkers by automatic variable selection based on aggregated data in iterative calls, which can be implemented under data protection constraints. The approach can be used to jointly analyze data distributed across several locations. To minimize the amount of transferred data and the number of calls, we also provide a heuristic variant of the approach. When performing global data standardization, the proposed method yields the same results as pooled individual-level data analysis. In a simulation study, the information loss introduced by local standardization is seen to be minimal. In a typical scenario, the heuristic decreases the number of data calls from more than 10 to 3, rendering manual data releases feasible. To make our approach widely available for application, we provide an implementation of the heuristic version incorporated in the DataSHIELD framework.\

stat.ML

Agito ergo sum: correlates of spatiotemporal motion characteristics during fMRI

The impact of in-scanner motion on functional magnetic resonance imaging (fMRI) data has a notorious reputation in the neuroimaging community. State-ofthe-art guidelines advise to scrub out excessively corrupted frames as assessed by a composite framewise displacement (FD) score, to regress out models of nuisance variables, and to include average FD as a covariate in group-level analyses. Here, we studied individual motion time courses at time points typically retained in fMRI analyses. We observed that even in this set of putatively clean time points, motion exhibited a very clear spatiotemporal structure, so that we could distinguish subjects into four groups of movers with varying characteristics Then, we showed that this spatiotemporal motion cartography tightly relates to a broad array of anthropometric, behavioral and clinical factors. Convergent results were obtained from two different analytical perspectives: univariate assessment of behavioral differences across mover subgroups unraveled defining markers, while subsequent multivariate analysis broadened the range of involved factors and clarified that multiple motion/behavior modes of covariance overlap in the data. Our results demonstrate that even the smaller episodes of motion typically retained in fMRI analyses carry structured, behaviorally relevant information. They call for further examinations of possible biases in current regression-based motion correction strategies.

q-bio.NC

Modeling Activity Tracker Data Using Deep Boltzmann Machines

Commercial activity trackers are set to become an essential tool in health research, due to increasing availability in the general population. The corresponding vast amounts of mostly unlabeled data pose a challenge to statistical modeling approaches. To investigate the feasibility of deep learning approaches for unsupervised learning with such data, we examine weekly usage patterns of Fitbit activity trackers with deep Boltzmann machines (DBMs). This method is particularly suitable for modeling complex joint distributions via latent variables. We also chose this specific procedure because it is a generative approach, i.e., artificial samples can be generated to explore the learned structure. We describe how the data can be preprocessed to be compatible with binary DBMs. The results reveal two distinct usage patterns in which one group frequently uses trackers on Mondays and Tuesdays, whereas the other uses trackers during the entire week. This exemplary result shows that DBMs are feasible and can be useful for modeling activity tracker data.

stat.ML