Searcharxiv⌕ Search

arXiv subjects

Mark van de Wiel

Publications and source records attributed to Mark van de Wiel.

4 recordsLinked to original sources

Bayesian Profile Regression with Linear Mixed Models (Profile-LMM) applied to Longitudinal Exposome Data

Exposure to diverse non-genetic factors, known as the exposome, is a critical determinant of health outcomes. However, analyzing the exposome presents significant methodological challenges, including: high collinearity among exposures, the longitudinal nature of repeated measurements, and potential complex interactions with individual characteristics. In this paper, we address these challenges by proposing a novel statistical framework that extends Bayesian profile regression. Our method integrates profile regression, which handles collinearity by clustering exposures into latent profiles, into a linear mixed model (LMM), a framework for longitudinal data analysis. This profile-LMM approach effectively accounts for within-person variability over time while also incorporating interactions between the latent exposure clusters and individual characteristics. We validate our method using simulated data, demonstrating its ability to accurately identify model parameters and recover the true latent exposure cluster structure. Finally, we apply this approach to a large longitudinal data set from the Lifelines cohort to identify combinations of exposures that are significantly associated with diastolic blood pressure.

stat.ME↗

Bayesian regression discontinuity design with unknown cutoff

The regression discontinuity design (RDD) is a quasi-experimental approach used to estimate the causal effects of an intervention assigned based on a cutoff criterion. RDD exploits the idea that close to the cutoff units below and above are similar; hence, they can be meaningfully compared. Consequently, the causal effect can be estimated only locally at the cutoff point. This makes the cutoff point an essential element of RDD. However, the exact cutoff location may not always be disclosed to the researchers, and even when it is, the actual location may deviate from the official one. As we illustrate on the application of RDD to the HIV treatment eligibility data, estimating the causal effect at an incorrect cutoff point leads to meaningless results. The method we present, LoTTA (Local Trimmed Taylor Approximation), can be applied both as an estimation and validation tool in RDD. We use a Bayesian approach to incorporate prior knowledge and uncertainty about the cutoff location in the causal effect estimation. At the same time, LoTTA is fitted globally to the whole data, whereas RDD is a local, boundary point estimation problem. In this work we address a natural question that arises: how to make Bayesian inference more local to render a meaningful and powerful estimate of the treatment effect?

stat.ME↗

A framework for interpretation and testing of sparse canonical correlations

In clinical and biomedical research, multiple high-dimensional datasets are nowadays routinely collected from omics and imaging devices. Multivariate methods, such as Canonical Correlation Analysis (CCA), integrate two (or more) datasets to discover and understand underlying biological mechanisms. For an explorative method like CCA, interpretation is key. We present a sparse CCA method based on soft-thresholding that produces near-orthogonal components, allows for browsing over various sparsity levels, and permutation-based hypothesis testing. Our soft-thresholding approach avoids tuning of a penalty parameter. Such tuning is computationally burdensome and may render unintelligible results. In addition, unlike alternative approaches, our method is less dependent on the initialisation. We examined the performance of our approach with simulations and illustrated its use on real cancer genomics data from drug sensitivity screens. Moreover, we compared its performance to Penalised Matrix Analysis (PMA), which is a popular alternative of sparse CCA with a focus on yielding interpretable results. Compared to PMA, our method offers improved interpretability of the results, while not compromising, or even improving, signal discovery. he software and simulation framework are available at https://github.com/nuria-sv/toscca.

stat.ME↗

Informative co-data learning for high-dimensional Horseshoe regression

High-dimensional data often arise from clinical genomics research to infer relevant predictors of a particular trait. A way to improve the predictive performance is to include information on the predictors derived from prior knowledge or previous studies. Such information is also referred to as ``co-data''. To this aim, we develop a novel Bayesian model for including co-data in a high-dimensional regression framework, called Informative Horseshoe regression (infHS). The proposed approach regresses the prior variances of the regression parameters on the co-data variables, improving variable selection and prediction. We implement both a Gibbs sampler and a Variational approximation algorithm. The former is suited for applications of moderate dimensions which, besides prediction, target posterior inference, whereas the computational efficiency of the latter allows handling a very large number of variables. We show the benefits from including co-data with a simulation study. Eventually, we demonstrate that infHS outperforms competing approaches for two genomics applications.

stat.ME↗