SearcharxivSearch

arXiv subjects

Roberta De Vito

Publications and source records attributed to Roberta De Vito.

15 recordsLinked to original sources

Multivariate Causal Effects: a Bayesian Causal Regression Factor Model

The impact of wildfire smoke on air quality is a growing concern, contributing to air pollution through a complex mixture of chemical species with important implications for public health. While previous studies have primarily focused on its association with total particulate matter (PM2.5), the causal relationship between wildfire smoke and the chemical composition of PM2.5 remains largely unexplored. Exposure to these chemical mixtures plays a critical role in shaping public health, yet capturing their relationships requires advanced statistical methods capable of modeling the complex dependencies among chemical species. To fill this gap, we propose a Bayesian causal regression factor model that estimates the multivariate causal effects of wildfire smoke on the concentration of 27 chemical species in PM2.5 across the United States. Our approach introduces two key innovations: (i) a causal inference framework for multivariate potential outcomes, and (ii) a novel Bayesian factor model that employs a probit stick-breaking process as prior for treatment-specific factor scores. By focusing on factor scores, our method addresses the missing data challenge common in causal inference and enables a flexible, data-driven characterization of the latent factor structure, which is crucial to capture the complex correlation among multivariate outcomes. Through Monte Carlo simulations, we show the model's accuracy in estimating the causal effects in multivariate outcomes and characterizing the treatment-specific latent structure. Finally, we apply our method to US air quality data, estimating the causal effect of wildfire smoke on 27 chemical species in PM2.5, providing a deeper understanding of their interdependencies.

stat.ME

Bayesian Nonparametric Causal Inference for High-Dimensional Nutritional Data via Factor-Based Exposure Mapping

Diet plays a crucial role in health, and understanding the causal effects of dietary patterns is essential for informing public health policy and personalized nutrition strategies. However, causal inference in nutritional epidemiology faces several challenges: (i) high-dimensional and correlated food/nutrient intake data induce massive treatment levels; (ii) nutritional studies are interested in latent dietary patterns rather than single food items; and (iii) the goal is to estimate heterogeneous causal effects of these dietary patterns on health outcomes. We address these challenges by introducing a sophisticated exposure mapping framework that reduces the high-dimensional treatment space via factor analysis and enables the identification of dietary patterns. We also extend the Bayesian Causal Forest to accommodate three ordered levels of dietary exposure, better capturing the complex structure of nutritional data and enabling estimation of heterogeneous causal effects. We evaluate the proposed method through extensive simulations and apply it to a multi-center epidemiological study of Hispanic/Latino adults residing in the US. Using high-dimensional dietary data, we identify six dietary patterns and estimate their causal link with two key health risk factors: body mass index and fasting insulin levels. Our findings suggest that higher consumption of plant lipid-antioxidant, plant-based, animal protein, and dairy product patterns is associated with reduced risk.

stat.ME

Estimating Gaussian graphical models of multi-study data with Multi-Study Factor Analysis

Network models are powerful tools for gaining new insights from complex biological data. Most lines of investigation in biology involve comparing datasets in the setting where the same predictors are measured across multiple studies or conditions (multi-study data). Consequently, the development of statistical tools for network modeling of multi-study data is a highly active area of research. Multi-study factor analysis (MSFA) is a method for estimation of latent variables (factors) in multi-study data. In this work, we generalize MSFA by adding the capacity to estimate Gaussian graphical models (GGMs). Our new tool, MSFA-X, is a framework for latent variable-based graphical modeling of shared and study-specific signals in multi-study data. We demonstrate through simulation that MSFA-X can recover shared and study-specific GGMs and outperforms a graphical lasso benchmark. We apply MSFA-X to analyze maternal response to an oral glucose tolerance test in targeted metabolomic profiles from the Hyperglycemia and Adverse Pregnancy Outcomes (HAPO) Study, identifying network-level differences in glucose metabolism between women with and without gestational diabetes mellitus.

stat.ME

Zero-Inflated Bayesian Multi-Study Infinite Non-Negative Matrix Factorization

Understanding the association between dietary patterns and health outcomes, such as the cancer risk, is crucial to inform public health guidelines and shaping future dietary interventions. However, dietary intake data present several statistical challenges: they are high-dimensional, often sparse with excess zeros, and exhibit heterogeneity driven by individual-level covariates. Non-Negative Matrix Factorization (NMF), commonly used to estimate patterns in high-dimensional count data, typically relies on Poisson assumptions and lacks the flexibility to fully address these complexities. Additionally, integrating data across multiple studies, such as case-control studies on cancer risk, requires models that can share information across sources while preserving study-specific structure. In this paper, we introduce a novel Bayesian NMF model that (i) jointly models multi-study count data to enable cross-study information sharing, (ii) incorporate a mixture component to account for zero inflation, and (iii) leverage flexible Bayesian non-parametric priors for characterizing the heterogeneity in pattern scores induced by the individual covariates. This structure allows for clustering of individuals based on dietary profiles, enabling downstream association analyses with health outcomes. Through extensive simulation studies, we demonstrate that our model significantly improves estimation accuracy compared to existing Bayesian NMF methods. We further illustrate its utility through an application to multiple case-control studies on diet and upper aero-digestive tract cancers, identifying nutritionally meaningful dietary patterns. An R package implementing our approach is available at https://github.com/blhansen/ZIMultiStudyNMF.

stat.ME

Missing data imputation using a truncated Gaussian infinite factor model with application to metabolomics data

Metabolomics is the study of small molecules in biological samples. Metabolomics data are typically high-dimensional and contain highly correlated variables and frequent missing values. Both missing at random (MAR) data, due to acquisition or processing errors, and missing not at random (MNAR) data, caused by values falling below detection thresholds, are common. Thus, imputation is a critical component of downstream analysis. Existing imputation methods generally assume one type of data missingness mechanism, or impute values outside the data's physical constraints. A novel truncated Gaussian infinite factor analysis (TGIFA) model is proposed to perform statistically principled and physically realistic imputation in metabolomics data. By incorporating truncated Gaussian assumptions, TGIFA respects the data's physical constraints, while leveraging an infinite latent factor framework to capture high-dimensional dependencies without pre-specifying the number of latent factors. Our Bayesian inference approach enables uncertainty quantification in both the values of the imputed data, and the missing data mechanism. A computationally efficient exchange algorithm enables scalable posterior inference via Markov Chain Monte Carlo. We validate TGIFA through a comprehensive simulation study and demonstrate its utility in a motivating urinary metabolomics dataset, where it yields useful imputations, with associated uncertainty quantification. Open-source R code, available at https://github.com/kfinucane/TGIFA, accompanies TGIFA.

stat.ME

Causal Inference for Latent Outcomes Learned with Factor Models

In many fields$\unicode{x2013}$including genomics, epidemiology, natural language processing, social and behavioral sciences, and economics$\unicode{x2013}$it is increasingly important to address causal questions in the context of factor models or representation learning. In this work, we investigate causal effects on $\textit{latent outcomes}$ derived from high-dimensional observed data using nonnegative matrix factorization. To the best of our knowledge, this is the first study to formally address causal inference in this setting. A central challenge is that estimating a latent factor model can cause an individual's learned latent outcome to depend on other individuals' treatments, thereby violating the standard causal inference assumption of no interference. We formalize this issue as $\textit{learning-induced interference}$ and distinguish it from interference present in a data-generating process. To address this, we propose a novel, intuitive, and theoretically grounded algorithm to estimate causal effects on latent outcomes while mitigating learning-induced interference and improving estimation efficiency. We establish theoretical guarantees for the consistency of our estimator and demonstrate its practical utility through simulation studies and an application to cancer mutational signature analysis. All baseline and proposed methods are available in our open-source R package, ${\tt causalLFO}$.

stat.ME

Bayesian integrative factor analysis methods, with application in nutrition and genomics data

High-dimensional data are crucial in biomedical research. Integrating such data from multiple studies is a critical process that relies on the choice of advanced statistical models, enhancing statistical power, reproducibility, and scientific insight compared to analyzing each study separately. Factor analysis (FA) is a core dimensionality reduction technique that models observed data through a small set of latent factors. Bayesian extensions of FA have recently emerged as powerful tools for multi-study integration, enabling researchers to disentangle shared biological signals from study-specific variability. In this tutorial, we provide a practical and comparative guide to five advanced Bayesian integrative factor models: Perturbed Factor Analysis (PFA), Bayesian Factor Regression with non-local spike-and-slab priors (MOM-SS), Subspace Factor Analysis (SUFA), Bayesian Multi-study Factor Analysis (BMSFA), and Bayesian Combinatorial Multi-study Factor Analysis (Tetris). To contextualize these methods, we also include two benchmark approaches: standard FA applied to pooled data (Stack FA) and FA applied separately to each study (Ind FA). We evaluate all methods through extensive simulations, assessing computational efficiency and accuracy in the estimation of loadings and number of factors. To bridge theory and practice, we present a full analytical workflow, with detailed R code, demonstrating how to apply these models to real-world datasets in nutrition and genomics. This tutorial is designed to guide applied researchers through the landscape of Bayesian integrative factor analysis, offering insights and tools for extracting interpretable, robust patterns from complex multi-source data. All code and resources are available at: https://github.com/Mavis-Liang/Bayesian_integrative_FA_tutorial

stat.AP

Sparse Bayesian Factor Models with Mass-Nonlocal Factor Scores

Bayesian factor models are widely used for dimensionality reduction and pattern discovery in high-dimensional datasets across diverse fields. These models typically focus on imposing priors on factor loading to induce sparsity and improve interpretability. However, factor scores, which play a critical role in individual-level associations with factors, have received less attention and are assumed to follow a standard normal distribution. This assumption oversimplifies the heterogeneity often observed in real-world applications. We propose the sparse Bayesian Factor model with MAss-Nonlocal factor scores (BFMAN), a novel framework that addresses these limitations by introducing a mass-nonlocal prior on factor scores. This prior allows for both exact zeros and flexible, nonlocal behavior, capturing individual-level sparsity and heterogeneity. The sparsity in the score matrix enables a robust and novel approach to determine the optimal number of factors. Model parameters are estimated via a fast and efficient Gibbs sampler. Extensive simulations demonstrate that BFMAN outperforms standard Bayesian factor models in factor recovery, sparsity detection, score estimation, and selection of the optimal number of factors. We apply BFMAN to the Hispanic Community Health Study/Study of Latinos, identifying meaningful dietary patterns and their associations with cardiovascular disease, showcasing the model's ability to uncover insights into complex nutritional data.

stat.ME

Multi-study factor regression model: an application in nutritional epidemiology

Diet is a risk factor for many diseases. In nutritional epidemiology, studying reproducible dietary patterns is critical to reveal important associations with health. However, it is challenging: diverse cultural and ethnic backgrounds may critically impact eating patterns, showing heterogeneity, leading to incorrect dietary patterns and obscuring the components shared across different groups or populations. Moreover, covariate effects generated from observed variables, such as demographics and other confounders, can further bias these dietary patterns. Identifying the shared and group-specific dietary components and covariate effects is essential to drive accurate conclusions. To address these issues, we introduce a new modeling factor regression, the Multi-Study Factor Regression (MSFR) model. The MSFR model analyzes different populations simultaneously, achieving three goals: capturing shared component(s) across populations, identifying group-specific structures, and correcting for covariate effects. We use this novel method to derive common and ethnic-specific dietary patterns in a multi-center epidemiological study in Hispanic/Latinos community. Our model improves the accuracy of common and group dietary signals and yields better prediction than other techniques, revealing significant associations with health. In summary, we provide a tool to integrate different groups, giving accurate dietary signals crucial to inform public health policy.

stat.AP

Bayesian Probit Multi-Study Non-negative Matrix Factorization for Mutational Signatures

Mutational signatures are patterns of somatic mutations in tumor genomes that provide insights into underlying mutagenic processes and cancer origin. Developing reliable methods for their estimation is of growing importance in cancer biology. Somatic mutation data are often collected for different cancer types, highlighting the need for multi-study approaches that enable joint analysis in a principled and integrative manner. Despite significant advancements, statistical models tailored for analyzing the genomes of multiple cancer types remain underexplored. In this work, we introduce a Bayesian Multi-Study Non-negative Matrix Factorization (NMF) approach that uses mixture modeling to incorporate sparsity in the exposure weights of each subject to mutational signatures, allowing for individual tumor profiles to be represented by a subset rather than all signatures, and making this subset depend on covariates. This allows for a) more precise ability to identify meaningful contributions of mutational signatures at the individual level; b) estimation of the prevalence of activity of signatures within a cancer type, defined by the proportion of tumor profiles where a certain signature is present; and c) de-novo identification of interpretable patient subtypes based on the mutational signatures present within their mutational profile. We apply our approach to the mutational profiles of tumors from seven different cancer types, demonstrating its ability to accurately estimate mutational signatures while uncovering both individual and tissue-specific differences. An R package implementing our method is available at https://github.com/blhansen/BAPmultiNMF.

stat.AP

Fast Variational Inference for Bayesian Factor Analysis in Single and Multi-Study Settings

Factors models are routinely used to analyze high-dimensional data in both single-study and multi-study settings. Bayesian inference for such models relies on Markov Chain Monte Carlo (MCMC) methods which scale poorly as the number of studies, observations, or measured variables increase. To address this issue, we propose variational inference algorithms to approximate the posterior distribution of Bayesian latent factor models using the multiplicative gamma process shrinkage prior. The proposed algorithms provide fast approximate inference at a fraction of the time and memory of MCMC-based implementations while maintaining comparable accuracy in characterizing the data covariance matrix. We conduct extensive simulations to evaluate our proposed algorithms and show their utility in estimating the model for high-dimensional multi-study gene expression data in ovarian cancers. Overall, our proposed approaches enable more efficient and scalable inference for factor models, facilitating their use in high-dimensional settings. An R package VIMSFA implementing our methods is available on GitHub (github.com/blhansen/VI-MSFA).

stat.ME

Bayesian Combinatorial Multi-Study Factor Analysis

Analyzing multiple studies allows leveraging data from a range of sources and populations, but until recently, there have been limited methodologies to approach the joint unsupervised analysis of multiple high-dimensional studies. A recent method, Bayesian Multi-Study Factor Analysis (BMSFA), identifies latent factors common to all studies, as well as latent factors specific to individual studies. However, BMSFA does not allow for partially shared factors, i.e. latent factors shared by more than one but less than all studies. We extend BMSFA by introducing a new method, Tetris, for Bayesian combinatorial multi-study factor analysis, which identifies latent factors that can be shared by any combination of studies. We model the subsets of studies that share latent factors with an Indian Buffet Process. We test our method with an extensive range of simulations, and showcase its utility not only in dimension reduction but also in covariance estimation. Finally, we apply Tetris to high-dimensional gene expression datasets to identify patterns in breast cancer gene expression, both within and across known classes defined by germline mutations.

stat.ME

Bayesian Ordinal Quantile Regression with a Partially Collapsed Gibbs Sampler

Unlike standard linear regression, quantile regression captures the relationship between covariates and the conditional response distribution as a whole, rather than only the relationship between covariates and the expected value of the conditional response. However, while there are well-established quantile regression methods for continuous variables and some forms of discrete data, there is no widely accepted method for ordinal variables, despite their importance in many medical contexts. In this work, we describe two existing ordinal quantile regression methods and demonstrate their weaknesses. We then propose a new method, Bayesian ordinal quantile regression with a partially collapsed Gibbs sampler (BORPS). We show superior results using BORPS versus existing methods on an extensive set of simulations. We further illustrate the benefits of our method by applying BORPS to the Fragile Families and Child Wellbeing Study data to tease apart associations with early puberty among both genders. Software is available at: GitHub.com/igrabski/borps.

stat.ME

Multi-study Factor Analysis

We introduce a novel class of factor analysis methodologies for the joint analysis of multiple studies. The goal is to separately identify and estimate 1) common factors shared across multiple studies, and 2) study-specific factors. We develop a fast Expectation Conditional-Maximization algorithm for parameter estimates and we provide a procedure for choosing the common and specific factor. We present simulations evaluating the performance of the method and we illustrate it by applying it to gene expression data in ovarian cancer. In both cases, we clarify the benefits of a joint analysis compared to the standard factor analysis. We hope to have provided a valuable tool to accelerate the pace at which we can combine unsupervised analysis across multiple studies, and understand the cross-study reproducibility of signal in multivariate data.

stat.AP

Bayesian Multi-study Factor Analysis for High-throughput Biological Data

This paper presents a new modeling strategy for joint unsupervised analysis of multiple high-throughput biological studies. As in Multi-study Factor Analysis, our goals are to identify both common factors shared across studies and study-specific factors. Our approach is motivated by the growing body of high-throughput studies in biomedical research, as exemplified by the comprehensive set of expression data on breast tumors considered in our case study. To handle high-dimensional studies, we extend Multi-study Factor Analysis using a Bayesian approach that imposes sparsity. Specifically, we generalize the sparse Bayesian infinite factor model to multiple studies. We also devise novel solutions for the identification of the loading matrices: we recover the loading matrices of interest ex-post, by adapting the orthogonal Procrustes approach. Computationally, we propose an efficient and fast Gibbs sampling approach. Through an extensive simulation analysis, we show that the proposed approach performs very well in a range of different scenarios, and outperforms standard Factor analysis in all the scenarios identifying replicable signal in unsupervised genomic applications. The results of our analysis of breast cancer gene expression across seven studies identified replicable gene patterns, clearly related to well-known breast cancer pathways. An R package is implemented and available on GitHub.

stat.AP