Searcharxiv⌕ Search

arXiv subjects

Amy H. Herring

Publications and source records attributed to Amy H. Herring.

At least 19 recordsLinked to original sources

Testing Additivity in Lead and Benzo[a]pyrene-induced Neurodegeneration in Caenorhabditis elegans

Exposure to environmental contaminants is a recognized cause of neurotoxicity, contributing to the onset of a broad range of neurological conditions. In realistic settings, such exposure involves com- plex mixtures, and the combined effect of their components may differ from what their individual effects would predict. Characterizing such interactions and testing them against a principled notion of additivity is central to assessing the neurotoxicological risk. We take up these questions for two widespread and independently neurotoxic pollutants, lead (Pb) and benzo[a]pyrene (BaP), through a novel C. elegans assay in which nematodes were subjected to single and joint exposures across a range of doses. Morphological damage is quantified on an ordinal scale at the level of individual dopaminergic neurons. To analyze these data, we model the full distribution of the ordinal response as a convex mixture between an unexposed and a maximally affected profile. The weight of this mixture varies with chemical doses, modeled flexibly via monotone splines and, for the joint effect, in a radial coordinate system. Additivity is assessed via a likelihood ratio test against established null models, and is calibrated via parametric bootstrap. Applied to the C. elegans assay, our analysis reveals a localized, asymmetric synergy between Pb and BaP, concentrated where moderate BaP meets high Pb exposure.

stat.AP↗

Order-Restricted Bayesian Ordinal Regression for the Modeling of Neuron Degeneration in Caenorhabditis elegans

Neuron degeneration is the underlying mechanism for the development of many diseases. Quantifying the association between increasing levels of toxic exposure and progressive neuronal damage is a critical component of understanding this development. We investigate this association by analyzing a novel dataset of ordinal neuronal damage scores derived from a series of toxicological assays of C. elegans, including variables such as toxicant concentration, maternal treatment, and direct chemical exposure. We propose a computationally efficient parameter-constrained Bayesian ordinal regression that captures the monotonic association between neuron damage scores and corresponding treatments. Power analysis via simulation studies reinforces the advantages of our model over standard alternatives used in existing work by practitioners. Analysis of the novel C. elegans assays indicates that maternal toxicity increases susceptibility in progeny, with the offspring generation exhibiting amplified neuronal damage upon later-life rotenone exposure even under mild parental developmental treatment.

stat.AP↗

Bayesian modeling of nearly mutually orthogonal processes

Functional factor analysis is an important dimension reduction method for functional and longitudinal data. Factor loadings give insight into patterns of variability of the observations, while latent factors provide a low-dimensional representation of the data that is useful for inferential tasks. Constraining the functional factor loadings to be mutually orthogonal is desirable for model parsimony but is computationally challenging. In this work, we introduce nearly mutually orthogonal processes, which can be used to effectively enforce mutual orthogonality of factor loadings while maintaining computational simplicity and efficiency. The joint distribution is governed by a penalty parameter that determines the degree to which the processes are mutually orthogonal and is related to ease of posterior computation. We demonstrate that our approach can be used for flexible and interpretable inference in an application to studying the effects of breastfeeding status, illness, and demographic factors on weight dynamics in early childhood. Code is available on GitHub: https://github.com/jamesmatuk/NeMO-FFA

stat.ME↗

Pathway-based Bayesian factor models for 'omics data

Interpreting RNA-sequencing data requires identifying coordinated gene expression patterns that correspond to biological pathways. Standard factor models provide useful dimension reduction but typically ignore existing pathway knowledge or incorporate it through restrictive assumptions, limiting interpretability, and reproducibility. Here, we develop Bayesian Analysis with gene-Sets Informed Latent space (BASIL), a scalable framework for analyzing transcriptomic data that integrates annotated gene sets into latent variable inference. BASIL places structured priors on factor loadings, shrinking them toward combinations of annotated gene sets, enhancing biological interpretability and stability, while simultaneously learning new unstructured components. BASIL provides accurate covariance estimates and uncertainty quantification, without resorting to computationally expensive Markov chain Monte Carlo sampling, by exploiting a pre-training approach that pre-estimates the latent factors. An automatic empirical Bayes procedure eliminates the need for manual hyperparameter tuning, promoting reproducibility and usability in practice. Applying BASIL to the global fever transcriptomic cohort uncovers interpretable host-response modules, with phosphoinositide signaling and interferon-driven inflammation emerging as key drivers of gene-expression variability.

stat.ME↗

Prenatal phthalate exposures and adiposity outcomes trajectories: a multivariate Bayesian factor regression approach

Experimental animal evidence and a growing body of observational studies suggest that prenatal exposure to phthalates may be a risk factor for childhood obesity. Using data from the Mount Sinai Children's Environmental Health Study (MSCEHS), which measured urinary phthalate metabolites (including MEP, MnBP, MiBP, MCPP, MBzP, MEHP, MEHHP, MEOHP, and MECPP) during the third trimester of pregnancy (between 25 and 40 weeks) of 382 mothers, we examined adiposity outcomes: body mass index (BMI), fat mass percentage, waist-to-hip ratio, and waist circumference, of 180 children between ages 4 and 9. We aimed to assess the effects of prenatal exposure to phthalates on these adiposity outcomes, with potential time-varying and sex-specific effects. We applied a novel Bayesian multivariate factor regression (BMFR) that (1) represents phthalate mixtures as latent factors, a DEHP and a non-DEHP factor, (2) borrows information across highly correlated adiposity outcomes to improve estimation precision, (3) models potentially non-linear time-varying effects of the latent factors on adiposity outcomes, and (4) fully quantifies uncertainty using state-of-the-art prior specifications. The results show that in boys, at younger ages (4-6), all phthalate components are associated with lower adiposity outcomes; however, after age 7, they are associated with higher outcomes. In girls, there is no evidence of associations between phthalate factors and adiposity outcomes.

stat.ME↗

Identifiable and interpretable nonparametric factor analysis

Factor models are widely used to reduce dimensionality in modeling high-dimensional data. However, there remains a need for models that can be reliably fit in modest sample sizes and are identifiable, interpretable, and flexible. To address this gap, we propose a NIFTY model that uses a linear factor structure with Gaussian residuals, but with a novel latent variable modeling structure. In particular, we model each latent variable as a one-dimensional nonlinear mapping of a uniform latent location. A key innovation is allowing different latent variables to be transformations of the same latent locations, accommodating intrinsic lower-dimensional nonlinear structures. Leveraging on pre-trained data obtained by diffusion maps and post-processing of MCMC samples, we obtain model identifiability. In addition, we softly constrain the empirical distribution of the latent locations to be close to uniform to address a latent posterior shift problem, which is common in factor models and can lead to substantial bias in parameter inferences, predictions, and generative modeling. We show good performance in density estimation and data visualization in simulations, and apply NIFTY to bird song data in an environmental monitoring application.

stat.ME↗

Bayesian Learning of Clinically Meaningful Sepsis Phenotypes in Northern Tanzania

Sepsis is a life-threatening condition caused by a dysregulated host response to infection. Recently, researchers have hypothesized that sepsis consists of a heterogeneous spectrum of distinct subtypes, motivating several studies to identify clusters of sepsis patients that correspond to subtypes, with the long-term goal of using these clusters to design subtype-specific treatments. Therefore, clinicians rely on clusters having a concrete medical interpretation, usually corresponding to clinically meaningful regions of the sample space that have a concrete implication to practitioners. In this article, we propose Clustering Around Meaningful Regions (CLAMR), a Bayesian clustering approach that explicitly models the medical interpretation of each cluster center. CLAMR favors clusterings that can be summarized via meaningful feature values, leading to medically significant sepsis patient clusters. We also provide details on measuring the effect of each feature on the clustering using Bayesian hypothesis tests, so one can assess what features are relevant for cluster interpretation. Our focus is on clustering sepsis patients from Moshi, Tanzania, where patients are younger and the prevalence of HIV infection is higher than in previous sepsis subtyping cohorts.

stat.AP↗

Low-rank longitudinal factor regression with application to chemical mixtures

Developmental epidemiology commonly focuses on assessing the association between multiple early life exposures and childhood health. Statistical analyses of data from such studies focus on inferring the contributions of individual exposures, while also characterizing time-varying and interacting effects. Such inferences are made more challenging by correlations among exposures, nonlinearity, and the curse of dimensionality. Motivated by studying the effects of prenatal bisphenol A (BPA) and phthalate exposures on glucose metabolism in adolescence using data from the ELEMENT study, we propose a low-rank longitudinal factor regression (LowFR) model for tractable inference on flexible longitudinal exposure effects. LowFR handles highly-correlated exposures using a Bayesian dynamic factor model, which is fit jointly with a health outcome via a novel factor regression approach. The model collapses on simpler and intuitive submodels when appropriate, while expanding to allow considerable flexibility in time-varying and interaction effects when supported by the data. After demonstrating LowFR's effectiveness in simulations, we use it to analyze the ELEMENT data and find that diethyl and dibutyl phthalate metabolite levels in trimesters 1 and 2 are associated with altered glucose metabolism in adolescence.

stat.AP↗

mpower: An R Package for Power Analysis of Exposure Mixture Studies via Monte Carlo Simulations

Estimating sample size and statistical power is an essential part of a good study design. This R package allows users to conduct power analysis based on Monte Carlo simulations in settings in which consideration of the correlations between predictors is important. It runs power analyses given a data generative model and an inference model. It can set up a data generative model that preserves dependence structures among variables given existing data (continuous, binary, or ordinal) or high-level descriptions of the associations. Users can generate power curves to assess the trade-offs between sample size, effect size, and power of a design. This paper presents tutorials and examples focusing on applications for environmental mixture studies when predictors tend to be moderately to highly correlated. It easily interfaces with several existing and newly developed analysis strategies for assessing associations between exposures and health outcomes. However, the package is sufficiently general to facilitate power simulations in a wide variety of settings.

stat.ME↗

Spatial predictions on physically constrained domains: Applications to Arctic sea salinity data

In this paper we predict sea surface salinity (SSS) in the Arctic Ocean based on satellite measurements. SSS is a crucial indicator for ongoing changes in the Arctic Ocean and can offer important insights about climate change. We particularly focus on areas of water mistakenly flagged as ice by satellite algorithms. To remove bias in the retrieval of salinity near sea ice, the algorithms use conservative ice masks, which result in considerable loss of data. We aim to produce realistic SSS values for such regions to obtain more complete understanding about the SSS surface over the Arctic Ocean and benefit future applications that may require SSS measurements near edges of sea ice or coasts. We propose a class of scalable nonstationary processes that can handle large data from satellite products and complex geometries of the Arctic Ocean. Barrier overlap-removal acyclic directed graph GP (BORA-GP) constructs sparse directed acyclic graphs (DAGs) with neighbors conforming to barriers and boundaries, enabling characterization of dependence in constrained domains. The BORA-GP models produce more sensible SSS values in regions without satellite measurements and show improved performance in various constrained domains in simulation studies compared to state-of-the-art alternatives. An R package is available at https://github.com/jinbora0720/boraGP.

stat.AP↗

Bayesian Matrix Completion for Hypothesis Testing

We aim to infer bioactivity of each chemical by assay endpoint combination, addressing sparsity of toxicology data. We propose a Bayesian hierarchical framework which borrows information across different chemicals and assay endpoints, facilitates out-of-sample prediction of activity for chemicals not yet assayed, quantifies uncertainty of predicted activity, and adjusts for multiplicity in hypothesis testing. Furthermore, this paper makes a novel attempt in toxicology to simultaneously model heteroscedastic errors and a nonparametric mean function, leading to a broader definition of activity whose need has been suggested by toxicologists. Real application identifies chemicals most likely active for neurodevelopmental disorders and obesity.

stat.AP↗

Perturbed factor analysis: Accounting for group differences in exposure profiles

In this article, we investigate group differences in phthalate exposure profiles using NHANES data. Phthalates are a family of industrial chemicals used in plastics and as solvents. There is increasing evidence of adverse health effects of exposure to phthalates on reproduction and neuro-development, and concern about racial disparities in exposure. We would like to identify a single set of low-dimensional factors summarizing exposure to different chemicals, while allowing differences across groups. Improving on current multi-group additive factor models, we propose a class of Perturbed Factor Analysis (PFA) models that assume a common factor structure after perturbing the data via multiplication by a group-specific matrix. Bayesian inference algorithms are defined using a matrix normal hierarchical model for the perturbation matrices. The resulting model is just as flexible as current approaches in allowing arbitrarily large differences across groups but has substantial advantages that we illustrate in simulation studies. Applying PFA to NHANES data, we learn common factors summarizing exposures to phthalates, while showing clear differences across groups.

stat.ME↗

Bayesian Hierarchical Factor Regression Models to Infer Cause of Death From Verbal Autopsy Data

In low-resource settings where vital registration of death is not routine it is often of critical interest to determine and study the cause of death (COD) for individuals and the cause-specific mortality fraction (CSMF) for populations. Post-mortem autopsies, considered the gold standard for COD assignment, are often difficult or impossible to implement due to deaths occurring outside the hospital, expense, and/or cultural norms. For this reason, Verbal Autopsies (VAs) are commonly conducted, consisting of a questionnaire administered to next of kin recording demographic information, known medical conditions, symptoms, and other factors for the decedent. This article proposes a novel class of hierarchical factor regression models that avoid restrictive assumptions of standard methods, allow both the mean and covariance to vary with COD category, and can include covariate information on the decedent, region, or events surrounding death. Taking a Bayesian approach to inference, this work develops an MCMC algorithm and validates the FActor Regression for Verbal Autopsy (FARVA) model in simulation experiments. An application of FARVA to real VA data shows improved goodness-of-fit and better predictive performance in inferring COD and CSMF over competing methods. Code and a user manual are made available at https://github.com/kelrenmor/farva.

stat.AP↗

Bayesian joint modeling of chemical structure and dose response curves

Today there are approximately 85,000 chemicals regulated under the Toxic Substances Control Act, with around 2,000 new chemicals introduced each year. It is impossible to screen all of these chemicals for potential toxic effects either via full organism in vivo studies or in vitro high-throughput screening (HTS) programs. Toxicologists face the challenge of choosing which chemicals to screen, and predicting the toxicity of as-yet-unscreened chemicals. Our goal is to describe how variation in chemical structure relates to variation in toxicological response to enable in silico toxicity characterization designed to meet both of these challenges. With our Bayesian partially Supervised Sparse and Smooth Factor Analysis ($\text{BS}^3\text{FA}$) model, we learn a distance between chemicals targeted to toxicity, rather than one based on molecular structure alone. Our model also enables the prediction of chemical dose-response profiles based on chemical structure (that is, without in vivo or in vitro testing) by taking advantage of a large database of chemicals that have already been tested for toxicity in HTS programs. We show superior simulation performance in distance learning and modest to large gains in predictive ability compared to existing methods. Results from the high-throughput screening data application elucidate the relationship between chemical structure and a toxicity-relevant high-throughput assay. An R package for $\text{BS}^3\text{FA}$ is available online at https://github.com/kelrenmor/bs3fa.

stat.AP↗

A generalized Bayes framework for probabilistic clustering

Loss-based clustering methods, such as k-means and its variants, are standard tools for finding groups in data. However, the lack of quantification of uncertainty in the estimated clusters is a disadvantage. Model-based clustering based on mixture models provides an alternative, but such methods face computational problems and large sensitivity to the choice of kernel. This article proposes a generalized Bayes framework that bridges between these two paradigms through the use of Gibbs posteriors. In conducting Bayesian updating, the log likelihood is replaced by a loss function for clustering, leading to a rich family of clustering methods. The Gibbs posterior represents a coherent updating of Bayesian beliefs without needing to specify a likelihood for the data, and can be used for characterizing uncertainty in clustering. We consider losses based on Bregman divergence and pairwise similarities, and develop efficient deterministic algorithms for point estimation along with sampling algorithms for uncertainty quantification. Several existing clustering algorithms, including k-means, can be interpreted as generalized Bayes estimators under our framework, and hence we provide a method of uncertainty quantification for these approaches.

stat.ME↗

Centered Partition Process: Informative Priors for Clustering

There is a very rich literature proposing Bayesian approaches for clustering starting with a prior probability distribution on partitions. Most approaches assume exchangeability, leading to simple representations in terms of Exchangeable Partition Probability Functions (EPPF). Gibbs-type priors encompass a broad class of such cases, including Dirichlet and Pitman-Yor processes. Even though there have been some proposals to relax the exchangeability assumption, allowing covariate-dependence and partial exchangeability, limited consideration has been given on how to include concrete prior knowledge on the partition. For example, we are motivated by an epidemiological application, in which we wish to cluster birth defects into groups and we have prior knowledge of an initial clustering provided by experts. As a general approach for including such prior knowledge, we propose a Centered Partition (CP) process that modifies the EPPF to favor partitions close to an initial one. Some properties of the CP prior are described, a general algorithm for posterior computation is developed, and we illustrate the methodology through simulation examples and an application to the motivating epidemiology study of birth defects.

stat.ME↗

Nonparametric Bayes models for mixed-scale longitudinal surveys

Modeling and computation for multivariate longitudinal surveys have proven challenging, particularly when data are not all continuous and Gaussian but contain discrete measurements. In many social science surveys, study participants are selected via complex survey designs such as stratified random sampling, leading to discrepancies between the sample and population, which are further compounded by missing data and loss to follow up. Survey weights are typically constructed to address these issues, but it is not clear how to include them in models. Motivated by data on sexual development, we propose a novel nonparametric approach for mixed-scale longitudinal data in surveys. In the proposed approach, the mixed-scale multivariate response is expressed through an underlying continuous variable with dynamic latent factors inducing time-varying associations. Bias from the survey design is adjusted for in posterior computation relying on a Markov chain Monte Carlo algorithm. The approach is assessed in simulation studies, and applied to the National Longitudinal Study of Adolescent to Adult Health.

stat.AP↗

Bayesian Local Extrema Splines

We consider the problem of shape restricted nonparametric regression on a closed set X ?\in R; where it is reasonable to assume the function has no more than H local extrema interior to X: Following a Bayesian approach we develop a nonparametric prior over a novel class of local extrema splines. This approach is shown to be consistent when modeling any continuously differentiable function within the class of functions considered, and is used to develop methods for hypothesis testing on the shape of the curve. Sampling algorithms are developed, and the method is applied in simulation studies and data examples where the shape of the curve is of interest.

stat.ME↗