SearcharxivSearch

arXiv subjects

Federica Stolf

Publications and source records attributed to Federica Stolf.

5 recordsLinked to original sources

Bayesian inference on beta diversity via feature allocation models with imperfect detection

Beta diversity quantifies variation in species composition across ecological communities and is fundamental for understanding biodiversity patterns across space and environmental gradients. Statistical inference on beta diversity is challenging: species occurrence data are high dimensional, many species remain unobserved despite extensive sampling, and surveys are subject to imperfect detection. Existing approaches are typically based on empirical dissimilarity indices with limited uncertainty quantification or on models that rely on unrealistic exchangeability and perfect detection assumptions. We introduce a new class of Bayesian feature allocation models for partially exchangeable species occurrence data with imperfect detection. The framework combines latent feature allocation models with occupancy-based detection mechanisms, allowing heterogeneous species compositions across sites, explicitly accounting for false negatives, and accommodating the discovery of previously unobserved species. We develop coherent probabilistic inference for species sharing and beta diversity, deriving explicit posterior and predictive distributions for between-community heterogeneity, including the number of shared species across sites and the number expected under future sampling. These analytical results yield interpretable posterior summaries of compositional heterogeneity and facilitate scalable inference in high-dimensional biodiversity studies. Simulation studies and an application to global fungal biodiversity data demonstrate improved inference on species sharing and between-community diversity.

stat.ME

Pathway-based Bayesian factor models for 'omics data

Interpreting RNA-sequencing data requires identifying coordinated gene expression patterns that correspond to biological pathways. Standard factor models provide useful dimension reduction but typically ignore existing pathway knowledge or incorporate it through restrictive assumptions, limiting interpretability, and reproducibility. Here, we develop Bayesian Analysis with gene-Sets Informed Latent space (BASIL), a scalable framework for analyzing transcriptomic data that integrates annotated gene sets into latent variable inference. BASIL places structured priors on factor loadings, shrinking them toward combinations of annotated gene sets, enhancing biological interpretability and stability, while simultaneously learning new unstructured components. BASIL provides accurate covariance estimates and uncertainty quantification, without resorting to computationally expensive Markov chain Monte Carlo sampling, by exploiting a pre-training approach that pre-estimates the latent factors. An automatic empirical Bayes procedure eliminates the need for manual hyperparameter tuning, promoting reproducibility and usability in practice. Applying BASIL to the global fever transcriptomic cohort uncovers interpretable host-response modules, with phosphoinositide signaling and interferon-driven inflammation emerging as key drivers of gene-expression variability.

stat.ME

Bayesian Adaptive Tucker Decompositions for Tensor Factorization

Tucker tensor decomposition offers a more effective representation for multiway data compared to the widely used PARAFAC model. However, its flexibility brings the challenge of selecting the appropriate latent multi-rank. To overcome the issue of pre-selecting the latent multi-rank, we introduce a Bayesian adaptive Tucker decomposition model that infers the multi-rank automatically via an infinite increasing shrinkage prior. The model introduces local sparsity in the core tensor, inducing rich and at the same time parsimonious dependency structures. Posterior inference proceeds via an efficient adaptive Gibbs sampler, supporting both continuous and binary data and allowing for straightforward missing data imputation when dealing with incomplete multiway data. We discuss fundamental properties of the proposed modeling framework, providing theoretical justification. Simulation studies and applications to chemometrics and complex ecological data offer compelling evidence of its advantages over existing tensor factorization methods.

stat.ME

Infinite joint species distribution models

Joint species distribution models are popular in ecology for modeling covariate effects on species occurrence, while characterizing cross-species dependence. Data consist of multivariate binary indicators of the occurrences of different species in each sample, along with sample-specific covariates. A key problem is that current models implicitly assume that the list of species under consideration is predefined and finite, while for highly diverse groups of organisms, it is impossible to anticipate which species will be observed in a study and discovery of unknown species is common. This article proposes a new modeling paradigm for statistical ecology, which generalizes traditional multivariate probit models to accommodate large numbers of rare species and new species discovery. We discuss theoretical properties of the proposed modeling paradigm and implement efficient algorithms for posterior computation. Simulation studies and applications to fungal biodiversity data provide compelling support for the new modeling class.

stat.ME

A hierarchical Bayesian non-asymptotic extreme value model for spatial data

Spatial maps of extreme precipitation are crucial in flood protection. With the aim of producing maps of precipitation return levels, we propose a novel approach to model a collection of spatially distributed time series where the asymptotic assumption, typical of the traditional extreme value theory, is relaxed. We introduce a Bayesian hierarchical model that accounts for the possible underlying variability in the distribution of event magnitudes and occurrences, which are described through latent temporal and spatial processes. Spatial dependence is characterized by geographical covariates and effects not fully described by the covariates are captured by spatial structure in the hierarchies. The performance of the approach is illustrated through simulation studies and an application to daily rainfall extremes across North Carolina (USA). The results show that we significantly reduce the estimation uncertainty with respect to state of the art techniques.

stat.ME