SearcharxivSearch

arXiv subjects

Massimo Ventrucci

Publications and source records attributed to Massimo Ventrucci.

11 recordsLinked to original sources

Bayesian Species Distribution Models using Hierarchical Decomposition Priors

Understanding the relative contributions of environmental, spatial, and temporal processes in shaping species distribution is a central objective in ecology. Bayesian species distribution models (SDMs) offer a flexible framework for this task, yet prior specification for variance components remains challenging. To address this issue, we adapt the Hierarchical Decomposition (HD) prior framework to latent Gaussian SDMs, enabling direct and transparent prior control over variance partitioning. The HD approach reparametrizes variances into a total variance and a set of interpretable proportions, structured through a decomposition tree that reflects both model architecture and ecologically meaningful groupings of effects. We discuss a principled approach for a default tree design tailored to SDMs and a practical workflow for the step-by-step implementation of the method. The framework is illustrated using presence--absence data for 39 demersal fish species from the NOAA Northeast Fisheries Science Center fall bottom trawl survey. Results demonstrate predictive performance comparable to established priors, while providing substantially improved interpretability and transparency in variance attribution and prior sensitivity analysis.

stat.AP

A Standardization Procedure to Incorporate Variance Partitioning Based Priors in Latent Gaussian Models

Latent Gaussian Models (LGMs) are a subset of Bayesian Hierarchical models where Gaussian priors, conditional on variance parameters, are assigned to all effects in the model. LGMs are employed in many fields for their flexibility and computational efficiency. However, practitioners find prior elicitation on the variance parameters challenging because of a lack of intuitive interpretation for them. Recently, several papers have tackled this issue by rethinking the model in terms of variance partitioning (VP) and assigning priors to parameters reflecting the relative contribution of each effect to the total variance. So far, the class of priors based on VP has been mainly deployed for random effects and fixed effects separately. This work presents a novel standardization procedure that expands the applicability of VP priors to a broader class of LGMs, including both fixed and random effects. We describe the steps required for standardization through various examples, with a particular focus on the popular class of intrinsic Gaussian Markov random fields (IGMRFs). The practical advantages of standardization are demonstrated with simulated data and a real dataset on survival analysis.

stat.ME

Informed Bayesian Finite Mixture Models via Asymmetric Dirichlet Priors

Finite mixture models are flexible methods that are commonly used for model-based clustering. A recent focus in the model-based clustering literature is to highlight the difference between the number of components in a mixture model and the number of clusters. The number of clusters is more relevant from a practical stand point, but to date, the focus of prior distribution formulation has been on the number of components. In light of this, we develop a finite mixture methodology that permits eliciting prior information directly on the number of clusters in an intuitive way. This is done by employing an asymmetric Dirichlet distribution as a prior on the weights of a finite mixture. Further, a penalized complexity motivated prior is employed for the Dirichlet shape parameter. We illustrate the ease to which prior information can be elicited via our construction and the flexibility of the resulting induced prior on the number of clusters. We also demonstrate the utility of our approach using numerical experiments and two real world data sets.

stat.ME

A comparison of priors for variance parameters in Bayesian basket trials

Phase II basket trials are popular tools to evaluate efficacy of a new treatment targeting genetic alteration common to a set of different cancer histologies. Efficient designs are obtained by pooling data from the different arms (e.g., cancer histologies) via Bayesian hierarchical modelling, with a variance parameter controlling the strength of shrinkage of each arm treatment effect to the overall treatment effect. One critical aspect of this approach is that prior choice on the variance plays a major role in determining the strength of shrinkage and impacts the operating characteristics of the design. We review the priors most commonly adopted in previous works and compare them with the recently introduced penalized complexity (PC) priors. Our simulation study shows comparable behaviour for the PC prior and the gold standard choice half-t prior, with the former performing better in the homogeneous scenario where all histologies respond similarly to the treatment. We argue that PC priors offer advantages over other priors because they allow the user to handle the degree of shrinkage by means of only one parameter and can be elicited based on clinical opinion when available.

stat.ME

Variance partitioning in spatio-temporal disease mapping models

Bayesian disease mapping, yet if undeniably useful to describe variation in risk over time and space, comes with the hurdle of prior elicitation on hard-to-interpret random effect precision parameters. We introduce a reparametrized version of the popular spatio-temporal interaction models, based on Kronecker product intrinsic Gaussian Markov Random Fields, that we name the variance partitioning (VP) model. The VP model includes a mixing parameter that balances the contribution of the main and interaction effects to the total (generalized) variance and enhances interpretability. The use of a penalized complexity prior on the mixing parameter aids in coding prior information in a intuitive way. We illustrate the advantages of the VP model using two case studies.

stat.ME

A spectral adjustment for spatial confounding

Adjusting for an unmeasured confounder is generally an intractable problem, but in the spatial setting it may be possible under certain conditions. In this paper, we derive necessary conditions on the coherence between the treatment variable of interest and the unmeasured confounder that ensure the causal effect of the treatment is estimable. We specify our model and assumptions in the spectral domain to allow for different degrees of confounding at different spatial resolutions. The key assumption that ensures identifiability is that confounding present at global scales dissipates at local scales. We show that this assumption in the spectral domain is equivalent to adjusting for global-scale confounding in the spatial domain by adding a spatially smoothed version of the treatment variable to the mean of the response variable. Within this general framework, we propose a sequence of confounder adjustment methods that range from parametric adjustments based on the Matern coherence function to more robust semi-parametric methods that use smoothing splines. These ideas are applied to areal and geostatistical data for both simulated and real datasets

stat.ME

A unified view on Bayesian varying coefficient models

Varying coefficient models are useful in applications where the effect of the covariate might depend on some other covariate such as time or location. Various applications of these models often give rise to case-specific prior distributions for the parameter(s) describing how much the coefficients vary. In this work, we introduce a unified view of varying coefficients models, arguing for a way of specifying these prior distributions that are coherent across various applications, avoid overfitting and have a coherent interpretation. We do this by considering varying coefficients models as a flexible extension of the natural simpler model and capitalising on the recently proposed framework of penalized complexity (PC) priors. We illustrate our approach in two spatial examples where varying coefficient models are relevant.

stat.ME

PC priors for residual correlation parameters in one-factor mixed models

Lack of independence in the residuals from linear regression motivates the use of random effect models in many applied fields. We start from the one-way anova model and extend it to a general class of one-factor Bayesian mixed models, discussing several correlation structures for the within group residuals. All the considered group models are parametrized in terms of a single correlation (hyper-)parameter, controlling the shrinkage towards the case of independent residuals (iid). We derive a penalized complexity (PC) prior for the correlation parameter of a generic group model. This prior has desirable properties from a practical point of view: i) it ensures appropriate shrinkage to the iid case; ii) it depends on a scaling parameter whose choice only requires a prior guess on the proportion of total variance explained by the grouping factor; iii) it is defined on a distance scale common to all group models, thus the scaling parameter can be chosen in the same manner regardless the adopted group model. We show the benefit of using these PC priors in a case study in community ecology where different group models are compared.

stat.ME

P-spline smoothing for spatial data collected worldwide

Spatial data collected worldwide at a huge number of locations are frequently used in environmental and climate studies. Spatial modelling for this type of data presents both methodological and computational challenges. In this work we illustrate a computationally efficient non parametric framework to model and estimate the spatial field while accounting for geodesic distances between locations. The spatial field is modelled via penalized splines (P-splines) using intrinsic Gaussian Markov Random Field (GMRF) priors for the spline coefficients. The key idea is to use the sphere as a surrogate for the Globe, then build the basis of B-spline functions on a geodesic grid system. The basis matrix is sparse and so is the precision matrix of the GMRF prior, thus computational efficiency is gained by construction. We illustrate the approach on a real climate study, where the goal is to identify the Intertropical Convergence Zone using high-resolution remote sensing data.

stat.ME

A note on intrinsic Conditional Autoregressive models for disconnected graphs

In this note we discuss (Gaussian) intrinsic conditional autoregressive (CAR) models for disconnected graphs, with the aim of providing practical guidelines for how these models should be defined, scaled and implemented. We show how these suggestions can be implemented in two examples on disease mapping.

stat.ME

Penalized complexity priors for degrees of freedom in Bayesian P-splines

Bayesian P-splines assume an intrinsic Gaussian Markov random field prior on the spline coefficients, conditional on a precision hyper-parameter $τ$. Prior elicitation of $τ$ is difficult. To overcome this issue we aim to building priors on an interpretable property of the model, indicating the complexity of the smooth function to be estimated. Following this idea, we propose Penalized Complexity (PC) priors for the number of effective degrees of freedom. We present the general ideas behind the construction of these new PC priors, describe their properties and show how to implement them in P-splines for Gaussian data.

stat.ME