SearcharxivSearch

arXiv subjects

Gianluca Mastrantonio

Publications and source records attributed to Gianluca Mastrantonio.

At least 19 recordsLinked to original sources

An interpretable family of projected normal distributions and a related copula model for Bayesian analysis of hypertoroidal data

This paper introduces two families of probability distributions for Bayesian analysis of hypertoroidal data. The first family consists of symmetric distributions derived from the projection of multivariate normal distributions under specific parameter constraints. This family is closed under marginalization and hence any marginal distribution belongs to a lower-dimensional case of the same family. In particular the univariate marginal of the family is the unimodal case of the projected normal distribution on the circle. The second family is a flexible extension of the copula case of the first family, which can accommodate any univariate marginal distributions. Unlike existing models derived via projection, both families have the common advantage that their parameters possess a clear and intuitive interpretation. The use of latent variables simplifies Bayesian estimation using Markov chain Monte Carlo algorithms. The usefulness of the proposed families is demonstrated through the analysis of a meteorological dataset.

stat.ME

Integrating Deep Learning and Spatial Statistics in Marine Ecosystem Monitoring

In ecology, photogrammetry is a crucial method for efficiently collecting non-destructive samples of natural environments. When estimating the spatial distribution of animals, detecting objects in large-scale images becomes crucial. Object detection models enable large-scale analysis but introduce uncertainty because detection probability depends on various factors. To address detection bias, we model the distribution of a species of benthic animals (holothurians) in an area of the Italian Tyrrhenian coast near Giglio Island using a Thinned Log-Gaussian Cox Process (LGCP). We assume that a "true" intensity function accurately describes the distribution, while the observed process, resulting from independent thinning, is represented by a degraded intensity. The detection function controls the thinning mechanism, influenced by the object's location and other detection-related features. We use manual identification of holothurians as our benchmark. We compare automatic detection with this benchmark, an unthinned LGCP, and the thinned model to highlight the improvements gained from the proposed approach.Our method allows researchers to use photogrammetry, automatically identify objects of interest, and correct biases and approximations caused by the observation process.

stat.OT

A new hierarchical distribution on arbitrary sparse precision matrices

We introduce a general strategy for defining distributions over the space of sparse symmetric positive definite matrices. Our method utilizes the Cholesky factorization of the precision matrix, imposing sparsity through constraints on its elements while preserving their independence and avoiding the numerical evaluation of normalization constants. In particular, we develop the S-Bartlett as a modified Bartlett decomposition, recovering the standard Wishart as a particular case. By incorporating a Spike-and-Slab prior to model graph sparsity, our approach facilitates Bayesian estimation through a tailored MCMC routine based on a Dual Averaging Hamiltonian Monte Carlo update. This framework extends naturally to the Generalized Linear Model setting, enabling applications to non-Gaussian outcomes via latent Gaussian variables. We test and compare the proposed S-Bartelett prior with the G-Wishart both on simulated and real data. Results highlight that the S-Bartlett prior offers a flexible alternative for estimating sparse precision matrices, with potential applications across diverse fields.

stat.ME

Modelling benthic animals in space and time using Bayesian Point Process with cross validation: the case of Holoturians

Understanding the spatial distribution of Holothurians is an essential task for ecosystem monitoring and sustainable management, particularly in the Mediterranean habitats. However, species distribution modeling is often complicated by the presence-only nature of the data and heterogeneous sampling designs. This study develops a spatio-temporal framework based on Log-Gaussian Cox Processes to analyze Holothurians' positions collected across nine survey campaigns conducted from 2022 to 2024 near Giglio Island, Italy. The surveys combined high-resolution photogrammetry with diver-based visual censuses, leading to varying detection probabilities across habitats, especially within Posidonia oceanica meadows. We adopt a model with a shared spatial Gaussian process component to accommodate this complexity, accounting for habitat structure, environmental covariates, and temporal variability. Model estimation is performed using Integrated Nested Laplace Approximation. We evaluate the predictive performances of alternative model specifications through a novel k-fold cross-validation strategy for point processes, using the Continuous Ranked Probability Score. Our approach provides a flexible and computationally efficient framework for integrating heterogeneous presence-only data in marine ecology and comparing the predictive ability of alternative models.

stat.AP

BayVel: A Bayesian Framework for RNA Velocity Estimation in Single-Cell Transcriptomics

RNA velocity is a model of gene expression dynamics designed to analyze single-cell RNA sequencing (scRNA-seq) data, and it has recently gained significant attention. However, despite its popularity, the model has raised several concerns, primarily related to three issues: its heavy dependence on data preprocessing, the need for post-processing of the results, and the limitations of the underlying statistical methodology. Current approaches, such as scVelo, suffer from notable statistical shortcomings. These include identifiability problems, reliance on heuristic preprocessing steps, and the absence of uncertainty quantification. To address these limitations, we propose BayVel, a Bayesian hierarchical model that directly models raw count data. BayVel resolves identifiability issues and provides posterior distributions for all parameters, including the RNA velocities themselves, without the need for any post processing. We evaluate BayVel's performance using simulated datasets. While scVelo fails to accurately reconstruct parameters, even when data are simulated directly from the model assumptions, BayVel demonstrates strong accuracy and robustness. This highlights BayVel as a statistically rigorous and reliable framework for studying transcriptional dynamics in the context of RNA velocity modeling. When applied to a real dataset of pancreatic epithelial cells previously analyzed with scVelo, BayVel does not replicate their findings, which appears to be strongly influenced by the postprocessing, supporting concerns raised in other studies about the reliability of scVelo.

stat.AP

Estimating the optimal time to perform a PET-PSMA exam in prostatectomized patients based on data from clinical practice

Prostatectomized patients are at risk of resurgence, and for this reason, during a follow-up period, they are monitored for Prostate Specific Antigen (PSA) growth, an indicator of tumor progression. The presence of tumors can be evaluated with an expensive exam, called Positron Emission Tomography with Prostate-Specific Membrane Antigen (PET-PSMA). To justify the high cost of the PET-PSMA and, at the same time, to contain the risk for the patient, this exam should be recommended only when the evidence of tumor progression is strong. With the aim of estimating the optimal time to recommend the exam based on the patient's history and collected data, we build a hierarchical Bayesian model that describes, jointly, the PSA growth curve and the probability of a positive PET-PSMA. With our proposal we process all past and present information about the patients PSA measurement and PET-PSMA results, in order to give an informed estimate of the optimal time, improving current practice.

stat.ME

Bayesian size-and-shape regression modelling

Building on Dryden et al. (2021), this note presents the Bayesian estimation of a regression model for size-and-shape response variables with Gaussian landmarks. Our proposal fits into the framework of Bayesian latent variable models and allows a highly flexible modelling framework.

stat.ME

Bayesian inference of Latent Spectral Shapes

This paper proposes a hierarchical spatial-temporal model for modelling the spectrograms of animal calls. The motivation stems from analyzing recordings of the so-called grunt calls emitted by various lemur species. Our goal is to identify a latent spectral shape that characterizes each species and facilitates measuring dissimilarities between them. The model addresses the synchronization of animal vocalizations, due to varying time-lengths and speeds, with non-stationary temporal patterns and accounts for periodic sampling artifacts produced by the time discretization of analog signals. The former is achieved through a synchronization function, and the latter is modeled using a circular representation of time. To overcome the curse of dimensionality inherent in the model's implementation, we employ the Nearest Neighbor Gaussian Process, and posterior samples are obtained using the Markov Chain Monte Carlo method. We apply the model to a real dataset comprising sounds from 8 different species. We define a representative sound for each species and compare them using a simple distance measure. Cross-validation is used to evaluate the predictive capability of our proposal and explore special cases. Additionally, a simulation example is provided to demonstrate that the algorithm is capable of retrieving the true parameters.

stat.AP

The modeling of multiple animals that share behavioral features

In this work, we propose a model that can be used to infer the behavior of multiple animals. Our proposal is defined as a set of hidden Markov models that are based on the sticky hierarchical Dirichlet process, with a shared base-measure, and a STAP emission distribution. The latent classifications are representative of the behavior assumed by the animals, which is described by the STAP parameters. Given the latent classifications, the animals are independent. As a result of the way we formalize the distribution over the STAP parameters, the animals may share, in different behaviors, the set or a subset of the parameters, thereby allowing us to investigate the similarities between them. The hidden Markov models, based on the Dirichlet process, allow us to estimate the number of latent behaviors for each animal, as a model parameter. This proposal is motivated by a real data problem, where the GPS coordinates of six Maremma Sheepdogs have been observed. Among the other results, we show that four dogs share most of the behavior characteristics, while two have specific behaviors.

stat.AP

Modeling animal movement with directional persistence and attractive points

GPS technology is currently easily accessible to researchers, and many animal movement datasets are available. Two of the main features that a model which describes an animal's path can possess are directional persistence and attraction to a point in space. In this work, we propose a new approach that can have both characteristics. Our proposal is a hidden Markov model with a new emission distribution. The emission distribution models the two aforementioned characteristics, while the latent state of the hidden Markov model is needed to account for the behavioral modes. We show that the model is easy to implement in a Bayesian framework. We estimate our proposal on the motivating data that represent GPS locations of a Maremma Sheepdog recorded in Australia. The obtained results are easily interpretable and we show that our proposal outperforms the main competitive model.

stat.AP

CircSpaceTime: an R package for spatial and spatio-temporal modeling of Circular data

CircSpaceTime is the only R package currently available that implements Bayesian models for spatial and spatio-temporal interpolation of circular data. Such data are often found in applications where, among the many, wind directions, animal movement directions, and wave directions are involved. To analyze such data we need models for observations at locations s and times t, as the so-called geostatistical models, providing structured dependence assumed to decay in distance and time. The approach we take begins with Gaussian processes defined for linear variables over space and time. Then, we use either wrapping or projection to obtain processes for circular data. The models are cast as hierarchical, with fitting and inference within a Bayesian framework. Altogether, this package implements work developed by a series of papers; the most relevant being Jona Lasinio, Gelfand, and Jona Lasinio (2012); Wang and Gelfand (2014); Mastrantonio, Jona Lasinio, and Gelfand (2016). All procedures are written using Rcpp. Estimates are obtained by MCMC allowing parallelized multiple chains run. The implementation of the proposed models is considerably improved on the simple routines adopted in the research papers. As original running examples, for the spatial and spatio-temporal settings, we use wind directions datasets over central Italy.

stat.AP

New formulation of the Logistic-Gaussian process to analyze trajectory tracking data

Improved communication systems, shrinking battery sizes and the price drop of tracking devices have led to an increasing availability of trajectory tracking data. These data are often analyzed to understand animal behavior. In this work, we propose a new model for interpreting the animal movent as a mixture of characteristic patterns, that we interpret as different behaviors. The probability that the animal is behaving according to a specific pattern, at each time instant, is non-parametrically estimated using the Logistic-Gaussian process. Owing to a new formalization and the way we specify the coregionalization matrix of the associated multivariate Gaussian process, our model is invariant with respect to the choice of the reference element and of the ordering of the probability vector components. We fit the model under a Bayesian framework, and show that the Markov chain Monte Carlo algorithm we propose is straightforward to implement. We perform a simulation study with the aim of showing the ability of the estimation procedure to retrieve the model parameters. We also test the performance of the information criterion we used to select the number of behaviors. The model is then applied to a real dataset where a wolf has been observed before and after procreation. The results are easy to interpret, and clear differences emerge in the two phases.

stat.AP

Semi-parametric Bayesian change-point model based on the Dirichlet process

In this work we introduce a semi-parametric Bayesian change-point model, defining its time dynamic as a latent Markov process based on the Dirichlet process. We treat the number of change point as a random variable and we estimate it during model fitting. Posterior inference is carried out using a Markov chain Monte Carlo algorithm based on a marginalized version of the proposed model. The model is illustrated using simulated examples and two real datasets, namely the coal- mining disasters, that is a widely used dataset for illustrative purpose, and a dataset of indoor radon recordings. With the simulated examples we show that the model is able to recover the parameters and number of change points, and we compare our results with the ones of the-state- of-the-art models, showing a clear improvement in terms of change points identification. The results obtained on the coal-mining disasters and radon data are coherent with previous literature.

stat.CO

A Hierarchical Multivariate Spatio-Temporal Model for Large Clustered Climate data with Annual Cycles

We present a multivariate hierarchical space-time model to describe the joint series of monthly extreme temperatures and amounts of rainfall. Data are available for 360 monitoring stations over 60 years, with missing data affecting almost all series. Model components account for spatio-temporal dependence with annual cycles, dependence on covariates and between responses. The very large amount of data is tackled modeling the spatio-temporal dependence by the nearest neighbor Gaussian process. Response multivariate dependencies are described using the linear model of coregionalization, while annual cycles are assessed by a circular representation of time. The proposed approach allows imputation of missing values and easy interpolation of climate surfaces at the national level. The motivation behind is the characterization of the so called ecoregions over the Italian territory. Ecoregions delineate broad and discrete ecologically homogeneous areas of similar potential as regards the climate, physiography, hydrography, vegetation and wildlife, and provide a geographic framework for interpreting ecological processes, disturbance regimes, vegetation patterns and dynamics. To now, the two main Italian macro-ecoregions are hierarchically arranged into 35 zones. The current climatic characterization of Italian ecoregions is based on data and bioclimatic indices for the period 1955-1985 and requires an appropriate update.

stat.AP

The joint projected normal and skew-normal: a distribution for poly-cylindrical data

The contribution of this work is the introduction of a multivariate circular-linear (or poly- cylindrical) distribution obtained by combining the projected and the skew-normal. We show the flexibility of our proposal, its property of closure under marginalization and how to quantify multivariate dependence. Due to a non-identifiability issue that our proposal inherits from the projected normal, a compu- tational problem arises. We overcome it in a Bayesian framework, adding suitable latent variables and showing that posterior samples can be obtained with a post-processing of the estimation algo- rithm output. Under specific prior choices, this approach enables us to implement a Markov chain Monte Carlo algorithm relying only on Gibbs steps, where the updates of the parameters are done as if we were working with a multivariate normal likelihood. The proposed approach can be also used with the projected normal. As a proof of concept, on simulated examples we show the ability of our algorithm in recovering the parameters values and to solve the identification problem. Then the proposal is used in a real data example, where the turning-angles (circular variables) and the logarithm of the step-lengths (linear variables) of four zebras are jointly modelled.

stat.ME

Distributions-oriented wind forecast verification by a hidden Markov model for multivariate circular-linear data

Winds from the North-West quadrant and lack of precipitation are known to lead to an increase of PM10 concentrations over a residential neighborhood in the city of Taranto (Italy). In 2012 the local government prescribed a reduction of industrial emissions by 10% every time such meteorological conditions are forecasted 72 hours in advance. Wind forecasting is addressed using the Weather Research and Forecasting (WRF) atmospheric simulation system by the Regional Environmental Protection Agency. In the context of distributions-oriented forecast verification, we propose a comprehensive model-based inferential approach to investigate the ability of the WRF system to forecast the local wind speed and direction allowing different performances for unknown weather regimes. Ground-observed and WRF-forecasted wind speed and direction at a relevant location are jointly modeled as a 4-dimensional time series with an unknown finite number of states characterized by homogeneous distributional behavior. The proposed model relies on a mixture of joint projected and skew normal distributions with time-dependent states, where the temporal evolution of the state membership follows a first order Markov process. Parameter estimates, including the number of states, are obtained by a Bayesian MCMC-based method. Results provide useful insights on the performance of WRF forecasts in relation to different combinations of wind speed and direction.

stat.AP

The wrapped skew Gaussian process for analyzing spatio-temporal data

We consider modeling of angular or directional data viewed as a linear variable wrapped onto a unit circle. In particular, we focus on the spatio-temporal context, motivated by a collection of wave directions obtained as computer model output developed dynamically over a collection of spatial locations. We propose a novel wrapped skew Gaussian process which enriches the class of wrapped Gaussian process. The wrapped skew Gaussian process enables more flexible marginal distributions than the symmetric ones arising under the wrapped Gaussian process and it allows straightforward interpretation of parameters. We clarify that replication through time enables criticism of the wrapped process in favor of the wrapped skew process. We formulate a hierarchical model incorporating this process and show how to introduce appropriate latent variables in order to enable efficient fitting to dynamic spatial directional data. We also show how to implement kriging and forecasting under this model. We provide a simulation example as a proof of concept as well as a real data example. Both examples reveal consequential improvement in predictive performance for the wrapped skew Gaussian specification compared with the earlier wrapped Gaussian version.

stat.ME

Spatio-temporal circular models with non-separable covariance structure

Circular data arise in many areas of application. Recently, there has been interest in looking at circular data collected separately over time and over space. Here, we extend some of this work to the spatio-temporal setting, introducing space-time dependence. We accommodate covariates, implement full kriging and forecasting, and also allow for a nugget which can be time dependent. We work within a Bayesian framework, introducing suitable latent variables to facilitate Markov chain Monte Carlo (MCMC) model fitting. The Bayesian framework enables us to implement full inference, obtaining predictive distributions for kriging and forecasting. We offer comparison between the less flexible but more interpretable wrapped Gaussian process and the more flexible but less interpretable projected Gaussian process. We do this illustratively using both simulated data and data from computer model output for wave directions in the Adriatic Sea off the coast of Italy.

stat.ME