SearcharxivSearch

arXiv subjects

Andrea Ongaro

Publications and source records attributed to Andrea Ongaro.

6 recordsLinked to original sources

Nonparametric Bayesian inference for the Gini-Simpson index

Many statistical problems concern the analysis of species distributions or, more generally, of discrete labeled quantities. Assessing species diversity constitutes a key step toward understanding population structure, and the Gini-Simpson index is among the most widely adopted diversity measures. In this manuscript, we examine several well-established nonparametric prior models for species frequencies and compare them with a newly proposed distribution within this framework. Specifically, we demonstrate that the conventional symmetric Dirichlet distribution leads to certain undesirable properties in terms of expectation and dispersion. These limitations can be mitigated by adopting an alternative symmetric Dirichlet specification, in which the parameter depends on the number of species. This modified formulation is characterized by analytical tractability and interpretability of its main summaries. Furthermore, when the number of distinct species is infinite, the classical Ferguson Dirichlet process exhibits unsatisfactory behavior compared to the more general Poisson-Dirichlet model. Notably, within this latter model, the posterior mean of the diversity index can be expressed as a convex combination of the optimal classical unbiased estimator and the prior expectation. Theoretical results are further supported by asymptotic analyses and systematically compared with their classical counterparts.

stat.ME

Dependent Dirichlet processes via thinning

When analyzing data from multiple sources, it is often convenient to strike a careful balance between two goals: capturing the heterogeneity of the samples and sharing information across them. We introduce a novel framework to model a collection of samples using dependent Dirichlet processes constructed through a thinning mechanism. The proposed approach modifies the stick-breaking representation of the Dirichlet process by thinning, that is, setting equal to zero a random subset of the beta random variables used in the original construction. This results in a collection of dependent random distributions that exhibit both shared and unique atoms, with the shared ones assigned distinct weights in each distribution. The generality of the construction allows expressing a wide variety of dependence structures among the elements of the generated random vectors. Moreover, its simplicity facilitates the characterization of several theoretical properties and the derivation of efficient computational methods for posterior inference. A simulation study illustrates how a modeling approach based on the proposed process reduces uncertainty in group-specific inferences while preventing excessive borrowing of information when the data indicate it is unnecessary. This added flexibility improves the accuracy of posterior inference, outperforming related state-of-the-art models. An application to the Collaborative Perinatal Project data highlights the model's capability to estimate group-specific densities and uncover a meaningful partition of the observations, both within and across samples, providing valuable insights into the underlying data structure.

stat.ME

BayesChange: an R package for Bayesian Change Point Analysis

We introduce BayesChange, a computationally efficient R package, built on C++, for Bayesian change point detection and clustering of observations sharing common change points. While many R packages exist for change point analysis, BayesChange offers methods not currently available elsewhere. The core functions are implemented in C++ to ensures computational efficiency, while an R user interface simplifies the package usage. The BayesChange package includes two R wrappers that integrate the C++ backend functions, along with S3 methods for summarizing the results. We present the theory beyond each method, the algorithms for posterior simulation and we illustrate the package's usage through synthetic examples.

stat.CO

Model-based clustering of time-dependent observations with common structural changes

We propose a novel model-based clustering approach for samples of time series. We assume as a unique commonality that two observations belong to the same group if structural changes in their behaviours happen at the same time. We resort to a latent representation of structural changes in each time series based on random orders to induce ties among different observations. Such an approach results in a general modeling strategy and can be combined with many time-dependent models known in the literature. Our studies have been motivated by an epidemiological problem, where we want to provide clusters of different countries of the European Union, where two countries belong to the same cluster if the spreading processes of the COVID-19 virus had structural changes at the same time.

stat.ME

Nested Compound Random Measures

Nested nonparametric processes are vectors of random probability measures widely used in the Bayesian literature to model the dependence across distinct, though related, groups of observations. These processes allow a two-level clustering, both at the observational and group levels. Several alternatives have been proposed starting from the nested Dirichlet process by Rodríguez et al. (2008). However, most of the available models are neither computationally efficient or mathematically tractable. In the present paper, we aim to introduce a range of nested processes that are mathematically tractable, flexible, and computationally efficient. Our proposal builds upon Compound Random Measures, which are vectors of dependent random measures early introduced by Griffin and Leisen (2017). We provide a complete investigation of theoretical properties of our model. In particular, we prove a general posterior characterization for vectors of Compound Random Measures, which is interesting per se and still not available in the current literature. Based on our theoretical results and the available posterior representation, we develop the first Ferguson & Klass algorithm for nested nonparametric processes. We specialize our general theorems and algorithms in noteworthy examples. We finally test the model's performance on different simulated scenarios, and we exploit the construction to study air pollution in different provinces of an Italian region (Lombardy). We empirically show how nested processes based on Compound Random Measures outperform other Bayesian competitors.

stat.ME

Contaminated Gibbs-type priors

Gibbs-type priors are widely used as key components in several Bayesian nonparametric models. By virtue of their flexibility and mathematical tractability, they turn out to be predominant priors in species sampling problems, clustering and mixture modelling. We introduce a new family of processes which extend the Gibbs-type one, by including a contaminant component in the model to account for the presence of anomalies (outliers) or an excess of observations with frequency one. We first investigate the induced random partition, the associated predictive distribution and we characterize the asymptotic behaviour of the number of clusters. All the results we obtain are in closed form and easily interpretable, as a noteworthy example we focus on the contaminated version of the Pitman-Yor process. Finally we pinpoint the advantage of our construction in different applied problems: we show how the contaminant component helps to perform outlier detection for an astronomical clustering problem and to improve predictive inference in a species-related dataset, exhibiting a high number of species with frequency one.

stat.ME