SearcharxivSearch

arXiv subjects

Stefano Favaro

Publications and source records attributed to Stefano Favaro.

At least 19 recordsLinked to original sources

Bayesian nonparametric inference for modal missing species and features

Species and feature sampling problems arise naturally whenever each observed unit is associated with one or more labels from a countable alphabet, and inference focuses on the unobserved portion of the distribution over such an alphabet. Within this framework, recent contributions have shifted attention from inference on the total probability mass over unseen labels to distribution-free confidence intervals for the largest unobserved label probability (i.e., the modal missing probability), thereby providing more refined information on whether the missing mass is concentrated on a few high-prevalence unseen labels or distributed across many negligible ones. Besides lacking a unified framework for species and features, these approaches employ a worst-case perspective that causes substantial information loss and produces overly-conservative intervals. We address these limitations through a unified model-based framework for inference on modal missing probabilities in both species and feature settings, which leverages a flexible Bayesian nonparametric formulation to localize uncertainty around models compatible with the observed data. This leads to sharper closed-form credible intervals that effectively exploit prior information and observed data, while preserving the theoretical frequentist properties and robustness of distribution-free intervals. Simulation studies confirm these improvements, while an organized crime application illustrates how our contribution has the potential to reshape law-enforcement decision-making in investigations.

stat.ME

Posterior contraction rates in Sobolev norms and Bayesian derivative estimation for infinite-dimensional exponential families

We study posterior contraction in positive-order Sobolev norms and Bayesian derivative estimation for infinite-dimensional exponential families. We embed the natural parameter in a Hilbert scale and model it via a standard Gaussian series prior expanded in the eigenbasis generating the scale. Under a two-sided link condition on the Fisher information and suitable local regularity assumptions, we show that smoothness-matching priors achieve minimax-optimal posterior contraction rates in any Hilbert scale norm up to the regularity of the ground truth. Our analysis builds on the novel approach to posterior contraction based on the Wasserstein distance recently introduced by Dolera et al. (2024). It combines refined Laplace-type estimates for infinite-dimensional integrals associated to the posterior kernels with a mixed-geometry estimate controlling their stability under fluctuations in the data, itself resting on a tailored Poincar\'e inequality for posterior distributions conditioned on neighbourhoods of the truth. We apply the general theory to density estimation with a logistic parametrisation, Poisson intensity estimation with an exponential link, and the Gaussian white-noise model, yielding minimax contraction rates in Sobolev norms across all three settings. In particular, these yield optimal recovery of density score functions and derivatives of Poisson intensities.

math.ST

The variance of the Pitman--Yor process: a Cifarelli--Regazzini identity and inversion formula

The celebrated Cifarelli--Regazzini identity for the Dirichlet process and its analytic inversion lie at the foundation of an elegant distributional theory for linear functionals of random probability measures, providing, in particular, an explicit formula for the density function of the Dirichlet mean. Nonlinear functionals of the Dirichlet process, by contrast, remain substantially less understood in the literature. In this paper, we develop a transform-and-inversion strategy beyond linear functionals by considering the variance of the Pitman--Yor process, which generalizes the Dirichlet process. We establish a Cifarelli--Regazzini identity for the Pitman--Yor variance, whose direct analytic inversion yields an explicit formula for its density function. The corresponding results for the Dirichlet variance are recovered as a limiting case.

math.PR

Merging of Bayes and quasi-Bayes empirical Bayes procedures for Poisson compound decisions

The Poisson compound decision problem is a long-standing problem in statistics, in which empirical Bayes methods are used to estimate Poisson means under a mixture model. We study this problem from the viewpoint of $g$-modeling, comparing two nonparametric strategies for estimating the unknown mixing distribution: a Bayesian empirical Bayes strategy, based on the Dirichlet process posterior, and a quasi-Bayesian empirical Bayes strategy, based on Newton's algorithm. The latter is computationally attractive, but its relationship with the Bayesian strategy requires theoretical justification. Under a Poisson mixture model with a ``true'', or oracle, mixing distribution, we establish concentration rates for the marginal probability mass functions induced by the Bayesian and quasi-Bayesian estimates. These rates are then translated into rates of decay for the corresponding regrets, interpreted as excess Bayes risks, and used to prove a frequentist merging result between the Bayesian and quasi-Bayesian empirical Bayes strategies. We also extend the analysis to the multidimensional Poisson compound decision problem. Numerical experiments on synthetic data illustrate that the quasi-Bayesian strategy achieves accuracy comparable to the Bayesian strategy, while requiring substantially fewer computational resources, especially in the multidimensional setting.

stat.ME

Bayesian Nonparametric Privacy-Preserving Synthetic Data Generation: I. Discrete Data

Synthetic data generation is a powerful approach to privacy-preserving statistical analysis, where data-release mechanisms are governed by a privacy-utility tradeoff: they should provide privacy guarantees while preserving the statistical utility of confidential data. We develop a Bayesian nonparametric framework for private synthetic data generation tailored to discrete data. Specifically, the confidential data are modeled as a random sample from an unknown discrete distribution endowed with a Pitman-Yor process prior, and synthetic data are generated from the corresponding posterior-predictive distribution. Since the Pitman-Yor process defines an almost surely discrete random probability measure, the resulting mechanism is naturally suited to data with ties and settings involving a potentially large, unknown, or growing number of categories. We study differential privacy guarantees of the Pitman-Yor posterior-predictive mechanism across the three regimes of the discount parameter $\sigma\in(-\infty,1)$. For $\sigma\in(0,1)$, we establish an instance-level $(\varepsilon,\delta)$-differential privacy guarantee. For $\sigma=0$ and $\sigma<0$, corresponding respectively to the Dirichlet process prior and to a parametric Dirichlet-Multinomial model, stronger guarantees are obtained, under suitable conditions on the released sample size. We also investigate statistical utility, or informativity, of the released data via the expected $1$-Wasserstein distance between the empirical distribution of the synthetic data and the "true" data-generating distribution. For $\sigma<0$ and $\sigma=0$, we prove consistency of the empirical distribution in this metric and derive explicit convergence rates, making precise the privacy-utility tradeoff: stronger privacy guarantees impose more restrictive choices of the released sample size, slowing down convergence to the "true" data-generating distribution.

math.ST

Quasi-Bayes empirical Bayes estimation of sums of random variables

The estimation of sums of functions of observable and unobservable variables is a long-standing problem in statistics with applications across many domains. Empirical Bayes methods provide a natural framework for this task under mixture models, but existing approaches often rely on restrictive parametric assumptions or apply only to limited classes of functionals in nonparametric settings. We propose a nonparametric methodology, referred to as quasi-Bayes empirical Bayes, that addresses these limitations through a recursive estimation of the mixing distribution based on Newton's algorithm. The resulting plug-in estimate of the target sum is computationally efficient, scalable, and applicable to a broad class of utility functions, while enabling uncertainty quantification via asymptotic credible intervals derived from a Gaussian central limit theorem. We establish large sample asymptotic theoretical guarantees by proving a merging between the quasi-Bayes and Bayes estimates and by showing consistency under a correctly specified frequentist model. Synthetic-data and real-data analyses demonstrate the practical accuracy and stability of the method, with performance comparable to, and in some cases better than, existing empirical Bayes procedures.

stat.ME

Asymptotic regimes for maximum likelihood estimation in the Ewens--Pitman model: When the strength parameter matters

We study the large sample asymptotic behaviour of the Maximum Likelihood Estimator of the discount and strength parameters $(\alpha,\theta)$ in the Ewens--Pitman model for random partitions, under mild assumptions on the data-generating mechanism. We show that four distinct regimes arise, depending on the limiting behaviour of the frequency spectrum. In particular, in contrast with previous work, we find that $\theta$ may play a crucial role asymptotically. We further show that the existing literature implicitly focuses on only two of these regimes, and we relate this restriction to the constraints imposed by infinite exchangeability. Under the latter, indeed, the number of distinct blocks and the frequency spectrum are necessarily tied by a rigid structural relation. We prove that this lack of flexibility can be overcome through what we call the scaled Ewens--Pitman model, in which $\theta$ is allowed to grow with the sample size $n$. Finally, we provide empirical evidence from real-world data showing that such extensions are needed to capture frequency spectra that fall outside the classical Ewens--Pitman framework.

math.ST

Elements of Conformal Prediction

Predictive inference is a fundamental task in statistics, traditionally addressed using parametric assumptions about the data distribution and detailed analyses of how models learn from data. In recent years, conformal prediction has emerged as an alternative framework that is well suited to modern applications involving high-dimensional data and complex machine learning models. Its appeal stems from being both distribution-free---relying mainly on symmetry assumptions such as exchangeability---and model-agnostic, treating the learning algorithm as a black box. Even under such limited assumptions, conformal prediction provides exact finite-sample guarantees, although these are typically marginal and require careful interpretation. This paper explains the core ideas of conformal prediction and reviews selected methods. Rather than offering an exhaustive survey, it aims to provide a clear conceptual entry point and a pedagogical overview of the field.

stat.ME

Confidence intervals for maximum unseen probabilities, with application to sequential sampling design

Discovery problems often require deciding whether additional sampling is needed to detect all categories whose prevalence exceeds a prespecified threshold. We study this question under a Bernoulli product (incidence) model, where categories are observed only through presence--absence across sampling units. Our inferential target is the \emph{maximum unseen probability}, the largest prevalence among categories not yet observed. We develop nonasymptotic, distribution-free upper confidence bounds for this quantity in two regimes: bounded alphabets (finite and known number of categories) and unbounded alphabets (countably infinite under a mild summability condition). We characterise the limits of data-independent worst-case bounds, showing that in the unbounded regime no nontrivial data-independent procedure can be uniformly valid. We then propose data-dependent bounds in both regimes and establish matching lower bounds demonstrating their near-optimality. We compare empirically the resulting procedures in both simulated and real datasets. Finally, we use these bounds to construct sequential stopping rules with finite-sample guarantees, and demonstrate robustness to contamination that introduces spurious low-prevalence categories.

stat.ME

A Central Limit Theorem for the Ewens-Pitman random partition in the large-$\theta$ regime via a martingale approach

The Ewens-Pitman model defines a distribution on random partitions of $\{1,\ldots,n\}$, with parameters $\alpha \in [0,1)$ and $\theta > -\alpha$; the case $\alpha=0$ reduces to the classical Ewens model from population genetics. We investigate the large-$n$ asymptotic behaviour of the Ewens-Pitman random partition in the nonstandard regime $\theta=\lambda n$ with $\lambda>0$, establishing joint fluctuation results for the total number of blocks $K_n^{\{n\}}$ and the counts $K_{r,n}^{\{n\}}$ of blocks of sizes $r=1,\dots,d$, for fixed $d\in\mathbb{N}$. In particular, for $\alpha\in[0,1)$ and $\theta=\lambda n$, our main result provides a strong law of large numbers and a central limit theorem for the $(d+1)$-dimensional vector $\mathbf{K}_{d,n}^{\{n\}} = \bigl(K_n^{\{n\}}, K_{1,n}^{\{n\}}, \dots, K_{d,n}^{\{n\}}\bigr)^T$ as $n \to \infty$. The proof exploits the Chinese restaurant sequential construction under $\theta=\lambda n$ and a central limit theorem for triangular arrays of martingales, extending techniques previously developed for the classical regime with fixed $\theta$. As corollaries of our results, we recover known asymptotics for $K_n^{\{n\}}$ and derive new strong laws and central limit theorems for each fixed $K_{r,n}^{\{n\}}$, thereby completing earlier weak-law results and providing a comprehensive asymptotic description of the Ewens-Pitman partition structure in the large-$\theta$ setting.

math.PR

A Gaussian process limit for the self-normalized Ewens-Pitman process

For an integer $n\geq1$, consider a random partition $\Pi_{n}$ of $\{1,\ldots,n\}$ into $K_{n}$ partition sets with $K_{r,n}$ partition subsets of size $r=1,\ldots,n$, and assume $\Pi_{n}$ distributed according to the Ewens-Pitman model with parameters $\alpha\in]0,1[$ and $\theta>-\alpha$. Although the large-$n$ asymptotic behaviors of $K_{n}$ and $K_{r,n}$ are well understood in terms of almost sure convergence and Gaussian fluctuations, much less is known about the asymptotic behavior of $P_{r,n}=K_{r,n}/K_n$ and of the self-normalized Ewens-Pitman process $(P_{1,n},P_{2,n},\dots)$. Motivated by the almost sure convergence of $(P_{1,n},P_{2,n},\dots)$ to the Sibuya distribution $p_{\alpha}=(p_{\alpha}(1),p_{\alpha}(2),\ldots)$, where $p_{\alpha}(r)$ is the probability mass at $r=1,2,\ldots$, we establish the $\ell^{2}$ distributional convergence \begin{displaymath} \sqrt{K_{n}}((P_{1,n},\,P_{2,n},\ldots)-p_{\alpha})\underset{n\rightarrow+\infty}{\overset{\cL}{\longrightarrow}}\mathcal{G}(\Gamma_\alpha), \end{displaymath} where $\mathcal{G}(\Gamma_\alpha)$ stands for a centered Gaussian process with covariance matrix $\Gamma_\alpha=diag(p_{\alpha}) - p_{\alpha} p_{\alpha}^T$. We apply our result to the estimation of the parameter

math.PR

Conformal Inference for Open-Set and Imbalanced Classification

This paper presents a conformal prediction method for classification in highly imbalanced and open-set settings, where there are many possible classes and not all may be represented in the data. Existing approaches require a finite, known label space and typically involve random sample splitting, which works well when there is a sufficient number of observations from each class. Consequently, they have two limitations: (i) they fail to provide adequate coverage when encountering new labels at test time, and (ii) they may become overly conservative when predicting previously seen labels. To obtain valid prediction sets in the presence of unseen labels, we compute and integrate into our predictions a new family of conformal p-values that can test whether a new data point belongs to a previously unseen class. We study these p-values theoretically, establishing their optimality, and uncover an intriguing connection with the classical Good--Turing estimator for the probability of observing a new species. To make more efficient use of imbalanced data, we also develop a selective sample splitting algorithm that partitions training and calibration data based on label frequency, leading to more informative predictions. Despite breaking exchangeability, this allows maintaining finite-sample guarantees through suitable re-weighting. With both simulated and real data, we demonstrate our method leads to prediction sets with valid coverage even in challenging open-set scenarios with infinite numbers of possible labels, and produces more informative predictions under extreme class imbalance.

stat.ML

Large-scale entity resolution via microclustering Ewens--Pitman random partitions

We introduce the microclustering Ewens--Pitman model for random partitions, obtained by scaling the strength parameter of the Ewens--Pitman model linearly with the sample size. The resulting random partition is shown to have the microclustering property, namely: the size of the largest cluster grows sub-linearly with the sample size, while the number of clusters grows linearly. By leveraging the interplay between the Ewens--Pitman random partition with the Pitman--Yor process, we develop efficient variational inference schemes for posterior computation in entity resolution. Our approach achieves a speed-up of three orders of magnitude over existing Bayesian methods for entity resolution, while maintaining competitive empirical performance.

stat.ME

Online activity prediction via generalized Indian buffet process models

Online A/B tests are the standard tool for data-driven decision-making at scale. Among the design choices with the largest impact on statistical power is the triggering mechanism: how many users to expose and for how long. This often requires forecasting user engagement, i.e., whether enough users will trigger, and when a target participation level will be reached, from limited pilot data. We introduce a Bayesian nonparametric model for predicting both new-user counts and total triggers, accommodating the heavy-tailed engagement patterns typical of web experiments. All predictive quantities can be computed without intensive numerical procedures such as MCMC or variational inference. We evaluate on three public datasets (over 450 public benchmark evaluations) and 1,774 proprietary A/B tests. In all the settings, our models show improved accuracy in forecasting new users, total triggers, and time to reach a target sample size compared with state-ofthe-art competitors, especially when only a few pilot days are observed.

stat.AP

A new look on large deviations and concentration inequalities for the Ewens-Pitman model

The Ewens-Pitman model is a probability distribution for random partitions of the set $[n]=\{1,\ldots,n\}$, parameterized by $\alpha\in[0,1)$ and $\theta>-\alpha$, with $\alpha=0$ corresponding to the Ewens model in population genetics. The goal of this paper is to provide an alternative and concise proof of the Feng-Hoppe large deviation principle for the number $K_{n}$ of partition sets in the Ewens-Pitman model with $\alpha\in(0,1)$ and $\theta>-\alpha$. Our approach leverages an integral representation of the moment-generating function of $K_{n}$ in terms of the (one-parameter) Mittag-Leffler function, along with a sharp asymptotic expansion of it. This approach significantly simplifies the original proof of Feng-Hoppe large deviation principle, as it avoids all the technical difficulties arising from a continuity argument with respect to rational and non-rational values of $\alpha$. Beyond large deviations for $K_{n}$, our approach allows to establish a sharp concentration inequality for $K_n$ involving the rate function of the large deviation principle.

math.PR

Student-t processes as infinite-width limits of posterior Bayesian neural networks

The asymptotic properties of Bayesian Neural Networks (BNNs) have been extensively studied, particularly regarding their approximations by Gaussian processes in the infinite-width limit. We extend these results by showing that posterior BNNs can be approximated by Student-t processes, which offer greater flexibility in modeling uncertainty. Specifically, we show that, if the parameters of a BNN follow a Gaussian prior distribution, and the variance of both the last hidden layer and the Gaussian likelihood function follows an Inverse-Gamma prior distribution, then the resulting posterior BNN converges to a Student-t process in the infinite-width limit. Our proof leverages the Wasserstein metric to establish control over the convergence rate of the Student-t process approximation.

stat.ML

Gaussian credible intervals in Bayesian nonparametric estimation of the unseen

The unseen-species problem assumes $n\geq1$ samples from a population of individuals belonging to different species, possibly infinite, and calls for estimating the number $K_{n,m}$ of hitherto unseen species that would be observed if $m\geq1$ new samples were collected from the same population. This is a long-standing problem in statistics, which has gained renewed relevance in biological and physical sciences, particularly in settings with large values of $n$ and $m$. In this paper, we adopt a Bayesian nonparametric approach to the unseen-species problem under the Pitman-Yor prior, and propose a novel methodology to derive large $m$ asymptotic credible intervals for $K_{n,m}$, for any $n\geq1$. By leveraging a Gaussian central limit theorem for the posterior distribution of $K_{n,m}$, our method improves upon competitors in two key aspects: firstly, it enables the full parameterization of the Pitman-Yor prior, including the Dirichlet prior; secondly, it avoids the need of Monte Carlo sampling, enhancing computational efficiency. We validate the proposed method on synthetic and real data, demonstrating that it improves the empirical performance of competitors by significantly narrowing the gap between asymptotic and exact credible intervals for any $m\geq1$.

stat.ME

Asymptotic behavior of clusters in hierarchical species sampling models

Consider a random sample of size $N$ from a hierarchical species sampling model. In this paper, we study the large $N$ asymptotic behavior of the number ${\bf K}_N$ of clusters in the random sample and the number ${\bf \widetilde M}_{\ell,N}$ of clusters represented by exactly $\ell$ latent first-level clusters in the first level of the hierarchical model. In particular, we establish almost sure and $L^p$ convergence for ${\bf \widetilde M}_{\ell,N}$, Gaussian fluctuations and the law of the iterated logarithm for ${\bf K}_N$, and large deviation principles for both ${\bf K}_N$ and ${\bf \widetilde M}_{\ell,N}$. Our approach relies on a random sample size (or random index) representation of the number of clusters through the corresponding non-hierarchical species sampling model.

math.PR