SearcharxivSearch

arXiv subjects

Pietro Rigo

Publications and source records attributed to Pietro Rigo.

At least 19 recordsLinked to original sources

Some cautionary tales about Bayesian predictive inference

Two misunderstandings, frequently arising in Bayesian predictive inference, are discussed. The first deals with the data generating mechanism, while the second consists in overestimating the role played by asymptotic exchangeability. Some consequences of such misunderstandings are highlighted through examples.

math.ST

Weak convergence of predictive distributions

Let $(X_n)$ be a sequence of random variables with values in a standard Borel space $S$. We investigate the condition \begin{gather}\label{x56w1q} E\bigl\{f(X_{n+1})\mid X_1,\ldots,X_n\bigr\}\,\quad\text{converges in probability,}\tag{*} \\\text{as }n\rightarrow\infty,\text{ for each bounded Borel function }f:S\rightarrow\mathbb{R}.\notag \end{gather} Some consequences of \eqref{x56w1q} are highlighted and various sufficient conditions for it are obtained. In particular, \eqref{x56w1q} is characterized in terms of stable convergence. Since \eqref{x56w1q} holds whenever $(X_n)$ is conditionally identically distributed, three weak versions of the latter condition are investigated as well. For each of such versions, our main goal is proving (or disproving) that \eqref{x56w1q} holds. Several counterexamples are given.

math.PR

Bayesian nonparametric inference on a Fr\'echet class

Let $(\mathcal{X},\mathcal{F},\mu)$ and $(\mathcal{Y},\mathcal{G},\nu)$ be probability spaces and $(Z_n)$ a sequence of random variables with values in $(\mathcal{X}\times\mathcal{Y},\,\mathcal{F}\otimes\mathcal{G})$. Let $\Gamma(\mu,\nu)$ be the collection of all probability measures $p$ on $\mathcal{F}\otimes\mathcal{G}$ such that $$p\bigl(A\times\mathcal{Y}\bigr)=\mu(A)\quad\text{and}\quad p\bigl(\mathcal{X}\times B\bigr)=\nu(B)\quad\text{for all }A\in\mathcal{F}\text{ and }B\in\mathcal{G}.$$ In this paper, we build some probability measures $\Pi$ on $\Gamma(\mu,\nu)$. In addition, for each such $\Pi$, we assume that $(Z_n)$ is exchangeable with de Finetti's measure $\Pi$ and we evaluate the conditional distribution $\Pi(\cdot\mid Z_1,\ldots,Z_n)$. In Bayesian nonparametrics, if $(Z_1,\ldots, Z_n)$ are the available data, $\Pi$ and $\Pi(\cdot\mid Z_1,\ldots, Z_n)$ can be regarded as the prior and the posterior, respectively. To support this interpretation, it suffices to think of a problem where the unknown probability distribution of some bivariate phenomenon is constrained to have marginals $\mu$ and $\nu$. Finally, analogous results are obtained for the set $\Gamma(\mu)$ of those probability measures on $\mathcal{F}\otimes\mathcal{G}$ with marginal $\mu$ on $\mathcal{F}$ (but arbitrary marginal on $\mathcal{G}$). That is, we introduce some priors on $\Gamma(\mu)$ and we evaluate the corresponding posteriors.

stat.ME

Knockoffs for exchangeable categorical covariates

Let $X=(X_1,\ldots,X_p)$ be a $p$-variate random vector and $F$ a fixed finite set. In a number of applications, mainly in genetics, it turns out that $X_i\in F$ for each $i=1,\ldots,p$. Despite the latter fact, to obtain a knockoff $\widetilde{X}$ (in the sense of \cite{CFJL18}), $X$ is usually modeled as an absolutely continuous random vector. While comprehensible from the point of view of applications, this approximate procedure does not make sense theoretically, since $X$ is supported by the finite set $F^p$. In this paper, explicit formulae for the joint distribution of $(X,\widetilde{X})$ are provided when $P(X\in F^p)=1$ and $X$ is exchangeable or partially exchangeable. In fact, when $X_i\in F$ for all $i$, there seem to be various reasons for assuming $X$ exchangeable or partially exchangeable. The robustness of $\widetilde{X}$, with respect to the de Finetti's measure $\pi$ of $X$, is investigated as well. Let $\mathcal{L}_\pi(\widetilde{X}\mid X=x)$ denote the conditional distribution of $\widetilde{X}$, given $X=x$, when the de Finetti's measure is $\pi$. It is shown that $$\norm{\mathcal{L}_{\pi_1}(\widetilde{X}\mid X=x)-\mathcal{L}_{\pi_2}(\widetilde{X}\mid X=x)}\le c(x)\,\norm{\pi_1-\pi_2}$$ where $\norm{\cdot}$ is total variation distance and $c(x)$ a suitable constant. Finally, a numerical experiment is performed. Overall, the knockoffs of this paper outperform the alternatives (i.e., the knockoffs obtained by giving $X$ an absolutely continuous distribution) as regards the false discovery rate but are slightly weaker in terms of power.

math.ST

Asymptotics of predictive distributions driven by sample means and variances

Let $\alpha_n(\cdot)=P\bigl(X_{n+1}\in\cdot\mid X_1,\ldots,X_n\bigr)$ be the predictive distributions of a sequence $(X_1,X_2,\ldots)$ of $p$-dimensional random vectors. Suppose $$\alpha_n= \mathcal{N} _p (M_n,Q_n)$$ where $M_n=\frac{1}{n}\sum_{i=1}^nX_i$ and $Q_n=\frac{1}{n}\sum_{i=1}^n(X_i-M_n)(X_i-M_n)^t$. Then, there is a random probability measure $\alpha$ on the Borel subsets of $\mathbb{R}^p$ such that $\lVert\alpha_n-\alpha\rVert\overset{a.s.}\longrightarrow 0$ where $\lVert\cdot\rVert$ is total variation distance. An explicit expression for $\alpha$ is provided and the convergence rate of $\lVert\alpha_n-\alpha\rVert$ is shown to be arbitrarily close to $n^{-1/2}$. Moreover, it is still true that $\lVert\alpha_n-\alpha\rVert\overset{a.s.}\longrightarrow 0$ even if $\alpha_n=\mathcal{L}(M_n,Q_n)$ where $\mathcal{L}$ belongs to a class of distributions much larger than the normal. The predictives $\alpha_n$ are useful in various frameworks, including Bayesian predictive inference and predictive resampling. Finally, the asymptotic behavior of copula-based predictive distributions (introduced in [13]) is investigated and a numerical experiment is performed.

math.ST

Generating knockoffs via conditional independence

Let $X$ be a $p$-variate random vector and $\widetilde{X}$ a knockoff copy of $X$ (in the sense of \cite{CFJL18}). A new approach for constructing $\widetilde{X}$ (henceforth, NA) has been introduced in \cite{JSPI}. NA has essentially three advantages: (i) To build $\widetilde{X}$ is straightforward; (ii) The joint distribution of $(X,\widetilde{X})$ can be written in closed form; (iii) $\widetilde{X}$ is often optimal under various criteria. However, for NA to apply, $X_1,\ldots, X_p$ should be conditionally independent given some random element $Z$. Our first result is that any probability measure $μ$ on $\mathbb{R}^p$ can be approximated by a probability measure $μ_0$ of the form $$μ_0\bigl(A_1\times\ldots\times A_p\bigr)=E\Bigl\{\prod_{i=1}^p P(X_i\in A_i\mid Z)\Bigr\}.$$ The approximation is in total variation distance when $μ$ is absolutely continuous, and an explicit formula for $μ_0$ is provided. If $X\simμ_0$, then $X_1,\ldots,X_p$ are conditionally independent. Hence, with a negligible error, one can assume $X\simμ_0$ and build $\widetilde{X}$ through NA. Our second result is a characterization of the knockoffs $\widetilde{X}$ obtained via NA. It is shown that $\widetilde{X}$ is of this type if and only if the pair $(X,\widetilde{X})$ can be extended to an infinite sequence so as to satisfy certain invariance conditions. The basic tool for proving this fact is de Finetti's theorem for partially exchangeable sequences. In addition to the quoted results, an explicit formula for the conditional distribution of $\widetilde{X}$ given $X$ is obtained in a few cases. In one of such cases, it is assumed $X_i\in\{0,1\}$ for all $i$.

math.ST

Some duality results for equivalence couplings and total variation

Let $(Ω,\mathcal{F})$ be a standard Borel space and $\mathcal{P}(\mathcal{F})$ the collection of all probability measures on $\mathcal{F}$. Let $E\subsetΩ\timesΩ$ be a measurable equivalence relation, that is, $E\in\mathcal{F}\otimes\mathcal{F}$ and the relation on $Ω$ defined as $x\sim y$ $\Leftrightarrow$ $(x,y)\in E$ is reflexive, symmetric and transitive. It is shown that there are two $σ$-fields $\mathcal{G}_0$ and $\mathcal{G}_1$ on $Ω$ such that, for all $μ,\,ν\in\mathcal{P}(\mathcal{F})$, $$\inf_{P\inΓ(μ,ν)}(1-P(E))=\norm{μ-ν}_{\mathcal{G}_1}\quad\text{and}\quad\min_{P\inΓ(μ,ν_0)}(1-P(E))=\norm{μ-ν}_{\mathcal{G}_0}.$$ Here, $ν_0\in\mathcal{P}(\mathcal{F})$ is a suitable probability measure satisfying $ν_0=ν$ on $\mathcal{G}_0$. Moreover, $\mathcal{G}_0\subset\mathcal{F}$ while $\mathcal{G}_1\subset\widehat{\mathcal{F}}$, where $\widehat{\mathcal{F}}$ is the universally measurable $σ$-field with respect to $\mathcal{F}$. However, for all $μ,\,ν\in\mathcal{P}(\mathcal{F})$, there is a $σ$-field $\mathcal{G}(μ,ν)\subset\mathcal{F}$ such that $$\inf_{P\inΓ(μ,ν)}(1-P(E))=\norm{μ-ν}_{\mathcal{G}(μ,ν)}.$$

math.PR

A probabilistic view on predictive constructions for Bayesian learning

Given a sequence $X=(X_1,X_2,\ldots)$ of random observations, a Bayesian forecaster aims to predict $X_{n+1}$ based on $(X_1,\ldots,X_n)$ for each $n\ge 0$. To this end, in principle, she only needs to select a collection $σ=(σ_0,σ_1,\ldots)$, called ``strategy" in what follows, where $σ_0(\cdot)=P(X_1\in\cdot)$ is the marginal distribution of $X_1$ and $σ_n(\cdot)=P(X_{n+1}\in\cdot\mid X_1,\ldots,X_n)$ the $n$-th predictive distribution. Because of the Ionescu-Tulcea theorem, $σ$ can be assigned directly, without passing through the usual prior/posterior scheme. One main advantage is that no prior probability is to be selected. In a nutshell, this is the predictive approach to Bayesian learning. A concise review of the latter is provided in this paper. We try to put such an approach in the right framework, to make clear a few misunderstandings, and to provide a unifying view. Some recent results are discussed as well. In addition, some new strategies are introduced and the corresponding distribution of the data sequence $X$ is determined. The strategies concern generalized Pólya urns, random change points, covariates and stationary sequences.

stat.ME

Finitely additive mass transportation

Some classical mass transportation problems are investigated in a finitely additive setting. Let $Ω=\prod_{i=1}^nΩ_i$ and $\mathcal{A}=\otimes_{i=1}^n\mathcal{A}_i$, where $(Ω_i,\mathcal{A}_i,μ_i)$ is a ($σ$-additive) probability space for $i=1,\ldots,n$. Let $c:Ω\rightarrow [0,\infty]$ be an $\mathcal{A}$-measurable cost function. Let $M$ be the collection of finitely additive probabilities on $\mathcal{A}$ with marginals $μ_1,\ldots,μ_n$. If couplings are meant as elements of $M$, most classical results of mass transportation theory, including duality and attainability of the Kantorovich inf, are valid without any further assumptions. Special attention is devoted to martingale transport. Let $(Ω_i,\mathcal{A}_i)=(\mathbb{R},\mathcal{B}(\mathbb{R}))$ for all $i$ and $$M_1=\bigl\{P\in M:P\ll P^*\text{ and }(π_1,\ldots,π_n)\text{ is a }P\text{-martingale}\}$$ where $P^*$ is a reference probability on $\mathcal{A}$. If $M_1\ne\emptyset$, then $$\int c\,dP=\inf_{Q\in M_1}\int c\,dQ\quad\quad\text{for some }P\in M_1.$$ Conditions for $M_1\ne\emptyset$ are given as well.

math.PR

Quantitative bounds in the central limit theorem for $m$-dependent random variables

For each $n\ge 1$, let $X_{n,1},\ldots,X_{n,N_n}$ be real random variables and $S_n=\sum_{i=1}^{N_n}X_{n,i}$. Let $m_n\ge 1$ be an integer. Suppose $(X_{n,1},\ldots,X_{n,N_n})$ is $m_n$-dependent, $E(X_{ni})=0$, $E(X_{ni}^2)<\infty$ and $σ_n^2:=E(S_n^2)>0$ for all $n$ and $i$. Then, \begin{gather*} d_W\Bigl(\frac{S_n}{σ_n},\,Z\Bigr)\le 30\,\bigl\{c^{1/3}+12\,U_n(c/2)^{1/2}\bigr\}\quad\quad\text{for all }n\ge 1\text{ and }c>0, \end{gather*} where $d_W$ is Wasserstein distance, $Z$ a standard normal random variable and $$U_n(c)=\frac{m_n}{σ_n^2}\,\sum_{i=1}^{N_n}E\Bigl[X_{n,i}^2\,1\bigl\{\abs{X_{n,i}}>c\,σ_n/m_n\bigr\}\Bigr].$$ Among other things, this estimate of $d_W\bigl(S_n/σ_n,\,Z\bigr)$ yields a similar estimate of $d_{TV}\bigl(S_n/σ_n,\,Z\bigr)$ where $d_{TV}$ is total variation distance.

math.PR

New perspectives on knockoffs construction

Let $Λ$ be the collection of all probability distributions for $(X,\widetilde{X})$, where $X$ is a fixed random vector and $\widetilde{X}$ ranges over all possible knockoff copies of $X$ (in the sense of \cite{CFJL18}). Three topics are developed in this paper: (i) A new characterization of $Λ$ is proved; (ii) A certain subclass of $Λ$, defined in terms of copulas, is introduced; (iii) The (meaningful) special case where the components of $X$ are conditionally independent is treated in depth. In real problems, after observing $X=x$, each of points (i)-(ii)-(iii) may be useful to generate a value $\widetilde{x}$ for $\widetilde{X}$ conditionally on $X=x$.

math.ST

Kernel based Dirichlet sequences

Let $X=(X_1,X_2,\ldots)$ be a sequence of random variables with values in a standard space $(S,\mathcal{B})$. Suppose \begin{gather*} X_1\simν\quad\text{and}\quad P\bigl(X_{n+1}\in\cdot\mid X_1,\ldots,X_n\bigr)=\frac{θν(\cdot)+\sum_{i=1}^nK(X_i)(\cdot)}{n+θ}\quad\quad\text{a.s.} \end{gather*} where $θ>0$ is a constant, $ν$ a probability measure on $\mathcal{B}$, and $K$ a random probability measure on $\mathcal{B}$. Then, $X$ is exchangeable whenever $K$ is a regular conditional distribution for $ν$ given any sub-$σ$-field of $\mathcal{B}$. Under this assumption, $X$ enjoys all the main properties of classical Dirichlet sequences, including Sethuraman's representation, conjugacy property, and convergence in total variation of predictive distributions. If $μ$ is the weak limit of the empirical measures, conditions for $μ$ to be a.s. discrete, or a.s. non-atomic, or $μ\llν$ a.s., are provided. Two CLT's are proved as well. The first deals with stable convergence while the second concerns total variation distance.

math.PR

Bayesian predictive inference without a prior

Let $(X_n:n\ge 1)$ be a sequence of random observations. Let $σ_n(\cdot)=P\bigl(X_{n+1}\in\cdot\mid X_1,\ldots,X_n\bigr)$ be the $n$-th predictive distribution and $σ_0(\cdot)=P(X_1\in\cdot)$ the marginal distribution of $X_1$. In a Bayesian framework, to make predictions on $(X_n)$, one only needs the collection $σ=(σ_n:n\ge 0)$. Because of the Ionescu-Tulcea theorem, $σ$ can be assigned directly, without passing through the usual prior/posterior scheme. One main advantage is that no prior probability has to be selected. In this paper, $σ$ is subjected to two requirements: (i) The resulting sequence $(X_n)$ is conditionally identically distributed, in the sense of Berti, Pratelli and Rigo (2004); (ii) Each $σ_{n+1}$ is a simple recursive update of $σ_n$. Various new $σ$ satisfying (i)-(ii) are introduced and investigated. For such $σ$, the asymptotics of $σ_n$, as $n\rightarrow\infty$, is determined. In some cases, the probability distribution of $(X_n)$ is also evaluated.

math.ST

On the almost sure convergence of sums

Two counterexamples, addressing questions raised in \cite{AD} and \cite{PZ}, are provided. Both counterexamples are related to chaoses. Let $F_n=Y_n+Z_n$. It may be that $F_n\overset{a.s.}\longrightarrow 0$, $F_n\overset{L_{2+δ}}\longrightarrow 0$ and $E\bigl\{\sup_n\,\abs{F_n}^δ\bigr\}<\infty$, where $δ>0$ and $Y_n$ and $Z_n$ belong to chaoses of uniformly bounded degree, and yet $Y_n$ fails to converge to 0 a.s.

math.PR

A note on duality theorems in mass transportation

The duality theory of the Monge-Kantorovich transport problem is investigated in an abstract measure theoretic framework. Let $(\mathcal{X},\mathcal{F},μ)$ and $(\mathcal{Y},\mathcal{G},ν)$ be any probability spaces and $c:\mathcal{X}\times\mathcal{Y}\rightarrow\mathbb{R}$ a measurable cost function such that $f_1+g_1\le c\le f_2+g_2$ for some $f_1,\,f_2\in L_1(μ)$ and $g_1,\,g_2\in L_1(ν)$. Define $α(c)=\inf_P\int c\,dP$ and $α^*(c)=\sup_P\int c\,dP$, where $\inf$ and $\sup$ are over the probabilities $P$ on $\mathcal{F}\otimes\mathcal{G}$ with marginals $μ$ and $ν$. Some duality theorems for $α(c)$ and $α^*(c)$, not requiring $μ$ or $ν$ to be perfect, are proved. As an example, suppose $\mathcal{X}$ and $\mathcal{Y}$ are metric spaces and $μ$ is separable. Then, duality holds for $α(c)$ (for $α^*(c)$) provided $c$ is upper-semicontinuous (lower-semicontinuous). Moreover, duality holds for both $α(c)$ and $α^*(c)$ if the maps $x\mapsto c(x,y)$ and $y\mapsto c(x,y)$ are continuous, or if $c$ is bounded and $x\mapsto c(x,y)$ is continuous. This improves the existing results in \cite{RR1995} if $c$ satisfies the quoted conditions and the cardinalities of $\mathcal{X}$ and $\mathcal{Y}$ do not exceed the continuum.

math.PR

Centers of probability measures without the mean

In the recent years, the notion of mixability has been developed with applications to optimal transportation, quantitative finance and operations research. An $n$-tuple of distributions is said to be jointly mixable if there exist $n$ random variables following these distributions and adding up to a constant, called center, with probability one. When the $n$ distributions are identical, we speak of complete mixability. If each distribution has finite mean, the center is obviously the sum of the means. In this paper, we investigate the set of centers of completely and jointly mixable distributions not having a finite mean. In addition to several results, we show the (possibly counterintuitive) fact that, for each $n \geq 2$, there exist $n$ standard Cauchy random variables adding up to a constant $C$ if and only if $$|C|\le\frac{n\,\log (n-1)}π.$$

math.PR

Asymptotics for randomly reinforced urns with random barriers

An urn contains black and red balls. Let $Z_n$ be the proportion of black balls at time $n$ and $0\leq L L$, then $b_n$ is replaced together with a random number $R_n$ of red balls. Otherwise, no additional balls are added, and $b_n$ alone is replaced. In this paper, we assume $R_n=B_n$. Then, under mild conditions, it is shown that $Z_n\overset{a.s.}\longrightarrow Z$ for some random variable $Z$, and \begin{gather*} D_n:=\sqrt{n}\,(Z_n-Z)\longrightarrow\mathcal{N}(0,σ^2)\quad\text{conditionally a.s.} \end{gather*} where $σ^2$ is a certain random variance. Almost sure conditional convergence means that \begin{gather*} P\bigl(D_n\in\cdot\mid\mathcal{G}_n\bigr)\overset{weakly}\longrightarrow\mathcal{N}(0,\,σ^2)\quad\text{a.s.} \end{gather*} where $P\bigl(D_n\in\cdot\mid\mathcal{G}_n\bigr)$ is a regular version of the conditional distribution of $D_n$ given the past $\mathcal{G}_n$. Thus, in particular, one obtains $D_n\longrightarrow\mathcal{N}(0,σ^2)$ stably. It is also shown that $L<Z<U$ a.s. and $Z$ has non-atomic distribution.

math.PR

Central limit theorems for an Indian buffet model with random weights

The three-parameter Indian buffet process is generalized. The possibly different role played by customers is taken into account by suitable (random) weights. Various limit theorems are also proved for such generalized Indian buffet process. Let $L_n$ be the number of dishes experimented by the first $n$ customers, and let $\overline{K}_n=(1/n)\sum_{i=1}^nK_i$ where $K_i$ is the number of dishes tried by customer $i$. The asymptotic distributions of $L_n$ and $\overline{K}_n$, suitably centered and scaled, are obtained. The convergence turns out to be stable (and not only in distribution). As a particular case, the results apply to the standard (i.e., nongeneralized) Indian buffet process.

math.PR