SearcharxivSearch

arXiv subjects

Emanuela Dreassi

Publications and source records attributed to Emanuela Dreassi.

11 recordsLinked to original sources

Some cautionary tales about Bayesian predictive inference

Two misunderstandings, frequently arising in Bayesian predictive inference, are discussed. The first deals with the data generating mechanism, while the second consists in overestimating the role played by asymptotic exchangeability. Some consequences of such misunderstandings are highlighted through examples.

math.ST

Variable selection via knockoffs in missing data settings with categorical predictors

Large-scale assessment data typically include numerous categorical variables, often affected by missing values. Motivated by the challenges arising in this framework, we extend the knockoffs method for selecting predictors to settings with missing values. Our proposal relies on a preliminary phase consisting of multiple imputations of missing values. Each imputed dataset is then processed using a suitable knockoff filter. We evaluate the performance of the proposed method through a simulation study, showing satisfactory results consistent with a recently advocated cutting-edge method. We apply the method to large-scale assessment data collected by INVALSI about test scores of Italian students in grade 5 with many background variables. This case study is challenging, as most predictors have unordered categories, a setting not taken into account by traditional knockoffs methods. In addition, some of the key predictors are affected by missing values. The model includes random effects to account for the multilevel structure of students nested into schools. Our proposal to implement the knockoffs method within a multiple imputation framework proves to be feasible, flexible and effective.

stat.ME

Bayesian nonparametric inference on a Fr\'echet class

Let $(\mathcal{X},\mathcal{F},\mu)$ and $(\mathcal{Y},\mathcal{G},\nu)$ be probability spaces and $(Z_n)$ a sequence of random variables with values in $(\mathcal{X}\times\mathcal{Y},\,\mathcal{F}\otimes\mathcal{G})$. Let $\Gamma(\mu,\nu)$ be the collection of all probability measures $p$ on $\mathcal{F}\otimes\mathcal{G}$ such that $$p\bigl(A\times\mathcal{Y}\bigr)=\mu(A)\quad\text{and}\quad p\bigl(\mathcal{X}\times B\bigr)=\nu(B)\quad\text{for all }A\in\mathcal{F}\text{ and }B\in\mathcal{G}.$$ In this paper, we build some probability measures $\Pi$ on $\Gamma(\mu,\nu)$. In addition, for each such $\Pi$, we assume that $(Z_n)$ is exchangeable with de Finetti's measure $\Pi$ and we evaluate the conditional distribution $\Pi(\cdot\mid Z_1,\ldots,Z_n)$. In Bayesian nonparametrics, if $(Z_1,\ldots, Z_n)$ are the available data, $\Pi$ and $\Pi(\cdot\mid Z_1,\ldots, Z_n)$ can be regarded as the prior and the posterior, respectively. To support this interpretation, it suffices to think of a problem where the unknown probability distribution of some bivariate phenomenon is constrained to have marginals $\mu$ and $\nu$. Finally, analogous results are obtained for the set $\Gamma(\mu)$ of those probability measures on $\mathcal{F}\otimes\mathcal{G}$ with marginal $\mu$ on $\mathcal{F}$ (but arbitrary marginal on $\mathcal{G}$). That is, we introduce some priors on $\Gamma(\mu)$ and we evaluate the corresponding posteriors.

stat.ME

Knockoffs for exchangeable categorical covariates

Let $X=(X_1,\ldots,X_p)$ be a $p$-variate random vector and $F$ a fixed finite set. In a number of applications, mainly in genetics, it turns out that $X_i\in F$ for each $i=1,\ldots,p$. Despite the latter fact, to obtain a knockoff $\widetilde{X}$ (in the sense of \cite{CFJL18}), $X$ is usually modeled as an absolutely continuous random vector. While comprehensible from the point of view of applications, this approximate procedure does not make sense theoretically, since $X$ is supported by the finite set $F^p$. In this paper, explicit formulae for the joint distribution of $(X,\widetilde{X})$ are provided when $P(X\in F^p)=1$ and $X$ is exchangeable or partially exchangeable. In fact, when $X_i\in F$ for all $i$, there seem to be various reasons for assuming $X$ exchangeable or partially exchangeable. The robustness of $\widetilde{X}$, with respect to the de Finetti's measure $\pi$ of $X$, is investigated as well. Let $\mathcal{L}_\pi(\widetilde{X}\mid X=x)$ denote the conditional distribution of $\widetilde{X}$, given $X=x$, when the de Finetti's measure is $\pi$. It is shown that $$\norm{\mathcal{L}_{\pi_1}(\widetilde{X}\mid X=x)-\mathcal{L}_{\pi_2}(\widetilde{X}\mid X=x)}\le c(x)\,\norm{\pi_1-\pi_2}$$ where $\norm{\cdot}$ is total variation distance and $c(x)$ a suitable constant. Finally, a numerical experiment is performed. Overall, the knockoffs of this paper outperform the alternatives (i.e., the knockoffs obtained by giving $X$ an absolutely continuous distribution) as regards the false discovery rate but are slightly weaker in terms of power.

math.ST

Generating knockoffs via conditional independence

Let $X$ be a $p$-variate random vector and $\widetilde{X}$ a knockoff copy of $X$ (in the sense of \cite{CFJL18}). A new approach for constructing $\widetilde{X}$ (henceforth, NA) has been introduced in \cite{JSPI}. NA has essentially three advantages: (i) To build $\widetilde{X}$ is straightforward; (ii) The joint distribution of $(X,\widetilde{X})$ can be written in closed form; (iii) $\widetilde{X}$ is often optimal under various criteria. However, for NA to apply, $X_1,\ldots, X_p$ should be conditionally independent given some random element $Z$. Our first result is that any probability measure $μ$ on $\mathbb{R}^p$ can be approximated by a probability measure $μ_0$ of the form $$μ_0\bigl(A_1\times\ldots\times A_p\bigr)=E\Bigl\{\prod_{i=1}^p P(X_i\in A_i\mid Z)\Bigr\}.$$ The approximation is in total variation distance when $μ$ is absolutely continuous, and an explicit formula for $μ_0$ is provided. If $X\simμ_0$, then $X_1,\ldots,X_p$ are conditionally independent. Hence, with a negligible error, one can assume $X\simμ_0$ and build $\widetilde{X}$ through NA. Our second result is a characterization of the knockoffs $\widetilde{X}$ obtained via NA. It is shown that $\widetilde{X}$ is of this type if and only if the pair $(X,\widetilde{X})$ can be extended to an infinite sequence so as to satisfy certain invariance conditions. The basic tool for proving this fact is de Finetti's theorem for partially exchangeable sequences. In addition to the quoted results, an explicit formula for the conditional distribution of $\widetilde{X}$ given $X$ is obtained in a few cases. In one of such cases, it is assumed $X_i\in\{0,1\}$ for all $i$.

math.ST

A probabilistic view on predictive constructions for Bayesian learning

Given a sequence $X=(X_1,X_2,\ldots)$ of random observations, a Bayesian forecaster aims to predict $X_{n+1}$ based on $(X_1,\ldots,X_n)$ for each $n\ge 0$. To this end, in principle, she only needs to select a collection $σ=(σ_0,σ_1,\ldots)$, called ``strategy" in what follows, where $σ_0(\cdot)=P(X_1\in\cdot)$ is the marginal distribution of $X_1$ and $σ_n(\cdot)=P(X_{n+1}\in\cdot\mid X_1,\ldots,X_n)$ the $n$-th predictive distribution. Because of the Ionescu-Tulcea theorem, $σ$ can be assigned directly, without passing through the usual prior/posterior scheme. One main advantage is that no prior probability is to be selected. In a nutshell, this is the predictive approach to Bayesian learning. A concise review of the latter is provided in this paper. We try to put such an approach in the right framework, to make clear a few misunderstandings, and to provide a unifying view. Some recent results are discussed as well. In addition, some new strategies are introduced and the corresponding distribution of the data sequence $X$ is determined. The strategies concern generalized Pólya urns, random change points, covariates and stationary sequences.

stat.ME

New perspectives on knockoffs construction

Let $Λ$ be the collection of all probability distributions for $(X,\widetilde{X})$, where $X$ is a fixed random vector and $\widetilde{X}$ ranges over all possible knockoff copies of $X$ (in the sense of \cite{CFJL18}). Three topics are developed in this paper: (i) A new characterization of $Λ$ is proved; (ii) A certain subclass of $Λ$, defined in terms of copulas, is introduced; (iii) The (meaningful) special case where the components of $X$ are conditionally independent is treated in depth. In real problems, after observing $X=x$, each of points (i)-(ii)-(iii) may be useful to generate a value $\widetilde{x}$ for $\widetilde{X}$ conditionally on $X=x$.

math.ST

Kernel based Dirichlet sequences

Let $X=(X_1,X_2,\ldots)$ be a sequence of random variables with values in a standard space $(S,\mathcal{B})$. Suppose \begin{gather*} X_1\simν\quad\text{and}\quad P\bigl(X_{n+1}\in\cdot\mid X_1,\ldots,X_n\bigr)=\frac{θν(\cdot)+\sum_{i=1}^nK(X_i)(\cdot)}{n+θ}\quad\quad\text{a.s.} \end{gather*} where $θ>0$ is a constant, $ν$ a probability measure on $\mathcal{B}$, and $K$ a random probability measure on $\mathcal{B}$. Then, $X$ is exchangeable whenever $K$ is a regular conditional distribution for $ν$ given any sub-$σ$-field of $\mathcal{B}$. Under this assumption, $X$ enjoys all the main properties of classical Dirichlet sequences, including Sethuraman's representation, conjugacy property, and convergence in total variation of predictive distributions. If $μ$ is the weak limit of the empirical measures, conditions for $μ$ to be a.s. discrete, or a.s. non-atomic, or $μ\llν$ a.s., are provided. Two CLT's are proved as well. The first deals with stable convergence while the second concerns total variation distance.

math.PR

Bayesian predictive inference without a prior

Let $(X_n:n\ge 1)$ be a sequence of random observations. Let $σ_n(\cdot)=P\bigl(X_{n+1}\in\cdot\mid X_1,\ldots,X_n\bigr)$ be the $n$-th predictive distribution and $σ_0(\cdot)=P(X_1\in\cdot)$ the marginal distribution of $X_1$. In a Bayesian framework, to make predictions on $(X_n)$, one only needs the collection $σ=(σ_n:n\ge 0)$. Because of the Ionescu-Tulcea theorem, $σ$ can be assigned directly, without passing through the usual prior/posterior scheme. One main advantage is that no prior probability has to be selected. In this paper, $σ$ is subjected to two requirements: (i) The resulting sequence $(X_n)$ is conditionally identically distributed, in the sense of Berti, Pratelli and Rigo (2004); (ii) Each $σ_{n+1}$ is a simple recursive update of $σ_n$. Various new $σ$ satisfying (i)-(ii) are introduced and investigated. For such $σ$, the asymptotics of $σ_n$, as $n\rightarrow\infty$, is determined. In some cases, the probability distribution of $(X_n)$ is also evaluated.

math.ST

A Bayesian semiparametric model for semicontinuous data

When the target variable exhibits a semicontinuous behaviour (i.e. a point mass in a single value and a continuous distribution elsewhere) parametric `two-part regression models' have been extensively used and investigated. In this paper, a semiparametric Bayesian two-part regression model for dealing with such variables is proposed. The model allows a semiparametric expression for the two part of the model by using Dirichlet processes. A motivating example (in the `small area estimation' framework) based on pseudo-real data on grapewine production in Tuscany, is used to evaluate the capabilities of the model. Results show a satisfactory performance of the suggested approach to model and predict semicontinuous data when parametric assumptions (distributional and/or relationship) are not reasonable.

stat.ME

Disease Mapping via Negative Binomial Regression M-quantiles

We introduce a semi-parametric approach to ecological regression for disease mapping, based on modelling the regression M-quantiles of a Negative Binomial variable. The proposed method is robust to outliers in the model covariates, including those due to measurement error, and can account for both spatial heterogeneity and spatial clustering. A simulation experiment based on the well-known Scottish lip cancer data set is used to compare the M-quantile modelling approach and a random effects modelling approach for disease mapping. This suggests that the M-quantile approach leads to predicted relative risks with smaller root mean square error than standard disease mapping methods. The paper concludes with an illustrative application of the M-quantile approach, mapping low birth weight incidence data for English Local Authority Districts for the years 2005-2010.

stat.ME