SearcharxivSearch

arXiv subjects

Rafael Bassi Stern

Publications and source records attributed to Rafael Bassi Stern.

11 recordsLinked to original sources

Posterior Invariance of Multiplicative Contrasts under Margin Constraints in Contingency Tables

Measures of association in contingency tables, such as odds ratios and their generalizations, are often studied under different sampling schemes that either fix or leave random the margins of the table. While classical results show that certain odds ratios are unaffected by constraining the margins, it is less clear when this invariance holds more generally. This paper studies posterior inference for a broad class of multiplicative contrasts of multinomial cell probabilities, which we refer to as generalized odds ratios, and addresses exactly when fixing a margin alters inference about them. We consider Bayesian inference under multinomial sampling and under models in which partition sums of the table are fixed in advance, and assume that the marginal and conditional parameters are independent a priori. Under additional mild assumptions, we show that the posterior distribution of a generalized odds ratio is invariant to fixing a margin if and only if the coefficients defining the contrast sum to zero within the margin.

math.ST

Kullback-Leibler Consistency of $p$-dimensional P\'olya Tree Posteriors and Differential Entropy Estimation

We exploit the multiplicative structure of P\'olya Tree priors to establish novel consistency results on $p$-dimensional trees, conditions to obtain Kullback-Leibler minimax contraction rates for univariate density estimation and a representation theorem of entropy functionals of P\'olya Tree posteriors. These results motivate a novel differential entropy estimator that is consistent under mild conditions on large dimensions.

math.ST

The e-value and the Full Bayesian Significance Test: Logical Properties and Philosophical Consequences

This article gives a conceptual review of the e-value, ev(H|X) -- the epistemic value of hypothesis H given observations X. This statistical significance measure was developed in order to allow logically coherent and consistent tests of hypotheses, including sharp or precise hypotheses, via the Full Bayesian Significance Test (FBST). Arguments of analysis allow a full characterization of this statistical test by its logical or compositional properties, showing a mutual complementarity between results of mathematical statistics and the logical desiderata lying at the foundations of this theory.

math.ST

Positive Polynomials on closed boxes

We present two different proofs that positive polynomials on closed boxes of $\mathbb{R}^2$ can be written as bivariate Bernstein polynomials with strictly positive coefficients. Both strategies can be extended to prove the analogous result for polynomials that are positive on closed boxes of $\mathbb{R}^n$, $n>2$.

math.CA

Conditional independence testing: a predictive perspective

Conditional independence testing is a key problem required by many machine learning and statistics tools. In particular, it is one way of evaluating the usefulness of some features on a supervised prediction problem. We propose a novel conditional independence test in a predictive setting, and show that it achieves better power than competing approaches in several settings. Our approach consists in deriving a p-value using a permutation test where the predictive power using the unpermuted dataset is compared with the predictive power of using dataset where the feature(s) of interest are permuted. We conclude that the method achives sensible results on simulated and real datasets.

stat.ML

Interpretable hypothesis tests

Although hypothesis tests play a prominent role in Science, their interpretation can be challenging. Three issues are (i) the difficulty in making an assertive decision based on the output of an hypothesis test, (ii) the logical contradictions that occur in multiple hypothesis testing, and (iii) the possible lack of practical importance when rejecting a precise hypothesis. These issues can be addressed through the use of agnostic tests and pragmatic hypotheses.

stat.ME

Quantification under prior probability shift: the ratio estimator and its extensions

The quantification problem consists of determining the prevalence of a given label in a target population. However, one often has access to the labels in a sample from the training population but not in the target population. A common assumption in this situation is that of prior probability shift, that is, once the labels are known, the distribution of the features is the same in the training and target populations. In this paper, we derive a new lower bound for the risk of the quantification problem under the prior shift assumption. Complementing this lower bound, we present a new approximately minimax class of estimators, ratio estimators, which generalize several previous proposals in the literature. Using a weaker version of the prior shift assumption, which can be tested, we show that ratio estimators can be used to build confidence intervals for the quantification problem. We also extend the ratio estimator so that it can: (i) incorporate labels from the target population, when they are available and (ii) estimate how the prevalence of positive labels varies according to a function of certain covariates.

stat.ML

Agnostic tests can control the type I and type II errors simultaneously

Despite its common practice, statistical hypothesis testing presents challenges in interpretation. For instance, in the standard frequentist framework there is no control of the type II error. As a result, the non-rejection of the null hypothesis cannot reasonably be interpreted as its acceptance. We propose that this dilemma can be overcome by using agnostic hypothesis tests, since they can control the type I and II errors simultaneously. In order to make this idea operational, we show how to obtain agnostic hypothesis in typical models. For instance, we show how to build (unbiased) uniformly most powerful agnostic tests and how to obtain agnostic tests from standard p-values. Also, we present conditions such that the above tests can be made logically coherent. Finally, we present examples of consistent agnostic hypothesis tests.

math.ST

Learning with many experts: model selection and sparsity

Experts classifying data are often imprecise. Recently, several models have been proposed to train classifiers using the noisy labels generated by these experts. How to choose between these models? In such situations, the true labels are unavailable. Thus, one cannot perform model selection using the standard versions of methods such as empirical risk minimization and cross validation. In order to allow model selection, we present a surrogate loss and provide theoretical guarantees that assure its consistency. Next, we discuss how this loss can be used to tune a penalization which introduces sparsity in the parameters of a traditional class of models. Sparsity provides more parsimonious models and can avoid overfitting. Nevertheless, it has seldom been discussed in the context of noisy labels due to the difficulty in model selection and, therefore, in choosing tuning parameters. We apply these techniques to several sets of simulated and real data.

stat.ME

Exchangeability and the Law of Maturity

The law of maturity is the belief that less-observed events are becoming mature and, therefore, more likely to occur in the future. Previous studies have shown that the assumption of infinite exchangeability contradicts the law of maturity. In particular, it has been shown that infinite exchangeability contradicts probabilistic descriptions of the law of maturity such as the gambler's belief and the belief in maturity. We show that the weaker assumption of finite exchangeability is compatible with both the gambler's belief and belief in maturity. We provide sufficient conditions under which these beliefs hold under finite exchangeability. These conditions are illustrated with commonly used parametric models.

math.ST

Coherence of countably many bets

De Finetti's betting argument is used to justify finitely additive probabilities when only finitely many bets are considered. Under what circumstances can countably many bets be used to justify countable additivity? In this framework, one faces issues such as the convergence of the returns of the bet. Generalizations of de Finetti's argument depend on what type of conditions on convergence are required of the bets under consideration. Two new such conditions are compared with others presented in the literature.

math.PR