SearcharxivSearch

arXiv subjects

Monica Musio

Publications and source records attributed to Monica Musio.

At least 19 recordsLinked to original sources

Bayesian Discrepancy Measure: Higher-order and Skewed approximations

The aim of this paper is to discuss both higher-order asymptotic expansions and skewed approximations for the Bayesian Discrepancy Measure for testing precise statistical hypotheses. In particular, we derive results on third-order asymptotic approximations and skewed approximations for univariate posterior distributions, also in the presence of nuisance parameters, demonstrating improved accuracy in capturing posterior shape with little additional computational cost over simple first-order approximations. For the third-order approximations, connections to frequentist inference via matching priors are highlighted. Moreover, the definition of the Bayesian Discrepancy Measure and the proposed methodology are extended to the multivariate setting, employing tractable skew-normal posterior approximations obtained via derivative matching at the mode. Accurate multivariate approximations for the Bayesian Discrepancy Measure are then derived by defining credible regions based on the Optimal Transport map, that transforms the skew-normal approximation to a standard multivariate normal distribution. The performance and practical benefits of these higher-order and skewed approximations are illustrated through two examples.

stat.ME

Individual Causation with Biased Data

We consider the problem of assessing whether, in an individual case, there is a causal relationship between an observed exposure and a response variable. When data are available on similar individuals we may be able to estimate prospective probabilities, but even under ideal conditions these are typically inadequate to identify the "probability of causation": instead we can only derive bounds for this. These bounds can be improved or amended when we have information on additional variables, such as mediators or covariates. When a covariate is unobserved or ignored, this will typically lead to biased inferences. We show by examples how serious such biases can be.

math.ST

A scientometric analysis of the effect of COVID-19 on the spread of research outputs

The spread of the Sars-COV-2 pandemic in 2020 had a huge impact on the life course of all of us. This rapid spread has also caused an increase in the research production in topics related to COVID-19 with regard to different aspects. Italy has, unfortunately, been one of the first countries to be massively involved in the outbreak of the disease. In this paper we present an extensive scientometric analysis of the research production both at global (entire literature produced in the first 2 years after the beginning of the pandemic) and local level (COVID-19 literature produced by authors with an Italian affiliation). Our results showed that US and China are the most active countries in terms of number of publications and that the number of collaborations between institutions varies according to geographical distance. Moreover, we identified the medical-biological as the fields with the greatest growth in terms of literature production. Furthermore, we also better explored the relationship between the number of citations and variables obtained from the data set (e.g. number of authors per article). Using multiple correspondence analysis and quantile regression we shed light on the role of journal topics and impact factor, the type of article, the field of study and how these elements affect citations.

cs.DL

Comparison of two coefficients of variation: a new Bayesian approach

The coefficient of variation is a useful indicator for comparing the spread of values between dataset with different units or widely different means. In this paper we address the problem of investigating the equality of the coefficients of variation from two independent populations. In order to do this we rely on the Bayesian Discrepancy Measure recently introduced in the literature. Computing this Bayesian measure of evidence is straightforward when the coefficient of variation is a function of a single parameter of the distribution. In contrast, it becomes difficult when it is a function of more parameters, often requiring the use of MCMC methods. We calculate the Bayesian Discrepancy Measure by considering a variety of distributions whose coefficients of variation depend on more than one parameter. We consider also applications to real data. As far as we know, some of the examined problems have not yet been covered in the literature.

stat.ME

Robust confidence distributions from proper scoring rules

A confidence distribution is a distribution for a parameter of interest based on a parametric statistical model. As such, it serves the same purpose for frequentist statisticians as a posterior distribution for Bayesians, since it allows to reach point estimates, to assess their precision, to set up tests along with measures of evidence, to derive confidence intervals, comparing the parameter of interest with other parameters from other studies, etc. A general recipe for deriving confidence distributions is based on classical pivotal quantities and their exact or approximate distributions. However, in the presence of model misspecifications or outlying values in the observed data, classical pivotal quantities, and thus confidence distributions, may be inaccurate. The aim of this paper is to discuss the derivation and application of robust confidence distributions. In particular, we discuss a general approach based on the Tsallis scoring rule in order to compute a robust confidence distribution. Examples and simulation results are discussed for some problems often encountered in practice, such as the two-sample heteroschedastic comparison, the receiver operating characteristic curves and regression models.

stat.ME

A new Bayesian discrepancy measure

The aim of this article is to make a contribution to the Bayesian procedure of testing precise hypotheses for parametric models. For this purpose, we define the Bayesian Discrepancy Measure that allows one to evaluate the suitability of a given hypothesis with respect to the available information (prior law and data). To summarise this information, the posterior median is employed, allowing a simple assessment of the discrepancy with a fixed hypothesis. The Bayesian Discrepancy Measure assesses the compatibility of a single hypothesis with the observed data, as opposed to the more common comparative approach where a hypothesis is rejected in favour of a competing hypothesis. The proposed measure of evidence has properties of consistency and invariance. After presenting the definition of the measure for a parameter of interest, both in the absence and in the presence of nuisance parameters, we illustrate some examples showing its conceptual and interpretative simplicity. Finally, we compare the BDT with the Full Bayesian Significance Test, a well-known Bayesian testing procedure for sharp hypotheses.

stat.ME

Interpreting the outcomes of research assessments: a geometrical approach

Research evaluations and comparison of the assessments of academic institutions (scientific areas, departments, universities etc.) are among the major issues in recent years in higher education systems. One method, followed by some national evaluation agencies, is to assess the research quality by the evaluation of a limited number of publications in a way that each publication is rated among $n$ classes. This method produces, for each institution, a distribution of the publications in the $n$ classes. In this paper we introduce a natural geometric way to compare these assessments by introducing an ad hoc distance from the performance of an institution to the best possible achievable assessment. Moreover, to avoid the methodological error of comparing non-homogeneous institutions, we introduce a {\em geometric score} based on such a distance. The latter represents the probability that an ideal institution, with the same configuration as the one under evaluation, performs worst. We apply our method, based on the geometric score, to rank, in two specific scientific areas, the Italian universities using the results of the evaluation exercise VQR 2011-2014.

cs.DL

Effects of Causes and Causes of Effects

We describe and contrast two distinct problem areas for statistical causality: studying the likely effects of an intervention ("effects of causes"), and studying whether there is a causal link between the observed exposure and outcome in an individual case ("causes of effects"). For each of these, we introduce and compare various formal frameworks that have been proposed for that purpose, including the decision-theoretic approach, structural equations, structural and stochastic causal models, and potential outcomes. It is argued that counterfactual concepts are unnecessary for studying effects of causes, but are needed for analysing causes of effects. They are however subject to a degree of arbitrariness, which can be reduced, though not in general eliminated, by taking account of additional structure in the problem.

math.ST

Bounding Causes of Effects with Mediators

Suppose X and Y are binary exposure and outcome variables, and we have full knowledge of the distribution of Y, given application of X. From this we know the average causal effect of X on Y. We are now interested in assessing, for a case that was exposed and exhibited a positive outcome, whether it was the exposure that caused the outcome. The relevant "probability of causation", PC, typically is not identified by the distribution of Y given X, but bounds can be placed on it, and these bounds can be improved if we have further information about the causal process. Here we consider cases where we know the probabilistic structure for a sequence of complete mediators between X and Y. We derive a general formula for calculating bounds on PC for any pattern of data on the mediators (including the case with no data). We show that the largest and smallest upper and lower bounds that can result from any complete mediation process can be obtained in processes with at most two steps. We also consider homogeneous processes with many mediators. PC can sometimes be identified as 0 with negative data, but it cannot be identified at 1 even with positive data on an infinite set of mediators. The results have implications for learning about causation from knowledge of general processes and of data on cases.

math.ST

The Hyv\"arinen scoring rule in Gaussian linear time series models

Likelihood-based estimation methods involve the normalising constant of the model distributions, expressed as a function of the parameter. However in many problems this function is not easily available, and then less efficient but more easily computed estimators may be attractive. In this work we study stationary time-series models, and construct and analyse "score-matching'' estimators, that do not involve the normalising constant. We consider two scenarios: a single series of increasing length, and an increasing number of independent series of fixed length. In the latter case there are two variants, one based on the full data, and another based on a sufficient statistic. We study the empirical performance of these estimators in three special cases, autoregressive (\AR), moving average (MA) and fractionally differenced white noise (\ARFIMA) models, and make comparisons with full and pairwise likelihood estimators. The results are somewhat model-dependent, with the new estimators doing well for $\MA$ and \ARFIMA\ models, but less so for $\AR$ models.

stat.ME

Causes of Effects via a Bayesian Model Selection Procedure

In causal inference, and specifically in the \textit{Causes of Effects} problem, one is interested in how to use statistical evidence to understand causation in an individual case, and so how to assess the so-called {\em probability of causation} (PC). The answer relies on the potential responses, which can incorporate information about what would have happened to the outcome as we had observed a different value of the exposure. However, even given the best possible statistical evidence for the association between exposure and outcome, we can typically only provide bounds for the PC. Dawid et al. (2016) highlighted some fundamental conditions, namely, exogeneity, comparability, and sufficiency, required to obtain such bounds, based on experimental data. The aim of the present paper is to provide methods to find, in specific cases, the best subsample of the reference dataset to satisfy such requirements. To this end, we introduce a new variable, expressing the desire to be exposed or not, and we set the question up as a model selection problem. The best model will be selected using the marginal probability of the responses and a suitable prior proposal over the model space. An application in the educational field is presented.

stat.ME

The Probability of Causation

Many legal cases require decisions about causality, responsibility or blame, and these may be based on statistical data. However, causal inferences from such data are beset by subtle conceptual and practical difficulties, and in general it is, at best, possible to identify the "probability of causation" as lying between certain empirically informed limits. These limits can be refined and improved if we can obtain additional information, from statistical or scientific data, relating to the internal workings of the causal processes. In this paper we review and extend recent work in this area, where additional information may be available on covariate and/or mediating variables.

math.ST

New bounds for the Probability of Causation in Mediation Analysis

An individual has been subjected to some exposure and has developed some outcome. Using data on similar individuals, we wish to evaluate, for this case, the probability that the outcome was in fact caused by the exposure. Even with the best possible experimental data on exposure and outcome, we typically can not identify this "probability of causation" exactly, but we can provide information in the form of bounds for it. Under appropriate assumptions, these bounds can be tightened if we can make other observations (e.g., on non-experimental cases), measure additional variables (e.g., covariates) or measure complete mediators. In this work we propose new bounds for the case that a third variable mediates partially the effect of the exposure on the outcome.

math.ST

A Note on Bayesian Model Selection for Discrete Data Using Proper Scoring Rules

We consider the problem of choosing between parametric models for a discrete observable, taking a Bayesian approach in which the within-model prior distributions are allowed to be improper. In order to avoid the ambiguity in the marginal likelihood function in such a case, we apply a homogeneous scoring rule. For the particular case of distinguishing between Poisson and Negative Binomial models, we conduct simulations that indicate that, applied prequentially, the method will consistently select the true model.

math.ST

Rejoinder to "Bayesian Model Selection Based on Proper Scoring Rules"

We are deeply appreciative of the initiative of the editor, Marina Vanucci, in commissioning a discussion of our paper, and extremely grateful to all the discussants for their insightful and thought-provoking comments. We respond to the discussions in alphabetical order [arXiv:1409.5291].

math.ST

Bayesian Model Selection Based on Proper Scoring Rules

Bayesian model selection with improper priors is not well-defined because of the dependence of the marginal likelihood on the arbitrary scaling constants of the within-model prior densities. We show how this problem can be evaded by replacing marginal log-likelihood by a homogeneous proper scoring rule, which is insensitive to the scaling constants. Suitably applied, this will typically enable consistent selection of the true model.

math.ST

Comparisons of Hyvärinen and pairwise estimators in two simple linear time series models

The aim of this paper is to compare numerically the performance of two estimators based on Hyvärinen's local homogeneous scoring rule with that of the full and the pairwise maximum likelihood estimators. In particular, two different model settings, for which both full and pairwise maximum likelihood estimators can be obtained, have been considered: the first order autoregressive model (AR(1)) and the moving average model (MA(1)). Simulation studies highlight very different behaviours for the Hyvärinen scoring rule estimators relative to the pairwise likelihood estimators in these two settings.

stat.ME