Searcharxiv⌕ Search

arXiv subjects

A. Philip Dawid

Publications and source records attributed to A. Philip Dawid.

At least 19 recordsLinked to original sources

Coherent Measures of Discrepancy, Uncertainty and Dependence, with Applications to Bayesian Predictive Experimental Design

We show how, associated with any decision problem, we may derive related functions measuring uncertainty, discrepancy and dependence of distributions. Such "coherent" functions have special properties, which we characterise, and each function essentially determines the others. The theory is applied to the Bayesian formulation of the problem of choosing an experiment in order to make a subsequent prediction. It is shown that coherent choice criteria may be based on any of the coherent functions, with related functions yielding identical solutions.

math.ST↗

Personalised Decision-Making without Counterfactuals

This article is a response to recent proposals by Pearl and others for a new approach to personalised treatment decisions, in contrast to the traditional one based on statistical decision theory. We argue that this approach is dangerously misguided and should not be used in practice.

stat.ME↗

A comparison of graphical methods in the case of the murder of Meredith Kercher

We compare three graphical methods for displaying evidence in a legal case: Wigmore Charts, Bayesian Networks and Chain Event Graphs. We find that these methods are aimed at three distinct audiences, respectively lawyers, forensic scientists and the police. The methods are illustrated using part of the evidence in the case of the murder of Meredith Kercher. More specifically, we focus on representing the list of propositions, evidence, testimony and facts given in the first trial against Raffaele Sollecito and Amanda Knox with these graphical methodologies.

stat.AP↗

On Learnability under General Stochastic Processes

Statistical learning theory under independent and identically distributed (iid) sampling and online learning theory for worst case individual sequences are two of the best developed branches of learning theory. Statistical learning under general non-iid stochastic processes is less mature. We provide two natural notions of learnability of a function class under a general stochastic process. We show that both notions are in fact equivalent to online learnability. Our results hold for both binary classification and regression.

stat.ML↗

Effects of Causes and Causes of Effects

We describe and contrast two distinct problem areas for statistical causality: studying the likely effects of an intervention ("effects of causes"), and studying whether there is a causal link between the observed exposure and outcome in an individual case ("causes of effects"). For each of these, we introduce and compare various formal frameworks that have been proposed for that purpose, including the decision-theoretic approach, structural equations, structural and stochastic causal models, and potential outcomes. It is argued that counterfactual concepts are unnecessary for studying effects of causes, but are needed for analysing causes of effects. They are however subject to a degree of arbitrariness, which can be reduced, though not in general eliminated, by taking account of additional structure in the problem.

math.ST↗

Decision-theoretic foundations for statistical causality

We develop a mathematical and interpretative foundation for the enterprise of decision-theoretic statistical causality (DT), which is a straightforward way of representing and addressing causal questions. DT reframes causal inference as "assisted decision-making", and aims to understand when, and how, I can make use of external data, typically observational, to help me solve a decision problem by taking advantage of assumed relationships between the data and my problem. The relationships embodied in any representation of a causal problem require deeper justification, which is necessarily context-dependent. Here we clarify the considerations needed to support applications of the DT methodology. Exchangeability considerations are used to structure the required relationships, and a distinction drawn between intention to treat and intervention to treat forms the basis for the enabling condition of "ignorability". We also show how the DT perspective unifies and sheds light on other popular formalisations of statistical causality, including potential responses and directed acyclic graphs.

math.ST↗

The Hyvärinen scoring rule in Gaussian linear time series models

Likelihood-based estimation methods involve the normalising constant of the model distributions, expressed as a function of the parameter. However in many problems this function is not easily available, and then less efficient but more easily computed estimators may be attractive. In this work we study stationary time-series models, and construct and analyse "score-matching'' estimators, that do not involve the normalising constant. We consider two scenarios: a single series of increasing length, and an increasing number of independent series of fixed length. In the latter case there are two variants, one based on the full data, and another based on a sufficient statistic. We study the empirical performance of these estimators in three special cases, autoregressive (\AR), moving average (MA) and fractionally differenced white noise (\ARFIMA) models, and make comparisons with full and pairwise likelihood estimators. The results are somewhat model-dependent, with the new estimators doing well for $\MA$ and \ARFIMA\ models, but less so for $\AR$ models.

stat.ME↗

A Note on Bayesian Model Selection for Discrete Data Using Proper Scoring Rules

We consider the problem of choosing between parametric models for a discrete observable, taking a Bayesian approach in which the within-model prior distributions are allowed to be improper. In order to avoid the ambiguity in the marginal likelihood function in such a case, we apply a homogeneous scoring rule. For the particular case of distinguishing between Poisson and Negative Binomial models, we conduct simulations that indicate that, applied prequentially, the method will consistently select the true model.

math.ST↗

A Note on Prediction Markets

In a prediction market, individuals can sequentially place bets on the outcome of a future event. This leaves a trail of personal probabilities for the event, each being conditional on the current individual's private background knowledge and on the previously announced probabilities of other individuals, which give partial information about their private knowledge. By means of theory and examples, we revisit some results in this area. In particular, we consider the case of two individuals, who start with the same overall probability distribution but different private information, and then take turns in updating their probabilities. We note convergence of the announced probabilities to a limiting value, which may or may not be the same as that based on pooling their private information.

math.ST↗

Extended Conditional Independence and Applications in Causal Inference

The goal of this paper is to integrate the notions of stochastic conditional independence and variation conditional independence under a more general notion of extended conditional independence. We show that under appropriate assumptions the calculus that applies for the two cases separately (axioms of a separoid) still applies for the extended case. These results provide a rigorous basis for a wide range of statistical concepts, including ancillarity and sufficiency, and, in particular, the Decision Theoretic framework for statistical causality, which uses the language and calculus of conditional independence in order to express causal properties and make causal inferences.

math.ST↗

Structural Markov graph laws for Bayesian model uncertainty

This paper considers the problem of defining distributions over graphical structures. We propose an extension of the hyper Markov properties of Dawid and Lauritzen [Ann. Statist. 21 (1993) 1272-1317], which we term structural Markov properties, for both undirected decomposable and directed acyclic graphs, which requires that the structure of distinct components of the graph be conditionally independent given the existence of a separating component. This allows the analysis and comparison of multiple graphical structures, while being able to take advantage of the common conditional independence constraints. Moreover, we show that these properties characterise exponential families, which form conjugate priors under sampling from compatible Markov distributions.

math.ST↗

Rejoinder to "Bayesian Model Selection Based on Proper Scoring Rules"

We are deeply appreciative of the initiative of the editor, Marina Vanucci, in commissioning a discussion of our paper, and extremely grateful to all the discussants for their insightful and thought-provoking comments. We respond to the discussions in alphabetical order [arXiv:1409.5291].

math.ST↗

Bayesian Model Selection Based on Proper Scoring Rules

Bayesian model selection with improper priors is not well-defined because of the dependence of the marginal likelihood on the arbitrary scaling constants of the within-model prior densities. We show how this problem can be evaded by replacing marginal log-likelihood by a homogeneous proper scoring rule, which is insensitive to the scaling constants. Suitably applied, this will typically enable consistent selection of the true model.

math.ST↗

A Commentary on Statistical Assessment of Violence Recidivism Risk

Increasing integration and availability of data on large groups of persons has been accompanied by proliferation of statistical and other algorithmic prediction tools in banking, insurance, marketiNg, medicine, and other FIelds (see e.g., Steyerberg (2009a;b)). Controversy may ensue when such tools are introduced to fields traditionally reliant on individual clinical evaluations. Such controversy has arisen about "actuarial" assessments of violence recidivism risk, i.e., the probability that someone found to have committed a violent act will commit another during a specified period. Recently Hart et al. (2007a) and subsequent papers from these authors in several reputable journals have claimed to demonstrate that statistical assessments of such risks are inherently too imprecise to be useful, using arguments that would seem to apply to statistical risk prediction quite broadly. This commentary examines these arguments from a technical statistical perspective, and finds them seriously mistaken in many particulars. They should play no role in reasoned discussions of violence recidivism risk assessment.

stat.ME↗

Stochastic Mechanistic Interaction

We propose a fully probabilistic formulation of the notion of mechanistic interaction (interaction in some fundamental mechanistic sense) between the effects of putative (possibly continuous) causal factors A and B on a binary outcome variable Y indicating 'survival' vs 'failure'. We define mechanistic interaction in terms of departure from a generalized 'noisy OR' model, under which the multiplicative causal effect of A (resp., B) on the probability of failure cannot be enhanced by manipulating B (resp., A). We present conditions under which mechanistic interaction in the above sense can be assessed via simple tests on excess risk or superadditivity, in a possibly retrospective regime of observation. These conditions are defined in terms of generalized conditional independence relationships (generalised because they may involve non-stochastic 'regime indicators') that can often be checked on a graphical representation of the problem. Inference about mechanistic interaction between direct, or path-specific, causal effects can be accommodated in the proposed framework. The method is illustrated with the aid of a study in experimental psychology.

stat.ME↗

Comparisons of Hyvärinen and pairwise estimators in two simple linear time series models

The aim of this paper is to compare numerically the performance of two estimators based on Hyvärinen's local homogeneous scoring rule with that of the full and the pairwise maximum likelihood estimators. In particular, two different model settings, for which both full and pairwise maximum likelihood estimators can be obtained, have been considered: the first order autoregressive model (AR(1)) and the moving average model (MA(1)). Simulation studies highlight very different behaviours for the Hyvärinen scoring rule estimators relative to the pairwise likelihood estimators in these two settings.

stat.ME↗

On Individual Risk

We survey a variety of possible explications of the term "Individual Risk." These in turn are based on a variety of interpretations of "Probability," including Classical, Enumerative, Frequency, Formal, Metaphysical, Personal, Propensity, Chance and Logical conceptions of Probability, which we review and compare. We distinguish between "groupist" and "individualist" understandings of Probability, and explore both "group to individual" (G2i) and "individual to group" (i2G) approaches to characterising Individual Risk. Although in the end that concept remains subtle and elusive, some pragmatic suggestions for progress are made.

stat.AP↗