SearcharxivSearch

arXiv subjects

Jeff Dominitz

Publications and source records attributed to Jeff Dominitz.

5 recordsLinked to original sources

Regret in Treatment Choice when Welfare Varies with an Uncertain Event: The Prediction-Threshold Problem

We study maximum regret (MR) of binary treatment choice in a population with observed covariates x, when welfare varies with an uncertain binary event. We consider decision making with plug-in probabilistic predictions of the event and pre-specified decision thresholds, which we term the prediction-threshold problem. The optimal treatment for persons with covariate value x is B if the conditional probability P(y=1|x) of a binary outcome y exceeds a particular x-specific threshold and is A otherwise. This structure is common in medical decision making and other contexts. Plug-in prediction uses data to estimate P(y|x) and acts as if the estimate is accurate. However, plug-in prediction is often performed with misspecified prediction models and conventional x-invariant thresholds. We use a combination of algebraic and computational analysis of limit and finite-sample MR to demonstrate how MR depends on the prediction model, the state space, and the thresholds used to choose treatments.

econ.EM

A Decision Theoretic Perspective on Artificial Superintelligence: Coping with Missing Data Problems in Prediction and Treatment Choice

Enormous attention and resources are being devoted to the quest for artificial general intelligence and, even more ambitiously, artificial superintelligence. We wonder about the implications for methodological research that aims to help decision makers cope with what econometricians call identification problems, inferential problems in empirical research that do not diminish as sample size grows. Of particular concern are missing data problems in prediction and treatment choice. Essentially all data collection intended to inform decision making is subject to missing data, which gives rise to identification problems. Thus far, we see no indication that the current dominant architecture of machine learning (ML)-based artificial intelligence (AI) systems will outperform humans in this context. In this paper, we explain why we have reached this conclusion and why we see the missing data problem as a cautionary case study in the quest for superintelligence more generally. We first discuss the concept of intelligence, focusing initially on some work by AI researchers, before presenting a decision-theoretic perspective that formalizes the connection between intelligence and identification problems. We next apply this perspective to two leading cases of missing data problems. Then we explain why we are skeptical that AI research is currently on a path toward machines doing better than humans at solving these identification problems.

econ.EM

Partial Identification of Mean Achievement in ILSA Studies with Multi-Stage Stratified Sample Design and Student Non-Participation

International large-scale assessment (ILSA) studies collect information across education systems with the objective of learning about the population-wide distribution of student achievement in the assessment. In this article, we study one of the most fundamental threats that these studies face when justifying the conclusions reached about these distributions: the identification problem that arises from student non-participation during data collection. Recognizing that ILSA studies have traditionally employed a narrow range of strategies to address non-participation, we examine this problem using tools developed within the framework of partial identification of probability distributions. We tailor this framework to the problem of non-participation when data are collected using a multi-stage stratified random sample design, as in most ILSA studies. We demonstrate this approach with application to the International Computer and Information Literacy Study in 2018. We show how to use the framework to assess mean achievement under reasonable and credible sets of assumptions about the non-participating population. We also provide examples of how these results may be reported by agencies that administer ILSA studies. By doing so, we bring to the field of ILSA an alternative strategy for identification, estimation, and reporting of population parameters of interest.

econ.EM

Using Total Margin of Error to Account for Non-Sampling Error in Election Polls: The Case of Nonresponse

The potential impact of non-sampling errors on election polls is well known, but measurement has focused on the margin of sampling error. Survey statisticians have long recommended measurement of total survey error by mean square error (MSE), which jointly measures sampling and non-sampling errors. We think it reasonable to use the square root of maximum MSE to measure the total margin of error (TME). Measurement of TME should encompass both sampling error and all forms of non-sampling error. We suggest that measurement of TME should be a standard feature in the reporting of polls. To provide a clear illustration, and because we believe the exceedingly low response rates commonly obtained by election polls to be a particularly worrisome source of potential error, we demonstrate how to measure the potential impact of nonresponse using the concept of TME. We first show how to measure TME when a pollster lacks any knowledge of the candidate preferences of nonrespondents. We then extend the analysis to settings where the pollster has partial knowledge that bounds the preferences of non-respondents. In each setting, we derive a simple poll estimate that approximately minimizes TME, a midpoint estimate, and compare it to a conventional poll estimate.

econ.EM

Comprehensive OOS Evaluation of Predictive Algorithms with Statistical Decision Theory

We argue that comprehensive out-of-sample (OOS) evaluation using statistical decision theory (SDT) should replace the current practice of K-fold and Common Task Framework validation in machine learning (ML) research on prediction. SDT provides a formal frequentist framework for performing comprehensive OOS evaluation across all possible (1) training samples, (2) populations that may generate training data, and (3) populations of prediction interest. Regarding feature (3), we emphasize that SDT requires the practitioner to directly confront the possibility that the future may not look like the past and to account for a possible need to extrapolate from one population to another when building a predictive algorithm. For specificity, we consider treatment choice using conditional predictions with alternative restrictions on the state space of possible populations that may generate training data. We discuss application of SDT to the problem of predicting patient illness to inform clinical decision making. SDT is simple in abstraction, but it is often computationally demanding to implement. We call on ML researchers, econometricians, and statisticians to expand the domain within which implementation of SDT is tractable.

econ.EM