SearcharxivSearch

arXiv subjects

Nick Huntington-Klein

Publications and source records attributed to Nick Huntington-Klein.

4 recordsLinked to original sources

Partial Identification of Causal Effects that Vary by Setting

The estimation of causal effects using quasiexperiments often relies on the use of unusual or serendipitous sources of exogenous variation. When the goal is estimating the same causal effects across many different settings, the same unusual exogenous variation often does not exist in all settings, and the only available form of identification is selection-on-observables, which relies on a conditional indepdendence assumption. Partial identification is especially valuable in this context, as it allows conditional independence to not hold perfectly. This paper proposes a method that sharpens the jointly identified set of causal effects across many settings by making use of unobserved relationships between omitted variable biases across settings.

econ.EM

Little Impact of ChatGPT Availability on High School Student Test Score Performance

In educational settings, AI can be used as a learning aid, but can also be used to avoid schoolwork, thereby passing classes while learning little. Many existing studies on the impact of AI on education focus on AI use in controlled settings or with specialized tools. In this paper, the dropoff in ChatGPT activity during non-school summer months in 2023 and 2024 is used to identify areas with heavy educational AI use and thus estimate the educational impact of AI as it is actually used. I find no meaningful impact of AI usage on high school test score averages in either direction. These results imply that, to the extent that high school students use AI to avoid learning, it either does not matter much for their test performance or is cancelled out by positive uses of AI in the aggregate.

econ.GN

Do LLMs Act as Repositories of Causal Knowledge?

Large language models (LLMs) offer the potential to automate a large number of tasks that previously have not been possible to automate, including some in science. There is considerable interest in whether LLMs can automate the process of causal inference by providing the information about causal links necessary to build a structural model. We use the case of confounding in the Coronary Drug Project (CDP), for which there are several studies listing expert-selected confounders that can serve as a ground truth. LLMs exhibit mediocre performance in identifying confounders in this setting, even though text about the ground truth is in their training data. Variables that experts identify as confounders are only slightly more likely to be labeled as confounders by LLMs compared to variables that experts consider non-confounders. Further, LLM judgment on confounder status is highly inconsistent across models, prompts, and irrelevant concerns like multiple-choice option ordering. LLMs do not yet have the ability to automate the reporting of causal links.

econ.EM

Linear Rescaling to Accurately Interpret Logarithms

The standard approximation of a natural logarithm in statistical analysis interprets a linear change of \(p\) in \(\ln(X)\) as a \((1+p)\) proportional change in \(X\), which is only accurate for small values of \(p\). I suggest base-\((1+p)\) logarithms, where \(p\) is chosen ahead of time. A one-unit change in \(\log_{1+p}(X)\) is exactly equivalent to a \((1+p)\) proportional change in \(X\). This avoids an approximation applied too broadly, makes exact interpretation easier and less error-prone, improves approximation quality when approximations are used, makes the change of interest a one-log-unit change like other regression variables, and reduces error from the use of \(\log(1+X)\).

econ.EM