SearcharxivSearch

arXiv subjects

George Perrett

Publications and source records attributed to George Perrett.

2 recordsLinked to original sources

Measurement Induced Confounding

A critical assumption of observational studies is that all confounding variables must be known and sufficiently adjusted for to estimate causal effects. An implicit, and often overlooked, aspect of this assumption is that all confounding variables have been measured without error. In the social and medical sciences, latent traits such as motivation, self-efficacy, and ability measures are likely confounding variables. Because latent traits are not directly observable, conventional approaches to adjust for them in observational studies rely on collecting responses to individual items on a test or survey instrument and then adjust for sum scores, measurement model-derived ability estimates, or item responses directly. Through a process we describe as measurement induced confounding, we show that measurement error propagates through the estimation process and that current conventional approaches to adjusting for latent traits in observational studies produce biased estimates of the average treatment effect with incorrectly calibrated coverage properties. A critical implication of this finding is that current observational studies that attempt to adjust for latent confounding variables likely put forth biased causal estimates with incorrect uncertainty intervals. We show that measurement induced confounding can be resolved through a Bayesian Joint Estimation approach that simultaneously estimates the measurement model, the treatment assignment model, and the response model.

stat.ME

Flaws in the LLM Automation Narrative

Large Language Models (LLMs) are increasingly described as performing at the level of human experts on knowledge economy tasks. These claims are primarily based on how LLMs perform on benchmarking tasks that measure average performance across standardized datasets. Primary limitations of many benchmarking tasks are that they often measure performance based on content directly included in LLM training data, and they frequently do not assess the reliability of LLM performance or the magnitude of LLM errors. However, in high stakes contexts, these qualities are critically important. Through a novel LLM benchmarking task that requires writing computer code to complete a data analysis task, we compare the performance of a frontier LLM against submissions from human experts and explicitly measure the variance of responses and the magnitude of errors. Our study reveals that the human experts perform better on average on a range of metrics and demonstrate less variability in performance. Our results provide evidence that LLMs do not consistently perform at the level of human experts and demonstrate the importance of measuring variance and assessing error magnitude in LLM benchmark evaluations.

stat.OT