SearcharxivSearch

arXiv subjects

Lo-Hua Yuan

Publications and source records attributed to Lo-Hua Yuan.

5 recordsLinked to original sources

Evaluating for the long term: Learnings from industry

Online platforms prioritize long-term business outcomes, yet typical experiments are far too short to measure these outcomes directly. Our goal in this paper is to collect and share industry knowledge on how to make decisions from short-term experiments that are better aligned with long-term outcomes. Based on a daylong workshop with 26 experts from 15 online platforms and 4 universities, we formulate a series of propositions that reflect current industry knowledge. Participants largely agreed that reversals of sign from short-run to long-run treatment effects are rare, with reversals concentrating in specific cases such as treatments involving content quality signals, hyper-monetization, and pricing. Although the magnitude of treatment effects can shift over time, a "univariate autosurrogate", corresponding to the short-run treatment effect on the long-run metric of interest, is often hard to beat. A recurring theme was the importance of surrogates that are not only (or even primarily) unbiased for true long-run outcomes, but that improve decision-making. Thus, participants generally agreed that simple, interpretable surrogates were generally preferable to elaborate but hard-to-explain surrogate indices. Participants also agreed that, due to concerns about confounding and transportability, experimentally-learned surrogates are generally preferable to observationally-learned surrogates. However, the drawback is that learning good surrogates from experiments typically requires a large, representative portfolio of long-run experiments that few platforms possess. We conclude that there is no substitute for a well-run long-term experiment, whether for learning surrogates or validating them, and we highlight open challenges including evolving treatments, persistent treatments not fully mediated by short-term proxies, and mismatch between experimental samples and the target population.

stat.AP

AI-Assisted Variance Reduction in Randomized Experiments

Generative AI and large language models can produce realistic predictions of human behavior from rich, unstructured inputs with little to no task-specific training data. Recent work uses these ``digital twin'' predictions to supplement human responses in surveys and experiments. We study the special case of using AI-generated predictions to reduce variance in randomized experiments. We argue that doing so requires no new estimators and that researchers can simply include AI predictions as covariates in standard regression adjustment, analogous to adjusting for a prognostic score. A benefit of this approach is a ``do no harm'' property whereby the adjusted estimator reverts to the unadjusted difference in means when predictions are uninformative. Other methods, such as variants of prediction-powered inference, do not have this guarantee. We provide implementation guidance, including how to obtain continuous scores from discrete LLM outputs and how to use LLMs to featurize unstructured inputs as auxiliary covariates. We demonstrate these ideas in simulations and three empirical applications: a survey mega-study, an email marketing A/B test, and a large-scale technology platform experiment. Overall, efficiency gains are real if modest, with greater benefits in studies that contain substantial text and other unstructured data. We also confirm the do no harm property empirically. Given these gains and limited costs, we recommend adjusting for AI-generated predictions as a regular empirical practice.

econ.EM

From Augmentation to Decomposition: A New Look at CUPED in 2023

Ten years ago, CUPED (Controlled Experiments Utilizing Pre-Experiment Data) mainstreamed the idea of variance reduction leveraging pre-experiment covariates. Since its introduction, it has been implemented, extended, and modernized by major online experimentation platforms. Many researchers and practitioners often interpret CUPED as a regression adjustment. In this article, we clarify its similarities and differences to regression adjustment and present CUPED as a more general augmentation framework which is closer to the spirit of the 2013 paper. We show that the augmentation view naturally leads to cleaner developments of variance reduction beyond simple average metrics, including ratio metrics and percentile metrics. Moreover, the augmentation view can go beyond using pre-experiment data and leverage in-experiment data, leading to significantly larger variance reduction. We further introduce metric decomposition using approximate null augmentation (ANA) as a mental model for in-experiment variance reduction. We study it under both a Bayesian framework and a frequentist optimal proxy metric framework. Metric decomposition arises naturally in conversion funnels, so this work has broad applicability.

stat.AP

Zero to Hero: Exploiting Null Effects to Achieve Variance Reduction in Experiments with One-sided Triggering

In online experiments where the intervention is only exposed, or "triggered", for a small subset of the population, it is critical to use variance reduction techniques to estimate treatment effects with sufficient precision to inform business decisions. Trigger-dilute analysis is often used in these situations, and reduces the sampling variance of overall intent-to-treat (ITT) effects by an order of magnitude equal to the inverse of the triggering rate; for example, a triggering rate of $5\%$ corresponds to roughly a $20x$ reduction in variance. To apply trigger-dilute analysis, one needs to know experimental subjects' triggering counterfactual statuses, i.e., the counterfactual behavior of subjects under both treatment and control conditions. In this paper, we propose an unbiased ITT estimator with reduced variance applicable for experiments where the triggering counterfactual status is only observed in the treatment group. Our method is based on the efficiency augmentation idea of CUPED and draws upon identification frameworks from the principal stratification and instrumental variables literature. The unbiasedness of our estimation approach relies on a testable assumption that the augmentation term used for covariate adjustment equals zero in expectation. Unlike traditional covariate adjustment or principal score modeling approaches, our estimator can incorporate both pre-experiment and in-experiment observations. We demonstrate through a real-world experiment and simulations that our estimator can remain unbiased and achieve precision improvements as large as if triggering status were fully observed, and in some cases can even outperform trigger-dilute analysis.

stat.ME

Identifying and Estimating Principal Causal Effects in Multi-site Trials

Randomized trials are often conducted with separate randomizations across multiple sites such as schools, voting districts, or hospitals. These sites can differ in important ways, including the site's implementation, local conditions, and the composition of individuals. An important question in practice is whether---and under what assumptions---researchers can leverage this cross-site variation to learn more about the intervention. We address these questions in the principal stratification framework, which describes causal effects for subgroups defined by post-treatment quantities. We show that researchers can estimate certain principal causal effects via the multi-site design if they are willing to impose the strong assumption that the site-specific effects are uncorrelated with the site-specific distribution of stratum membership. We motivate this approach with a multi-site trial of the Early College High School Initiative, a unique secondary education program with the goal of increasing high school graduation rates and college enrollment. Our analyses corroborate previous studies suggesting that the initiative had positive effects for students who would have otherwise attended a low-quality high school, although power is limited.

stat.ME