SearcharxivSearch

arXiv subjects

Brett R. Gordon

Publications and source records attributed to Brett R. Gordon.

5 recordsLinked to original sources

Characterizing and Minimizing Divergent Delivery in Meta Advertising Experiments

Many digital platforms offer advertisers experimentation tools like Meta's Lift and A/B tests to optimize their ad campaigns. Lift tests compare outcomes between users eligible to see ads versus users in a no-ad control group. In contrast, A/B tests compare users exposed to alternative ad configurations, absent any control group. The latter setup raises the prospect of divergent delivery: ad delivery algorithms may target different ad variants to different audience segments. This complicates causal interpretation because results may reflect both ad content effectiveness and changes to audience composition. We offer three key contributions. First, we make clear that divergent delivery is specific to A/B tests and intentional, informing advertisers about ad performance in practice. Second, we measure divergent delivery at scale, considering 3,204 Lift tests and 181,890 A/B tests. Lift tests show no meaningful audience imbalance, confirming their causal validity, while A/B tests show clear imbalance, as expected. Third, we demonstrate that campaign configuration choices can reduce divergent delivery in A/B tests, lessening algorithmic influence on results. While no configuration guarantees eliminating divergent delivery entirely, we offer evidence-based guidance for those seeking more generalizable insights about ad content in A/B tests.

econ.GN

Amazon Ads Multi-Touch Attribution

Amazon's new Multi-Touch Attribution (MTA) solution allows advertisers to measure how each touchpoint across the marketing funnel contributes to a conversion. This gives advertisers a more comprehensive view of their Amazon Ads performance across objectives when multiple ads influence shopping decisions. Amazon MTA uses a combination of randomized controlled trials (RCTs) and machine learning (ML) models to allocate credit for Amazon conversions across Amazon Ads touchpoints in proportion to their value, i.e., their likely contribution to shopping decisions. ML models trained purely on observational data are easy to scale and can yield precise predictions, but the models might produce biased estimates of ad effects. RCTs yield unbiased ad effects but can be noisy. Our MTA methodology combines experiments, ML models, and Amazon's shopping signals in a thoughtful manner to inform attribution credit allocation.

econ.EM

Predicted Incrementality by Experimentation (PIE) for Ad Measurement

Randomized controlled trials (RCTs) provide the most credible estimates of advertising incrementality but are difficult to scale. We propose Predicted Incrementality by Experimentation (PIE), which reframes ad measurement as a campaign-level prediction problem. PIE uses a sample of RCTs to learn a mapping from campaign features to causal effects, then applies it to campaigns not run as RCTs. Because the RCTs identify the causal effects, PIE can incorporate post-determined features -- campaign-level aggregates such as test-group outcomes, exposure rates, and last-click conversions, computed after campaign completion. These metrics reflect the consumer behaviors that generate treatment effects, so they carry predictive information about incrementality even though they would be invalid controls in a causal model. Using 2,226 Meta ad experiments, PIE achieves an out-of-sample $R^2 = 0.88$ for incremental conversions per dollar, compared to $R^2 = 0.19$ for industry-standard 7-day last-click attribution. In a decision-making framework, PIE disagrees with RCT-based decisions in only 8-12% of campaigns, compared to 12-20% for last-click attribution. We conclude that PIE can help scale causal measurement from a limited number of RCTs to a large set of non-experimental campaigns.

econ.EM

Multicell experiments for marginal treatment effect estimation of digital ads

Randomized experiments with treatment and control groups are an important tool to measure the impacts of interventions. However, in experimental settings with one-sided noncompliance extant empirical approaches may not produce the estimands a decision maker needs to solve the problem of interest. For example, these experimental designs are common in digital advertising settings but typical methods do not yield effects that inform the intensive margin: how many consumers should be reached or how much should be spent on a campaign. We propose a solution that combines a novel multicell experimental design with modern estimation techniques that enables decision makers to solve problems with an intensive margin. Our design is straightforward to implement and does not require additional budget. We illustrate our method through simulations calibrated using an advertising experiment at Facebook, demonstrating its superior performance in various scenarios and its advantage over direct optimization approaches.

econ.EM

Close Enough? A Large-Scale Exploration of Non-Experimental Approaches to Advertising Measurement

Despite their popularity, randomized controlled trials (RCTs) are not always available for the purposes of advertising measurement. Non-experimental data is thus required. However, Facebook and other ad platforms use complex and evolving processes to select ads for users. Therefore, successful non-experimental approaches need to "undo" this selection. We analyze 663 large-scale experiments at Facebook to investigate whether this is possible with the data typically logged at large ad platforms. With access to over 5,000 user-level features, these data are richer than what most advertisers or their measurement partners can access. We investigate how accurately two non-experimental methods -- double/debiased machine learning (DML) and stratified propensity score matching (SPSM) -- can recover the experimental effects. Although DML performs better than SPSM, neither method performs well, even using flexible deep learning models to implement the propensity and outcome models. The median RCT lifts are 29%, 18%, and 5% for the upper, middle, and lower funnel outcomes, respectively. Using DML (SPSM), the median lift by funnel is 83% (173%), 58% (176%), and 24% (64%), respectively, indicating significant relative measurement errors. We further characterize the circumstances under which each method performs comparatively better. Overall, despite having access to large-scale experiments and rich user-level data, we are unable to reliably estimate an ad campaign's causal effect.

econ.EM