SearcharxivSearch

arXiv subjects

Dominik Papies

Publications and source records attributed to Dominik Papies.

3 recordsLinked to original sources

Double Machine Learning meets Panel Data -- Promises, Pitfalls, and Potential Solutions

Estimating causal effect using machine learning (ML) algorithms can help to relax functional form assumptions if used within appropriate frameworks. However, most of these frameworks assume settings with cross-sectional data, whereas researchers often have access to panel data, which in traditional methods helps to deal with unobserved heterogeneity between units. In this paper, we explore how we can adapt double/debiased machine learning (DML) (Chernozhukov et al., 2018) for panel data in the presence of unobserved heterogeneity. This adaptation is challenging because DML's cross-fitting procedure assumes independent data and the unobserved heterogeneity is not necessarily additively separable in settings with nonlinear observed confounding. We assess the performance of several intuitively appealing estimators in a variety of simulations. While we find violations of the cross-fitting assumptions to be largely inconsequential for the accuracy of the effect estimates, many of the considered methods fail to adequately account for the presence of unobserved heterogeneity. However, we find that using predictive models based on the correlated random effects approach (Mundlak, 1978) within DML leads to accurate coefficient estimates across settings, given a sample size that is large relative to the number of observed confounders. We also show that the influence of the unobserved heterogeneity on the observed confounders plays a significant role for the performance of most alternative methods.

econ.EM

Does TikTok Promote or Cannibalize Music Streaming? Estimands and Identification with Heavy-Tailed Outcomes

We study how TikTok affects demand for music on paid streaming platforms. We use Universal Music Group's (UMG) global withdrawal of its catalog from TikTok as a quasi-natural experiment. Recent work using this setting reaches mixed conclusions about whether TikTok promotes or cannibalizes streaming demand. We show that these findings can be reconciled by making the estimand explicit: with heavy-tailed exposure and outcomes, common difference-in-differences (DiD) implementations in levels, logs, and Poisson answer different economic questions. In our data, the top 10% of songs account for 96% of TikTok creations and 76% of Spotify streams, which makes the distinction between the typical song and the economically consequential song central. We find that removing TikTok access lowers Spotify demand for UMG titles, with losses concentrated among viral songs and little economically meaningful change for the long tail. Because the viral head accounts for a disproportionate share of listening and revenue, these losses drive aggregate implications. A TikTok creator-side analysis shows that some activity reallocates toward non-UMG audio when UMG content is unavailable. This substitution is limited in magnitude but economically relevant for interpreting the treatment effect because streaming compensation depends on relative stream shares. Finally, using the 2025 U.S. TikTok outage, which affected all labels symmetrically and is not subject to the label-specific spillover concern as the UMG withdrawal, we find corroborating evidence that disruptions to TikTok access reduce monetized streaming. We also provide a practitioner companion that guides the choice of DiD estimands, estimators, and diagnostics in heavy-tailed outcome settings.

econ.GN

Estimating Causal Effects with Double Machine Learning -- A Method Evaluation

The estimation of causal effects with observational data continues to be a very active research area. In recent years, researchers have developed new frameworks which use machine learning to relax classical assumptions necessary for the estimation of causal effects. In this paper, we review one of the most prominent methods - "double/debiased machine learning" (DML) - and empirically evaluate it by comparing its performance on simulated data relative to more traditional statistical methods, before applying it to real-world data. Our findings indicate that the application of a suitably flexible machine learning algorithm within DML improves the adjustment for various nonlinear confounding relationships. This advantage enables a departure from traditional functional form assumptions typically necessary in causal effect estimation. However, we demonstrate that the method continues to critically depend on standard assumptions about causal structure and identification. When estimating the effects of air pollution on housing prices in our application, we find that DML estimates are consistently larger than estimates of less flexible methods. From our overall results, we provide actionable recommendations for specific choices researchers must make when applying DML in practice.

stat.ML