Searcharxiv⌕ Search

arXiv subjects

Paul S. Clarke

Publications and source records attributed to Paul S. Clarke.

4 recordsLinked to original sources

Double Machine Learning with High-dimensional Interactive Fixed Effects

Factor structures are central to empirical work in economics and finance, and are usually used to model time-varying unobserved heterogeneity through interactive fixed effects (IFE). Existing IFE estimators rest on low-dimensional and linear specifications in the covariates, assumptions which are increasingly restrictive in applications drawing on rich datasets with controls of unknown functional form. This paper develops a Double Machine Learning estimator for the high-dimensional partially linear panel model with interactive fixed effects (panel DML-IFE). The method combines projection-based defactorisation of the data, in the spirit of Common Correlated Effects (CCE), with a Neyman-orthogonal score function and cross-fitting procedure, and accommodates low-rank factor structures in outcomes and treatments alongside high-dimensional, potentially nonlinear covariate effects estimated by machine learning algorithms. Monte Carlo simulations show that panel DML-IFE outperforms conventional IFE estimator outside the correctly-specified linear case, with bias reduction driven primarily by the time and covariate dimensions. An empirical application to U.S. stock returns shows that several effects documented under linear specifications lose statistical significance once high-dimensional nonlinear confounding and the presence of IFE are jointly accounted for.

econ.EM↗

Double Machine Learning for Static Panel Data with Instrumental Variables: New Method and Applications

Panel data methods are widely used in empirical analysis to address unobserved heterogeneity, but causal inference remains challenging when treatments are endogenous and confounding variables high-dimensional and potentially nonlinear. Standard instrumental variables (IV) estimators, such as two-stage least squares (2SLS), become unreliable when instrument validity requires flexibly conditioning on many covariates with potentially non-linear effects. This paper develops a Double Machine Learning estimator for static panel models with endogenous treatments (panel IV DML), and introduces weak-identification diagnostics for it. We revisit three influential migration studies that use shift-share instruments. In these settings, instrument validity depends on a rich covariate adjustment. In one application, panel IV DML strengthens the predictive power of the instrument and broadly confirms 2SLS results. In the other cases, flexible adjustment makes the instruments weak, leading to substantially more cautious causal inference than conventional 2SLS. Monte Carlo evidence supports these findings, showing that panel IV DML improves estimation accuracy under strong instruments and delivers more reliable inference under weak identification.

econ.EM↗

Double Machine Learning for Static Panel Models with Fixed Effects

Recent advances in causal inference have seen the development of methods which make use of the predictive power of machine learning algorithms. In this paper, we develop novel double machine learning (DML) procedures for panel data in which these algorithms are used to approximate high-dimensional and nonlinear nuisance functions of the covariates. Our new procedures are extensions of the well-known correlated random effects, within-group and first-difference estimators from linear to nonlinear panel models, specifically, Robinson (1988)'s partially linear regression model with fixed effects and unspecified nonlinear confounding. Our simulation study assesses the performance of these procedures using different machine learning algorithms. We use our procedures to re-estimate the impact of minimum wage on voting behaviour in the UK. From our results, we recommend the use of first-differencing because it imposes the fewest constraints on the distribution of the fixed effects, and an ensemble learning strategy to ensure optimum estimator accuracy.

econ.EM↗

Estimating Structural Mean Models with Multiple Instrumental Variables Using the Generalised Method of Moments

Instrumental variables analysis using genetic markers as instruments is now a widely used technique in epidemiology and biostatistics. As single markers tend to explain only a small proportion of phenotypic variation, there is increasing interest in using multiple genetic markers to obtain more precise estimates of causal parameters. Structural mean models (SMMs) are semiparametric models that use instrumental variables to identify causal parameters. Recently, interest has started to focus on using these models with multiple instruments, particularly for multiplicative and logistic SMMs. In this paper we show how additive, multiplicative and logistic SMMs with multiple orthogonal binary instrumental variables can be estimated efficiently in models with no further (continuous) covariates, using the generalised method of moments (GMM) estimator. We discuss how the Hansen J-test can be used to test for model misspecification, and how standard GMM software routines can be used to fit SMMs. We further show that multiplicative SMMs, like the additive SMM, identify a weighted average of local causal effects if selection is monotonic. We use these methods to reanalyse a study of the relationship between adiposity and hypertension using SMMs with two genetic markers as instruments for adiposity. We find strong effects of adiposity on hypertension.

stat.ME↗