Searcharxiv⌕ Search

arXiv subjects

Malte S. Kurz

Publications and source records attributed to Malte S. Kurz.

7 recordsLinked to original sources

DoubleML -- An Object-Oriented Implementation of Double Machine Learning in R

The R package DoubleML implements the double/debiased machine learning framework of Chernozhukov et al. (2018). It provides functionalities to estimate parameters in causal models based on machine learning methods. The double machine learning framework consist of three key ingredients: Neyman orthogonality, high-quality machine learning estimation and sample splitting. Estimation of nuisance components can be performed by various state-of-the-art machine learning methods that are available in the mlr3 ecosystem. DoubleML makes it possible to perform inference in a variety of causal models, including partially linear and interactive regression models and their extensions to instrumental variable estimation. The object-oriented implementation of DoubleML enables a high flexibility for the model specification and makes it easily extendable. This paper serves as an introduction to the double machine learning framework and the R package DoubleML. In reproducible code examples with simulated and real data sets, we demonstrate how DoubleML users can perform valid inference based on machine learning methods.

stat.ML↗

Vine copula based knockoff generation for high-dimensional controlled variable selection

Vine copulas are a flexible tool for high-dimensional dependence modeling. In this article, we discuss the generation of approximate model-X knockoffs with vine copulas. It is shown how Gaussian knockoffs can be generalized to Gaussian copula knockoffs. A convenient way to parametrize Gaussian copulas are partial correlation vines. We discuss how completion problems for partial correlation vines are related to Gaussian knockoffs. A natural generalization of partial correlation vines are vine copulas which are well suited for the generation of approximate model-X knockoffs. We discuss a specific D-vine structure which is advantageous to obtain vine copula knockoff models. In a simulation study, we demonstrate that vine copula knockoff models are effective and powerful for high-dimensional controlled variable selection.

stat.ME↗

Testing the simplifying assumption in high-dimensional vine copulas

Testing the simplifying assumption in high-dimensional vine copulas is a difficult task. Tests must be based on estimated observations and check constraints on high-dimensional distributions. So far, corresponding tests have been limited to single conditional copulas with a low-dimensional set of conditioning variables. We propose a novel testing procedure that is computationally feasible for high-dimensional data sets and that exhibits a power that decreases only slightly with the dimension. By discretizing the support of the conditioning variables and incorporating a penalty in the test statistic, we mitigate the curse of dimensionality by looking for the possibly strongest deviation from the simplifying assumption. The use of a decision tree renders the test computationally feasible for large dimensions. We derive the asymptotic distribution of the test and analyze its finite sample performance in an extensive simulation study. An application of the test to four real data sets is provided.

stat.ME↗

DoubleML -- An Object-Oriented Implementation of Double Machine Learning in Python

DoubleML is an open-source Python library implementing the double machine learning framework of Chernozhukov et al. (2018) for a variety of causal models. It contains functionalities for valid statistical inference on causal parameters when the estimation of nuisance parameters is based on machine learning methods. The object-oriented implementation of DoubleML provides a high flexibility in terms of model specifications and makes it easily extendable. The package is distributed under the MIT license and relies on core libraries from the scientific Python ecosystem: scikit-learn, numpy, pandas, scipy, statsmodels and joblib. Source code, documentation and an extensive user guide can be found at https://github.com/DoubleML/doubleml-for-py and https://docs.doubleml.org.

stat.ML↗

Distributed Double Machine Learning with a Serverless Architecture

This paper explores serverless cloud computing for double machine learning. Being based on repeated cross-fitting, double machine learning is particularly well suited to exploit the high level of parallelism achievable with serverless computing. It allows to get fast on-demand estimations without additional cloud maintenance effort. We provide a prototype Python implementation \texttt{DoubleML-Serverless} for the estimation of double machine learning models with the serverless computing platform AWS Lambda and demonstrate its utility with a case study analyzing estimation times and costs.

cs.DC↗

The partial vine copula: A dependence measure and approximation based on the simplifying assumption

Simplified vine copulas (SVCs), or pair-copula constructions, have become an important tool in high-dimensional dependence modeling. So far, specification and estimation of SVCs has been conducted under the simplifying assumption, i.e., all bivariate conditional copulas of the vine are assumed to be bivariate unconditional copulas. We introduce the partial vine copula (PVC) which provides a new multivariate dependence measure and which plays a major role in the approximation of multivariate distributions by SVCs. The PVC is a particular SVC where to any edge a j-th order partial copula is assigned and constitutes a multivariate analogue of the bivariate partial copula. We investigate to what extent the PVC describes the dependence structure of the underlying copula. We show that the PVC does not minimize the Kullback-Leibler divergence from the true copula and that the best approximation satisfying the simplifying assumption is given by a vine pseudo-copula. However, under regularity conditions, step-wise estimators of pair-copula constructions converge to the PVC irrespective of whether the simplifying assumption holds or not. Moreover, we elucidate why the PVC is the best feasible SVC approximation in practice.

stat.ME↗

The partial copula: Properties and associated dependence measures

The partial correlation coefficient is a commonly used measure to assess the conditional dependence between two random variables. We provide a thorough explanation of the partial copula, which is a natural generalization of the partial correlation coefficient, and investigate several of its properties. In addition, properties of some associated partial dependence measures are examined.

stat.ME↗