Searcharxiv⌕ Search

arXiv subjects

Niels Richard Hansen

Publications and source records attributed to Niels Richard Hansen.

At least 19 recordsLinked to original sources

Outcome-adapted Automatic Debiased Machine Learning

Parameters of interest in causal inference, such as treatment or policy effects, can often be expressed as linear functionals of an outcome regression function. Automatic debiased machine learning (AutoDML) is a unified framework for obtaining asymptotically normal estimators of such parameters, which requires estimation of both a regression function and a Riesz representer. Existing AutoDML neural network architectures, such as RieszNet and MADNet, use a shared intermediate covariate representation. However, it remains unclear whether this shared representation should be predictive of the Riesz representer or the outcome. We show that a shared representation of the covariates that preserves predictive power of the outcome while discarding information about the Riesz representer is asymptotically more efficient than the baseline AutoDML estimator that uses all covariates. Motivated by these results, we propose the outcome-adapted AutoDML estimator and establish its asymptotic behavior in a sample splitting framework. We provide a neural network implementation of the estimator that learns a sparse representation of the covariates that is predictive of the outcome but not predictive of the Riesz representer. We demonstrate the efficiency gains of our estimator over existing alternatives on synthetic data and achieve state-of-the-art estimation accuracy on the semi-synthetic IHDP benchmark dataset.

stat.ME↗

Identifiability and Estimation in Continuous Lyapunov Models

Cross-sectional observations from a dynamical system can be modeled via steady-state distributions of Markov processes. The major challenge is then to determine whether the process parameters can be identified and estimated from the steady-state distributions. We study this problem for continuous Lyapunov models that arise as steady-state distributions of the solution to a multivariate stochastic differential equation, whose linear drift matrix is parametrized by a directed graph. We derive equations for the cumulant tensors of any order for this distribution, which generalize the well-known covariance Lyapunov equation. Under a non-Gaussianity assumption we prove generic identifiability of the drift matrix for any connected graph using the equations for the higher-order cumulants. Based on the identifiability result, we propose a new semiparametric estimator of the drift matrix, and we derive its asymptotic distribution. A simulation study demonstrates the asymptotic validity of the estimator but shows that it is only accurate for relatively large sample sizes, illustrating the hardness of the unconstrained estimation problem.

math.ST↗

Efficient adjustment for complex covariates: Gaining efficiency with DOPE

Covariate adjustment is a ubiquitous method used to estimate the average treatment effect (ATE) from observational data. Assuming a known graphical structure of the data generating model, recent results give graphical criteria for optimal adjustment, which enables efficient estimation of the ATE. However, graphical approaches are challenging for high-dimensional and complex data, and it is not straightforward to specify a meaningful graphical model of non-Euclidean data such as texts. We propose a new framework that accommodates adjustment for any subset of information expressed by the covariates, and we show that the information that is minimally sufficient for prediction of the outcome given the treatment is also most efficient for adjustment. Based on our theoretical results, we propose the Debiased Outcome-adapted Propensity Estimator (DOPE) for efficient estimation of the ATE, and we provide asymptotic results for DOPE under general conditions. Compared to the augmented inverse propensity weighted (AIPW) estimator, DOPE can retain its efficiency even when the covariates are highly predictive of treatment. We illustrate this with a single-index model, and with an implementation of DOPE based on neural networks, we demonstrate its performance on simulated and real data. Our results show that DOPE provides an efficient and robust methodology for ATE estimation in various observational settings.

math.ST↗

A Trek Rule for the Lyapunov Equation

The Lyapunov equation is a linear matrix equation characterizing the cross-sectional steady-state covariance matrix of a Gaussian Markov process. We show a new version of the trek rule for this equation, which links the graphical structure of the drift of the process to the entries of the steady-state covariance matrix. In general, the trek rule is a power series expansion of the covariance matrix in the entries of the drift and volatility matrices. For acyclic models it simplifies to a polynomial in the off-diagonal entries of the drift matrix. Using the trek rule we can give relatively explicit formulas for the entries of the covariance matrix for some special cases of the drift matrix. These results illustrate notable differences between covariance models entailed by the Lyapunov equation and those entailed by linear additive noise models. To further explore differences and similarities between these two model classes, we use the trek rule to derive a new lower bound on the marginal variances in the acyclic case. This sheds light on the phenomenon, well known for the linear additive noise model, that the variances in the acyclic case tend to increase along a topological ordering of the variables.

math.ST↗

Substitute adjustment via recovery of latent variables

The deconfounder was proposed as a method for estimating causal parameters in a context with multiple causes and unobserved confounding. It is based on recovery of a latent variable from the observed causes. We disentangle the causal interpretation from the statistical estimation problem and show that the deconfounder in general estimates adjusted regression target parameters. It does so by outcome regression adjusted for the recovered latent variable termed the substitute. We refer to the general algorithm, stripped of causal assumptions, as substitute adjustment. We give theoretical results to support that substitute adjustment estimates adjusted regression parameters when the regressors are conditionally independent given the latent variable. We also introduce a variant of our substitute adjustment algorithm that estimates an assumption-lean target parameter with minimal model assumptions. We then give finite sample bounds and asymptotic results supporting substitute adjustment estimation in the case where the latent variable takes values in a finite set. A simulation study illustrates finite sample properties of substitute adjustment. Our results support that when the latent variable model of the regressors hold, substitute adjustment is a viable method for adjusted regression.

math.ST↗

Identifiability in Continuous Lyapunov Models

The recently introduced graphical continuous Lyapunov models provide a new approach to statistical modeling of correlated multivariate data. The models view each observation as a one-time cross-sectional snapshot of a multivariate dynamic process in equilibrium. The covariance matrix for the data is obtained by solving a continuous Lyapunov equation that is parametrized by the drift matrix of the dynamic process. In this context, different statistical models postulate different sparsity patterns in the drift matrix, and it becomes a crucial problem to clarify whether a given sparsity assumption allows one to uniquely recover the drift matrix parameters from the covariance matrix of the data. We study this identifiability problem by representing sparsity patterns by directed graphs. Our main result proves that the drift matrix is globally identifiable if and only if the graph for the sparsity pattern is simple (i.e., does not contain directed two-cycles). Moreover, we present a necessary condition for generic identifiability and provide a computational classification of small graphs with up to 5 nodes.

math.ST↗

Nonparametric Conditional Local Independence Testing

Conditional local independence is an asymmetric independence relation among continuous time stochastic processes. It describes whether the evolution of one process is directly influenced by another process given the histories of additional processes, and it is important for the description and learning of causal relations among processes. We develop a model-free framework for testing the hypothesis that a counting process is conditionally locally independent of another process. To this end, we introduce a new functional parameter called the Local Covariance Measure (LCM), which quantifies deviations from the hypothesis. Following the principles of double machine learning, we propose an estimator of the LCM and a test of the hypothesis using nonparametric estimators and sample splitting or cross-fitting. We call this test the (cross-fitted) Local Covariance Test ((X)-LCT), and we show that its level and power can be controlled uniformly, provided that the nonparametric estimators are consistent with modest rates. We illustrate the theory by an example based on a marginalized Cox model with time-dependent covariates, and we show in simulations that when double machine learning is used in combination with cross-fitting, then the test works well without restrictive parametric assumptions.

math.ST↗

Soft Maximin Estimation for Heterogeneous Data

Extracting a common robust signal from data divided into heterogeneous groups can be difficult when each group -- in addition to the signal -- can contain large, unique variation components. Previously, maximin estimation has been proposed as a robust estimation method in the presence of heterogeneous noise. We propose soft maximin estimation as a computationally attractive alternative aimed at striking a balance between pooled estimation and (hard) maximin estimation. The soft maximin method provides a range of estimators, controlled by a parameter $ζ>0$, that interpolates pooled least squares estimation and maximin estimation. By establishing relevant theoretical properties we argue that the soft maximin method is both statistically sensibel and computationally attractive. We also demonstrate, on real and simulated data, that the soft maximin estimator can offer improvements over both pooled OLS and hard maximin in terms of predictive performance and computational complexity. A time and memory efficient implementation is provided in the R package \verb+SMME+ available on CRAN.

stat.ME↗

Local Independence Testing for Point Processes

Constraint based causal structure learning for point processes require empirical tests of local independence. Existing tests require strong model assumptions, e.g. that the true data generating model is a Hawkes process with no latent confounders. Even when restricting attention to Hawkes processes, latent confounders are a major technical difficulty because a marginalized process will generally not be a Hawkes process itself. We introduce an expansion similar to Volterra expansions as a tool to represent marginalized intensities. Our main theoretical result is that such expansions can approximate the true marginalized intensity arbitrarily well. Based on this we propose a test of local independence and investigate its properties in real and simulated data.

stat.ME↗

Testing Conditional Independence via Quantile Regression Based Partial Copulas

The partial copula provides a method for describing the dependence between two random variables $X$ and $Y$ conditional on a third random vector $Z$ in terms of nonparametric residuals $U_1$ and $U_2$. This paper develops a nonparametric test for conditional independence by combining the partial copula with a quantile regression based method for estimating the nonparametric residuals. We consider a test statistic based on generalized correlation between $U_1$ and $U_2$ and derive its large sample properties under consistency assumptions on the quantile regression procedure. We demonstrate through a simulation study that the resulting test is sound under complicated data generating distributions. Moreover, in the examples considered the test is competitive to other state-of-the-art conditional independence tests in terms of level and power, and it has superior power in cases with conditional variance heterogeneity of $X$ and $Y$ given $Z$.

math.ST↗

Graphical modeling of stochastic processes driven by correlated errors

We study a class of graphs that represent local independence structures in stochastic processes allowing for correlated error processes. Several graphs may encode the same local independencies and we characterize such equivalence classes of graphs. In the worst case, the number of conditions in our characterizations grows superpolynomially as a function of the size of the node set in the graph. We show that deciding Markov equivalence is coNP-complete which suggests that our characterizations cannot be improved upon substantially. We prove a global Markov property in the case of a multivariate Ornstein-Uhlenbeck process which is driven by correlated Brownian motions.

math.ST↗

Sparse Network Estimation for Dynamical Spatio-temporal Array Models

Neural field models represent neuronal communication on a population level via synaptic weight functions. Using voltage sensitive dye (VSD) imaging it is possible to obtain measurements of neural fields with a relatively high spatial and temporal resolution. The synaptic weight functions represent functional connectivity in the brain and give rise to a spatio-temporal dependence structure. We present a stochastic functional differential equation for modeling neural fields, which leads to a vector autoregressive model of the data via basis expansions of the synaptic weight functions and time and space discretization. Fitting the model to data is a pratical challenge as this represents a large scale regression problem. By using a 1-norm penalty in combination with localized basis functions it is possible to learn a sparse network representation of the functional connectivity of the brain, but still, the explicit construction of a design matrix can be computationally prohibitive. We demonstrate that by using tensor product basis expansions, the computation of the penalized estimator via a proximal gradient algorithm becomes feasible. It is crucial for the computations that the data is organized in an array as is the case for the three dimensional VSD imaging data. This allows for the use of array arithmetic that is both memory and time efficient.The proposed method is implemented and showcased in the R package dynamo available from CRAN.

stat.ME↗

Graphical continuous Lyapunov models

The linear Lyapunov equation of a covariance matrix parametrizes the equilibrium covariance matrix of a stochastic process. This parametrization can be interpreted as a new graphical model class, and we show how the model class behaves under marginalization and introduce a method for structure learning via $\ell_1$-penalized loss minimization. Our proposed method is demonstrated to outperform alternative structure learning algorithms in a simulation study, and we illustrate its application for protein phosphorylation network reconstruction.

stat.ML↗

Markov equivalence of marginalized local independence graphs

Symmetric independence relations are often studied using graphical representations. Ancestral graphs or acyclic directed mixed graphs with $m$-separation provide classes of symmetric graphical independence models that are closed under marginalization. Asymmetric independence relations appear naturally for multivariate stochastic processes, for instance in terms of local independence. However, no class of graphs representing such asymmetric independence relations, which is also closed under marginalization, has been developed. We develop the theory of directed mixed graphs with $μ$-separation and show that this provides a graphical independence model class which is closed under marginalization and which generalizes previously considered graphical representations of local independence. For statistical applications, it is pivotal to characterize graphs that induce the same independence relations as such a Markov equivalence class of graphs is the object that is ultimately identifiable from observational data. Our main result is that for directed mixed graphs with $μ$-separation each Markov equivalence class contains a maximal element which can be constructed from the independence relations alone. Moreover, we introduce the directed mixed equivalence graph as the maximal graph with edge markings. This graph encodes all the information about the edges that is identifiable from the independence relations, and furthermore it can be computed efficiently from the maximal graph.

math.ST↗

Learning Large Scale Ordinary Differential Equation Systems

Learning large scale nonlinear ordinary differential equation (ODE) systems from data is known to be computationally and statistically challenging. We present a framework together with the adaptive integral matching (AIM) algorithm for learning polynomial or rational ODE systems with a sparse network structure. The framework allows for time course data sampled from multiple environments representing e.g. different interventions or perturbations of the system. The algorithm AIM combines an initial penalised integral matching step with an adapted least squares step based on solving the ODE numerically. The R package episode implements AIM together with several other algorithms and is available from CRAN. It is shown that AIM achieves state-of-the-art network recovery for the in silico phosphoprotein abundance data from the eighth DREAM challenge with an AUROC of 0.74, and it is demonstrated via a range of numerical examples that AIM has good statistical properties while being computationally feasible even for large systems.

math.ST↗

A comment on Stein's unbiased risk estimate for reduced rank estimators

In the framework of matrix valued observables with low rank means, Stein's unbiased risk estimate (SURE) can be useful for risk estimation and for tuning the amount of shrinkage towards low rank matrices. This was demonstrated by Candès et al. (2013) for singular value soft thresholding, which is a Lipschitz continuous estimator. SURE provides an unbiased risk estimate for an estimator whenever the differentiability requirements for Stein's lemma are satisfied. Lipschitz continuity of the estimator is sufficient, but it is emphasized that differentiability Lebesgue almost everywhere isn't. The reduced rank estimator, which gives the best approximation of the observation with a fixed rank, is an example of a discontinuous estimator for which Stein's lemma actually applies. This was observed by Mukherjee et al. (2015), but the proof was incomplete. This brief note gives a sufficient condition for Stein's lemma to hold for estimators with discontinuities, which is then shown to be fulfilled for a class of spectral function estimators including the reduced rank estimator. Singular value hard thresholding does, however, not satisfy the condition, and Stein's lemma does not apply to this estimator.

math.ST↗

Degrees of Freedom for Piecewise Lipschitz Estimators

A representation of the degrees of freedom akin to Stein's lemma is given for a class of estimators of a mean value parameter in $\mathbb{R}^n$. Contrary to previous results our representation holds for a range of discontinues estimators. It shows that even though the discontinuities form a Lebesgue null set, they cannot be ignored when computing degrees of freedom. Estimators with discontinuities arise naturally in regression if data driven variable selection is used. Two such examples, namely best subset selection and lasso-OLS, are considered in detail in this paper. For lasso-OLS the general representation leads to an estimate of the degrees of freedom based on the lasso solution path, which in turn can be used for estimating the risk of lasso-OLS. A similar estimate is proposed for best subset selection. The usefulness of the risk estimates for selecting the number of variables is demonstrated via simulations with a particular focus on lasso-OLS.

math.ST↗

Penalized estimation in large-scale generalized linear array models

Large-scale generalized linear array models (GLAMs) can be challenging to fit. Computation and storage of its tensor product design matrix can be impossible due to time and memory constraints, and previously considered design matrix free algorithms do not scale well with the dimension of the parameter vector. A new design matrix free algorithm is proposed for computing the penalized maximum likelihood estimate for GLAMs, which, in particular, handles nondifferentiable penalty functions. The proposed algorithm is implemented and available via the R package \verb+glamlasso+. It combines several ideas -- previously considered separately -- to obtain sparse estimates while at the same time efficiently exploiting the GLAM structure. In this paper the convergence of the algorithm is treated and the performance of its implementation is investigated and compared to that of \verb+glmnet+ on simulated as well as real data. It is shown that the computation time for

stat.CO↗