SearcharxivSearch

arXiv subjects

Elina Robeva

Publications and source records attributed to Elina Robeva.

At least 19 recordsLinked to original sources

Generalization of Pearl's Front-Door Criterion

Pearl's front-door criterion provides a set of sufficient conditions for estimating the total causal effect from observational data in the presence of latent confounding, using the functional P(y | do(x := x*)) = \sum_z P(z | x*) \sum_x P(y | x, z) P(x). An open question is whether these conditions can be generalized to be both necessary and sufficient for the validity of this functional, similar to the generalization achieved for the back-door adjustment criterion by Shpitser. In this paper, we present a new, weakened set of graph-based conditions sufficient for the front-door formula to estimate the total causal effect, expanding the scope of problems amenable to front-door identification.

math.ST

Identifiability in Graphical Discrete Lyapunov Models

In this paper, we study discrete Lyapunov models, which consist of steady-state distributions of first-order vector autoregressive models. The parameter matrix of such a model encodes a directed graph whose vertices correspond to the components of the random vector. This combinatorial framework naturally allows for cycles in the graph structure. We focus on the fundamental problem of identifying the entries of the parameter matrix. In contrast to the classical setting, we assume non-Gaussian error terms, which allows us to use the higher-order cumulants of the model. In this setup, we show generic identifiability for directed acyclic graphs with self-loops at each vertex and show how to express the parameters as a rational function of the cumulants. Furthermore, we establish local identifiability for all directed graphs containing self loops at each vertex and no isolated vertices. Finally, we provide first results on the defining equations of the models, showing model equivalence for certain graphs and paving the way towards structure learning.

math.ST

Algebraic geometry of rational neural networks

We study the expressivity of rational neural networks (RationalNets) through the lens of algebraic geometry. We consider rational functions that arise from a given RationalNet to be tuples of fractions of homogeneous polynomials of fixed degrees. For a given architecture, the neuromanifold is the set of all such expressible tuples. For RationalNets with one hidden layer and fixed activation function $1/x$, we characterize the dimension of the neuromanifold and provide defining equations for some architectures. We also propose algorithms that determine whether a given rational function belongs to the neuromanifold. For deep binary RationalNets, i.e., RationalNets all of whose layers except potentially for the last one are binary, we classify when the Zarisky closure of the neuromanifold equals the whole ambient space, and give bounds on its dimensions.

math.AG

Causal Discovery for Linear Non-Gaussian Models with Disjoint Cycles

The paradigm of linear structural equation modeling readily allows one to incorporate causal feedback loops in the model specification. These appear as directed cycles in the common graphical representation of the models. However, the presence of cycles entails difficulties such as the fact that models need no longer be characterized by conditional independence relations. As a result, learning cyclic causal structures remains a challenging problem. In this paper, we offer new insights on this problem in the context of linear non-Gaussian models. First, we precisely characterize when two directed graphs determine the same linear non-Gaussian model. Next, we take up a setting of cycle-disjoint graphs, for which we are able to show that simple quadratic and cubic polynomial relations among low-order moments of a non-Gaussian distribution allow one to locate source cycles. Complementing this with a strategy of decorrelating cycles and multivariate regression allows one to infer a block-topological order among the directed cycles, which leads to a {consistent and computationally efficient algorithm} for learning causal structures with disjoint cycles.

math.ST

Path-Dependent SDEs: Solutions and Parameter Estimation

We develop a consistent method for estimating the parameters of a rich class of path-dependent SDEs, called signature SDEs, which can model general path-dependent phenomena. Path signatures are iterated integrals of a given path with the property that any sufficiently nice function of the path can be approximated by a linear functional of its signatures. This is why we model the drift and diffusion of our signature SDE as linear functions of path signatures. We provide conditions that ensure the existence and uniqueness of solutions to a general signature SDE. We then introduce the Expected Signature Matching Method (ESMM) for linear signature SDEs, which enables inference of the signature-dependent drift and diffusion coefficients from observed trajectories. Furthermore, we prove that ESMM is consistent: given sufficiently many samples and Picard iterations used by the method, the parameters estimated by the ESMM approach the true parameter with arbitrary precision. Finally, we demonstrate on a variety of empirical simulations that our ESMM accurately infers the drift and diffusion parameters from observed trajectories. While parameter estimation is often restricted by the need for a suitable parametric model, this work makes progress toward a completely general framework for SDE parameter estimation, using signature terms to model arbitrary path-independent and path-dependent processes.

math.ST

Gradient-flow SDEs have unique transient population dynamics

Identifying the drift and diffusion of an SDE from its population dynamics is a notoriously challenging task. Researchers in machine learning and single-cell biology have only been able to prove a partial identifiability result: for potential-driven SDEs, the gradient-flow drift can be identified from temporal marginals if the Brownian diffusivity is already known. Existing methods therefore assume that the diffusivity is known a priori, despite it being unknown in practice. We dispel the need for this assumption by providing a complete characterization of identifiability: the gradient-flow drift and Brownian diffusivity are jointly identifiable from temporal marginals if and only if the process is observed outside of equilibrium. Given this fundamental result, we propose nn-APPEX, the first Schrodinger Bridge-based inference method that can simultaneously learn the drift and diffusion of a gradient-flow SDE solely from observed marginals. Extensive experiments show that nn-APPEX's ability to adjust its diffusion estimate enables accurate inference, while previous Schrodinger Bridge methods obtain biased drift estimates due to their assumed, and likely incorrect, diffusion.

math.PR

Algebraic Constraints for Linear Acyclic Causal Models

In this paper we study the space of second- and third-order moment tensors of random vectors which satisfy a Linear Non-Gaussian Acyclic Model (LiNGAM). In such a causal model each entry $X_i$ of the random vector $X$ corresponds to a vertex $i$ of a directed acyclic graph $G$ and can be expressed as a linear combination of its direct causes $\{X_j: j\to i\}$ and random noise. For any directed acyclic graph $G$, we show that a random vector $X$ arises from a LiNGAM with graph $G$ if and only if certain easy-to-construct matrices, whose entries are second- and third-order moments of $X$, drop rank. This determinantal characterization extends previous results proven for polytrees and generalizes the well-known local Markov property for Gaussian models.

math.ST

Identifying Drift, Diffusion, and Causal Structure from Temporal Snapshots

Stochastic differential equations (SDEs) are a fundamental tool for modelling dynamic processes, including gene regulatory networks (GRNs), contaminant transport, financial markets, and image generation. However, learning the underlying SDE from data is a challenging task, especially if individual trajectories are not observable. Motivated by burgeoning research in single-cell datasets, we present the first comprehensive approach for jointly identifying the drift and diffusion of an SDE from its temporal marginals. Assuming linear drift and additive diffusion, we show that non-identifiability can only arise if the initial distribution possesses generalized rotational symmetries. We further prove that even if this condition holds, the drift and diffusion can almost always be recovered from the marginals. Additionally, we show that the causal graph of any SDE with additive diffusion can be recovered from the identified SDE parameters. To complement this theory, we adapt entropy-regularized optimal transport to handle anisotropic diffusion, and introduce APPEX (Alternating Projection Parameter Estimation from $X_0$), an iterative algorithm designed to estimate the drift, diffusion, and causal graph of an additive noise SDE, solely from temporal marginals. We show that APPEX iteratively decreases Kullback-Leibler divergence to the true solution, and demonstrate its effectiveness on simulated data from linear additive noise SDEs.

stat.ML

Causal Discovery of Linear Non-Gaussian Causal Models with Unobserved Confounding

We consider linear non-Gaussian structural equation models that involve latent confounding. In this setting, the causal structure is identifiable, but, in general, it is not possible to identify the specific causal effects. Instead, a finite number of different causal effects result in the same observational distribution. Most existing algorithms for identifying these causal effects use overcomplete independent component analysis (ICA), which often suffers from convergence to local optima. Furthermore, the number of latent variables must be known a priori. To address these issues, we propose an algorithm that operates recursively rather than using overcomplete ICA. The algorithm first infers a source, estimates the effect of the source and its latent parents on their descendants, and then eliminates their influence from the data. For both source identification and effect size estimation, we use rank conditions on matrices formed from higher-order cumulants. We prove asymptotic correctness under the mild assumption that locally, the number of latent variables never exceeds the number of observed variables. Simulation studies demonstrate that our method achieves comparable performance to overcomplete ICA even though it does not know the number of latents in advance.

stat.ML

Log-concave Density Estimation with Independent Components

We propose a method for estimating a log-concave density on $\mathbb R^d$ from samples, under the assumption that there exists an orthogonal transformation that makes the components of the random vector independent. While log-concave density estimation is hard both computationally and statistically, the independent components assumption alleviates both issues, while still maintaining a large non-parametric class. We prove that under mild conditions, at most $\tilde{\mathcal{O}}(\epsilon^{-4})$ samples (suppressing constants and log factors) suffice for our proposed estimator to be within $\epsilon$ of the original density in squared Hellinger distance. On the computational front, while the usual log-concave maximum likelihood estimate can be obtained via a finite-dimensional convex program, it is slow to compute -- especially in higher dimensions. We demonstrate through numerical experiments that our estimator can be computed efficiently, making it more practical to use.

math.ST

Causal Structure Learning in Directed, Possibly Cyclic, Graphical Models

We consider the problem of learning a directed graph $G^\star$ from observational data. We assume that the distribution which gives rise to the samples is Markov and faithful to the graph $G^\star$ and that there are no unobserved variables. We do not rely on any further assumptions regarding the graph or the distribution of the variables. Particularly, we allow for directed cycles in $G^\star$ and work in the fully non-parametric setting. Given the set of conditional independence statements satisfied by the distribution, we aim to find a directed graph which satisfies the same $d$-separation statements as $G^\star$. We propose a hybrid approach consisting of two steps. We first find a partially ordered partition of the vertices of $G^\star$ by optimizing a certain score in a greedy fashion. We prove that any optimal partition uniquely characterizes the Markov equivalence class of $G^\star$. Given an optimal partition, we propose an algorithm for constructing a graph in the Markov equivalence class of $G^\star$ whose strongly connected components correspond to the elements of the partition, and which are partially ordered according to the partial order of the partition. Our algorithm comes in two versions -- one which is provably correct and another one which performs fast in practice.

math.ST

Multivariate Super-Resolution without Separation

In this paper we study the high-dimensional super-resolution imaging problem. Here we are given an image of a number of point sources of light whose locations and intensities are unknown. The image is pixelized and is blurred by a known point-spread function arising from the imaging device. We encode the unknown point sources and their intensities via a nonnegative measure and we propose a convex optimization program to find it. Assuming the device's point-spread function is component-wise decomposable, we show that the optimal solution is the true measure in the noiseless case, and it approximates the true measure well in the noisy case with respect to the generalized Wasserstein distance. Our main assumption is that the components of the point-spread function form a Tchebychev system ($T$-system) in the noiseless case and a $T^*$-system in the noisy case, mild conditions that are satisfied by Gaussian point-spread functions. Our work is a generalization to all dimensions of the work (Eftekhari, Bendory, & Tang, 2021) where the same analysis is carried out in 2 dimensions. We resolve an open problem posed in (Schiebinger, Robeva, & Recht, 2018) in the case when the point-spread function decomposes.

math.OC

Log-concave density estimation in undirected graphical models

We study the problem of maximum likelihood estimation of densities that are log-concave and lie in the graphical model corresponding to a given undirected graph $G$. We show that the maximum likelihood estimate (MLE) is the product of the exponentials of several tent functions, one for each maximal clique of $G$. While the set of log-concave densities in a graphical model is infinite-dimensional, our results imply that the MLE can be found by solving a finite-dimensional convex optimization problem. We provide an implementation and a few examples. Furthermore, we show that the MLE exists and is unique with probability 1 as long as the number of sample points is larger than the size of the largest clique of $G$ when $G$ is chordal. We show that the MLE is consistent when the graph $G$ is a disjoint union of cliques. Finally, we discuss the conditions under which a log-concave density in the graphical model of $G$ has a log-concave factorization according to $G$.

math.ST

Ultra-marginal Feature Importance: Learning from Data with Causal Guarantees

Scientists frequently prioritize learning from data rather than training the best possible model; however, research in machine learning often prioritizes the latter. Marginal contribution feature importance (MCI) was developed to break this trend by providing a useful framework for quantifying the relationships in data. In this work, we aim to improve upon the theoretical properties, performance, and runtime of MCI by introducing ultra-marginal feature importance (UMFI), which uses dependence removal techniques from the AI fairness literature as its foundation. We first propose axioms for feature importance methods that seek to explain the causal and associative relationships in data, and we prove that UMFI satisfies these axioms under basic assumptions. We then show on real and simulated data that UMFI performs better than MCI, especially in the presence of correlated interactions and unrelated features, while partially learning the structure of the causal graph and reducing the exponential runtime of MCI to super-linear.

stat.ML

Third-Order Moment Varieties of Linear Non-Gaussian Graphical Models

In this paper we study linear non-Gaussian graphical models from the perspective of algebraic statistics. These are acyclic causal models in which each variable is a linear combination of its direct causes and independent noise. The underlying directed causal graph can be identified uniquely via the set of second and third order moments of all random vectors that lie in the corresponding model. Our focus is on finding the algebraic relations among these moments for a given graph. We show that when the graph is a polytree these relations form a toric ideal. We construct explicit trek-matrices associated to 2-treks and 3-treks in the graph. Their entries are covariances and third order moments and their 2-minors define our model set-theoretically. Furthermore, we prove that their 2-minors also generate the vanishing ideal of the model. Finally, we describe the polytopes of third order moments and the ideals for models with hidden variables.

math.ST

Robust Eigenvectors of Symmetric Tensors

The tensor power method generalizes the matrix power method to higher order arrays, or tensors. Like in the matrix case, the fixed points of the tensor power method are the eigenvectors of the tensor. While every real symmetric matrix has an eigendecomposition, the vectors generating a symmetric decomposition of a real symmetric tensor are not always eigenvectors of the tensor. In this paper we show that whenever an eigenvector is a generator of the symmetric decomposition of a symmetric tensor, then (if the order of the tensor is sufficiently high) this eigenvector is robust, i.e., it is an attracting fixed point of the tensor power method. We exhibit new classes of symmetric tensors whose symmetric decomposition consists of eigenvectors. Generalizing orthogonally decomposable tensors, we consider equiangular tight frame decomposable and equiangular set decomposable tensors. Our main result implies that such tensors can be decomposed using the tensor power method.

math.NA

The Set of Orthogonal Tensor Trains

In this paper we study the set of tensors that admit a special type of decomposition called an orthogonal tensor train decomposition. Finding equations defining varieties of low-rank tensors is generally a hard problem, however, the set of orthogonally decomposable tensors is defined by appealing quadratic equations. The tensors we consider are an extension of orthogonally decomposable tensors. We show that they are defined by similar quadratic equations, as well as an interesting higher-degree additional equation.

math.AG

Learning Linear Non-Gaussian Graphical Models with Multidirected Edges

In this paper we propose a new method to learn the underlying acyclic mixed graph of a linear non-Gaussian structural equation model given observational data. We build on an algorithm proposed by Wang and Drton, and we show that one can augment the hidden variable structure of the recovered model by learning {\em multidirected edges} rather than only directed and bidirected ones. Multidirected edges appear when more than two of the observed variables have a hidden common cause. We detect the presence of such hidden causes by looking at higher order cumulants and exploiting the multi-trek rule. Our method recovers the correct structure when the underlying graph is a bow-free acyclic mixed graph with potential multi-directed edges.

cs.LG