SearcharxivSearch

arXiv subjects

Anna Seigal

Publications and source records attributed to Anna Seigal.

At least 19 recordsLinked to original sources

Tensor-based second-order causal discovery

Causal discovery seeks to uncover the causal dependencies among variables. For this purpose, we propose an algorithm called Tensor-based Second-order Causal Discovery (TSCD). Its input is a tensor obtained from the covariance matrices of observational and interventional data. Assuming the causal dependencies follow a linear structural equation model on a directed acyclic graph (DAG), TSCD outputs the DAG and the functions on its edges, requiring only that the noise variables are uncorrelated. We also implement a version of the approach for nonlinear models. Our focus on second-order statistics (via the covariance matrices) is motivated by their statistical and computational efficiency relative to higher-order moments, their identifiability relative to first-order statistics, and that they work regardless of whether the variables are Gaussian. We show that TSCD has identifiable causal order and parameters from a number of interventions that is logarithmic in the number of variables. Experiments show that TSCD is robust to noise, competitive with existing methods, and scales to hundreds of variables.

stat.ML

Causal discovery under mean independence and linearity

Causal discovery methods such as LiNGAM identify causal structure from observational data by assuming mutually independent disturbances. This assumption is fragile: shared volatility, common scale effects, or other forms of dependence can cause the methods to recover the wrong causal order, even with infinite data. We introduce the Linear Mean-Independent Acyclic Model (LiMIAM), which replaces full independence with weaker one-sided mean-independence restrictions on the disturbances. Under finite-order consequences of these restrictions, source nodes are generically identifiable, and hence a compatible causal order can be recovered recursively. Our proof is constructive and leads to DirectLiMIAM, a sequential residual-based algorithm for causal discovery under dependent noise. In simulations with mean-independent but dependent disturbances, DirectLiMIAM outperforms LiNGAM methods. A large-scale empirical application to the oil market highlights the implausibility of the independence assumption and the ability of DirectLiMIAM to recover a realistic causal ordering, from policy to production and from prices to inflation.

stat.ME

Multi-context principal component analysis

Principal component analysis (PCA) is a tool to capture factors that explain variation in data. Across domains, data are now collected across multiple contexts (for example, individuals with different diseases, cells of different types, or words across texts). While the factors explaining variation in data are undoubtedly shared across subsets of contexts, no tools currently exist to systematically recover such factors. We develop multi-context principal component analysis (MCPCA), a theoretical and algorithmic framework that decomposes data into factors shared across subsets of contexts. Applied to gene expression, MCPCA reveals axes of variation shared across subsets of cancer types and an axis whose variability in tumor cells, but not mean, is associated with lung cancer progression. Applied to contextualized word embeddings from language models, MCPCA maps stages of a debate on human nature, revealing a discussion between science and fiction over decades. These axes are not found by combining data across contexts or by restricting to individual contexts. MCPCA is a principled generalization of PCA to address the challenge of understanding factors underlying data across contexts.

stat.ML

Multi-subspace power method for decomposing partially symmetric tensors

We present an algorithm for low rank decomposition of tensors of any symmetry type, from fully asymmetric to fully symmetric. It recovers the decomposition one summand at a time via the higher-order power method. This approach is known to fail in general: there need not be a relationship between the summands of a decomposition and the (partially symmetric) singular vector tuples (pSVTs) of the tensor. Our approach overcomes this problem by transforming the input to a tensor with orthonormal slices, via orthogonalization of a flattening. The summands of the decomposition of the original tensor can be recovered from the pSVTs of this new transformed tensor. We introduce a shifted power method for computing pSVTs and prove its global convergence. Numerical experiments demonstrate that our algorithm achieves higher accuracy and faster runtime than existing methods.

math.NA

Beyond independent component analysis: identifiability and algorithms

Independent Component Analysis (ICA) is a classical method for recovering latent variables with useful identifiability properties. For independent variables, cumulant tensors are diagonal; relaxing independence yields tensors whose zero structure generalizes diagonality. These models have been the subject of recent work in non-independent component analysis. We show that pairwise mean independence answers the question of how much one can relax independence: it is identifiable, any weaker notion is non-identifiable, and it contains the models previously studied as special cases. Our results apply to distributions with the required zero pattern at any cumulant tensor. We propose an algebraic recovery algorithm based on least-squares optimization over the orthogonal group. Simulations highlight robustness: enforcing full independence can harm estimation, while pairwise mean independence enables more stable recovery. These findings extend the classical ICA framework and provide a rigorous basis for blind source separation beyond independence.

math.ST

Orthogonal eigenvectors and singular vectors of tensors

The spectral theorem says that a real symmetric matrix has an orthogonal basis of eigenvectors and that, for a matrix with distinct eigenvalues, the basis is unique (up to signs). In this paper, we study the symmetric tensors with an orthogonal basis of eigenvectors and show that, for a generic such tensor, the orthogonal basis is unique. This resolves a conjecture by Mesters and Zwiernik. We also study the non-symmetric setting. The singular value decomposition says that a real matrix has an orthogonal basis of singular vector pairs and that, for a matrix with distinct singular values, the basis is unique (up to signs). We describe the tensors with an orthogonal basis of singular vectors and show that a generic such tensor has a unique orthogonal basis, with one exceptional format: order four binary tensors. We use these results to propose a new tensor decomposition that generalizes an orthogonally decomposable decomposition and specializes the Tucker decomposition.

math.SP

Decomposing tensors via rank-one approximations

Matrices can be decomposed via rank-one approximations: the best rank-one approximation is a singular vector pair, and the singular value decomposition writes a matrix as a sum of singular vector pairs. The singular vector tuples of a tensor are the critical points of its best rank-one approximation problem. In this paper, we study tensors that can be decomposed via successive rank-one approximations: compute a singular vector tuple, subtract it off, compute a singular vector tuple of the new deflated tensor, and repeat. The number of terms in such a decomposition may exceed the tensor rank. Moreover, the decomposition may depend on the order in which terms are subtracted. We show that the decomposition is valid independent of order if and only if all singular vectors in the process are orthogonal in at least two factors. We study the variety of such tensors. We lower bound its dimension, showing that it is significantly larger than the variety of odeco tensors.

math.AG

A Real Generalized Trisecant Trichotomy

The classical trisecant lemma says that a general chord of a non-degenerate space curve is not a trisecant; that is, the chord only meets the curve in two points. The generalized trisecant lemma extends the result to higher-dimensional varieties. It states that the linear space spanned by general points on a projective variety intersects the variety in exactly these points, provided the dimension of the linear space is smaller than the codimension of the variety and that the variety is irreducible, reduced, and non-degenerate. We prove a real analogue of the generalized trisecant lemma, which takes the form of a trichotomy. Along the way, we characterize the possible numbers of real intersection points between a real projective variety and a complimentary dimension real linear space. We show that any integer of correct parity between a minimum and a maximum number can be achieved. We then specialize to Segre-Veronese varieties, where our results apply to the identifiability of independent component analysis, tensor decomposition and to typical tensor ranks.

math.AG

Linear causal disentanglement via higher-order cumulants

Linear causal disentanglement is a recent method in causal representation learning to describe a collection of observed variables via latent variables with causal dependencies between them. It can be viewed as a generalization of both independent component analysis and linear structural equation models. We study the identifiability of linear causal disentanglement, assuming access to data under multiple contexts, each given by an intervention on a latent variable. We show that one perfect intervention on each latent variable is sufficient and in the worst case necessary to recover parameters under perfect interventions, generalizing previous work to allow more latent than observed variables. We give a constructive proof that computes parameters via a coupled tensor decomposition. For soft interventions, we find the equivalence class of latent graphs and parameters that are consistent with observed data, via the study of a system of polynomial equations. Our results hold assuming the existence of non-zero higher-order cumulants, which implies non-Gaussianity of variables.

stat.ML

Contrastive independent component analysis

In recent years, there has been growing interest in jointly analyzing a foreground dataset, representing an experimental group, and a background dataset, representing a control group. The goal of such contrastive investigations is to identify salient features in the experimental group relative to the control. Independent component analysis (ICA) is a powerful tool for learning independent patterns in a dataset. We generalize it to contrastive ICA (cICA). For this purpose, we devise a new linear algebra based tensor decomposition algorithm, which is more expressive but just as efficient and identifiable as other linear algebra based algorithms. We establish the identifiability of cICA and demonstrate its performance in finding patterns and visualizing data, using synthetic, semi-synthetic, and real-world datasets, comparing the approach to existing methods.

math.ST

Identifiability of overcomplete independent component analysis

Independent component analysis (ICA) studies mixtures of independent latent sources. An ICA model is identifiable if the mixing can be recovered uniquely. It is well-known that ICA is identifiable if and only if at most one source is Gaussian. However, this applies only to the setting where the number of sources is at most the number of observations. In this paper, we generalize the identifiability of ICA to the overcomplete setting, where the number of sources exceeds the number of observations. We give an if and only if characterization of the identifiability of overcomplete ICA. The proof studies linear spaces of rank one symmetric matrices. For generic mixing, we present an identifiability condition in terms of the number of sources and the number of observations. We use our identifiability results to design an algorithm to recover the mixing matrix from data and apply it to synthetic data and two real datasets.

math.ST

Complete collineations for maximum likelihood estimation

We import the algebro-geometric notion of a complete collineation into the study of maximum likelihood estimation in directed Gaussian graphical models. A complete collineation produces a perturbation of sample data, which we call a stabilisation of the sample. While a maximum likelihood estimate (MLE) may not exist or be unique given sample data, it is always unique given a stabilisation. We relate the MLE given a stabilisation to the MLE given original sample data, when one exists, providing necessary and sufficient conditions for the MLE given a stabilisation to be one given the original sample. For linear regression models, we show that the MLE given any stabilisation is the minimal norm choice among the MLEs given an original sample. We show that the MLE has a well-defined limit as the stabilisation of a sample tends to the original sample, and that the limit is an MLE given the original sample, when one exists. Finally, we study which MLEs given a sample can arise as such limits. We reduce this to a question regarding the non-emptiness of certain algebraic varieties.

math.ST

Linear Causal Disentanglement via Interventions

Causal disentanglement seeks a representation of data involving latent variables that relate to one another via a causal model. A representation is identifiable if both the latent model and the transformation from latent to observed variables are unique. In this paper, we study observed variables that are a linear transformation of a linear latent causal model. Data from interventions are necessary for identifiability: if one latent variable is missing an intervention, we show that there exist distinct models that cannot be distinguished. Conversely, we show that a single intervention on each latent variable is sufficient for identifiability. Our proof uses a generalization of the RQ decomposition of a matrix that replaces the usual orthogonal and upper triangular conditions with analogues depending on a partial order on the rows of the matrix, with partial order determined by a latent causal model. We corroborate our theoretical results with a method for causal disentanglement that accurately recovers a latent causal model.

stat.ML

Rectifiable paths with polynomial log-signature are straight lines

The signature of a rectifiable path is a tensor series in the tensor algebra whose coefficients are definite iterated integrals of the path. The signature characterises the path up to a generalised form of reparametrisation. It is a classical result of K. T. Chen that the log-signature (the logarithm of the signature) is a Lie series. A Lie series is polynomial if it has finite degree. We show that the log-signature is polynomial if and only if the path is a straight line up to reparametrisation. Consequently, the log-signature of a rectifiable path either has degree one or infinite support. Though our result pertains to rectifiable paths, the proof uses results from rough path theory, in particular that the signature characterises a rough path up to reparametrisation.

math.RA

Supermodular Rank: Set Function Decomposition and Optimization

We define the supermodular rank of a function on a lattice. This is the smallest number of terms needed to decompose it into a sum of supermodular functions. The supermodular summands are defined with respect to different partial orders. We characterize the maximum possible value of the supermodular rank and describe the functions with fixed supermodular rank. We analogously define the submodular rank. We use submodular decompositions to optimize set functions. Given a bound on the submodular rank of a set function, we formulate an algorithm that splits an optimization problem into submodular subproblems. We show that this method improves the approximation ratio guarantees of several algorithms for monotone set function maximization and ratio of set functions minimization, at a computation overhead that depends on the submodular rank.

math.CO

Multilinear Hyperquiver Representations

We count singular vector tuples of a system of tensors assigned to the edges of a directed hypergraph. To do so, we study the generalisation of quivers to directed hypergraphs. Assigning vector spaces to the nodes of a hypergraph and multilinear maps to its hyperedges gives a hyperquiver representation. Hyperquiver representations generalise quiver representations (where all hyperedges are edges) and tensors (where there is only one multilinear map). The singular vectors of a hyperquiver representation are a compatible assignment of vectors to the nodes. We compute the dimension and degree of the variety of singular vectors of a sufficiently generic hyperquiver representation. Our formula specialises to known results that count the singular vectors and eigenvectors of a generic tensor. Lastly, we study a hypergraph generalisation of the inverse tensor eigenvalue problem and solve it algorithmically.

math.AG

Differential Equations for Gaussian Statistical Models with Rational Maximum Likelihood Estimator

We study multivariate Gaussian statistical models whose maximum likelihood estimator (MLE) is a rational function of the observed data. We establish a one-to-one correspondence between such models and the solutions to a nonlinear first-order partial differential equation (PDE). Using our correspondence, we reinterpret familiar classes of models with rational MLE, such as directed (and decomposable undirected) Gaussian graphical models. We also find new models with rational MLE. For linear concentration models with rational MLE, we show that homaloidal polynomials from birational geometry lead to solutions to the PDE. We thus shed light on the problem of classifying Gaussian models with rational MLE by relating it to the open problem in birational geometry of classifying homaloidal polynomials.

math.AG

Lower bounds on the rank and symmetric rank of real tensors

We lower bound the rank of a tensor by a linear combination of the ranks of three of its unfoldings, using Sylvester's rank inequality. In a similar way, we lower bound the symmetric rank by a linear combination of the symmetric ranks of three unfoldings. Lower bounds on the rank and symmetric rank of tensors are important for finding counterexamples to Comon's conjecture. A real counterexample to Comon's conjecture is a tensor whose real rank and real symmetric rank differ. Previously, only one real counterexample was known. We divide the construction into three steps. The first step involves linear spaces of binary tensors. The second step considers a linear space of larger decomposable tensors. The third step is to verify a conjecture that lower bounds the symmetric rank, on a tensor of interest. We use the construction to build an order six real tensor whose real rank and real symmetric rank differ.

math.AG