SearcharxivSearch

arXiv subjects

Benjamin Hollering

Publications and source records attributed to Benjamin Hollering.

At least 19 recordsLinked to original sources

Log Canonical Models and Positive Geometries

Constructing log canonical compactifications of open varieties is a central problem in birational geometry. Finding a natural coordinate system and obtaining the equations of these models is difficult in general. We show that for a large class of varieties explicit coordinates for the log canonical model are provided by canonical forms of positive geometries, and use this to compute the equations of these models. Our theory applies, for instance, to complements of hyperplane arrangements, cubic surfaces with lines removed, and the moduli space of marked cubic del Pezzo surfaces.

math.AG

Efficient Symbolic Computations for Identifying Causal Effects

Determining identifiability of causal effects from observational data under latent confounding is a central challenge in causal inference. For linear structural causal models, identifiability of causal effects is decidable through symbolic computation. However, standard approaches based on Gr\"obner bases become computationally infeasible beyond small settings due to their doubly exponential complexity. In this work, we study how to practically use symbolic computation for deciding rational identifiability. In particular, we present an efficient algorithm that provably finds the lowest degree identifying formulas. For a causal effect of interest, if there exists an identification formula of a prespecified maximal degree, our algorithm returns such a formula in quasi-polynomial time.

stat.ML

Landau Analysis in the Grassmannian

Momentum twistors for scattering amplitudes in particle physics are lines in three-space. We develop Landau analysis for Feynman integrals in this setting. The resulting discriminants and resultants are identified with Hurwitz and Chow forms of incidence varieties in products of Grassmannians. We study their degrees and factorizations, and the kinematic regimes in which the fibers of the Landau map are rational or real. Identifying this map with the amplituhedron map on positroid varieties, and the associated recursions with promotion maps, yields a geometric mechanism for the emergence of positivity and cluster structures in planar N=4 super Yang-Mills theory.

math.AG

Positivity and Cluster Structures in Landau Analysis

Landau analysis in momentum twistor space can be formulated as the study of varieties of lines in three-dimensional projective space, together with their projections and discriminants. Within this framework, we define enumerative invariants (LS degrees) that count leading singularities. Leading Landau singularities (LS discriminants) arise as discriminants detecting the collision of leading singularities. We uncover a recursive mechanism underlying Landau singularities, governed by substitution maps between Grassmannians. Applying this framework, we prove positivity and factorization into cluster variables for the LS discriminant of a large class of Landau diagrams at arbitrary loop order. This provides a first-principles explanation for the emergence of positivity and cluster algebra structures in the singularities of planar N=4 super Yang-Mills theory.

hep-th

Algebraic Statistics in OSCAR

We introduce the AlgebraicStatistics section of the OSCAR computer algebra system. We give an overview of its extensible design and highlight its features including serialization of data types for sharing results and creating databases, and state-of-the-art implicitization algorithms.

stat.CO

Varieties of Lines in 3-Space

We consider configurations of lines in 3-space with incidences prescribed by a graph. This defines a subvariety in a product of Grassmannians. Leveraging a connection with rigidity theory in the plane, for any graph, we determine the dimension of the incidence variety and characterize when it is irreducible or a complete intersection. We study its multidegree and the family of Schubert problems it encodes. Our spanning-tree coordinates enable efficient symbolic computations. We also provide numerical irreducible decompositions for incidence varieties with up to eight lines. These constructions with lines play a key role in the Landau analysis of scattering amplitudes in particle physics.

math.CO

Structural Identifiability of Graphical Continuous Lyapunov Models

We prove two characterizations of model equivalence of acyclic graphical continuous Lyapunov models (GCLMs) with uncorrelated noise. The first result shows that two graphs are model equivalent if and only if they have the same skeleton and equivalent induced 4-node subgraphs. We also give a transformational characterization via structured edge reversals. The two theorems are Lyapunov analogues of celebrated results for Bayesian networks by Verma and Pearl, and Chickering, respectively. Our results have broad consequences for the theory of causal inference of GCLMs. First, we find that model equivalence classes of acyclic GCLMs refine the corresponding classes of Bayesian networks. Furthermore, we obtain polynomial-time algorithms to test model equivalence and structural identifiability of given directed acyclic graphs.

math.ST

Parke-Taylor varieties

Parke-Taylor functions are certain rational functions on the Grassmannian of lines encoding MHV amplitudes in particle physics. For $n$ particles there are $n!$ Parke-Taylor functions, corresponding to all orderings of the particles. Linear relations between these functions have been extensively studied in the last years. We here describe all non-linear polynomial relations between these functions in a simple combinatorial way and study the variety parametrized by them, called the Parke-Taylor variety. We show that the Parke-Taylor variety is linearly isomorphic to the log canonical embedding of the moduli space $\overline{\mathcal{M}}_{0,n}$ due to Keel and Tevelev, and that the intersection with the algebraic torus recovers the open part, $\mathcal{M}_{0,n}$. We give an explicit description of this isomorphism. Unlike the log canonical embedding, this Parke-Taylor embedding respects the symmetry of the $n$ marked points and is constructed in a single-step procedure, avoiding the intermediate embedding into a product of projective spaces.

math.AG

A PC Algorithm for Max-Linear Bayesian Networks

Max-linear Bayesian networks (MLBNs) are a relatively recent class of structural equation models which arise when the random variables involved have heavy-tailed distributions. Unlike most directed graphical models, MLBNs are typically not faithful to d-separation and thus classical causal discovery algorithms such as the PC algorithm or greedy equivalence search can not be used to accurately recover the true graph structure. In this paper, we begin the study of constraint-based discovery algorithms for MLBNs given an oracle for testing conditional independence in the true, unknown graph. We show that if the oracle is given by the $\ast$-separation criteria in the true graph, then the PC algorithm remains consistent despite the presence of additional CI statements implied by $\ast$-separation. We also introduce a new causal discovery algorithm named "PCstar" which assumes faithfulness to $C^\ast$-separation and is able to orient additional edges which cannot be oriented with only d- or $\ast$-separation.

stat.ML

Polyhedral Aspects of Maxoids

The conditional independence (CI) relation of a distribution in a max-linear Bayesian network depends on its weight matrix through the $C^\ast$-separation criterion. These CI~models, which we call maxoids, are compositional graphoids which are in general not representable by Gaussian random variables. We prove that every maxoid can be obtained from a transitively closed weighted DAG and show that the stratification of generic weight matrices by their maxoids yields a polyhedral~fan. We also use this connection to polyhedral geometry to develop an algorithm for solving the conditional independence implication problem for maxoids.

math.CO

Conditional Independence in Stationary Diffusions

Stationary distributions of multivariate diffusion processes have recently been proposed as probabilistic models of causal systems in statistics and machine learning. Motivated by these developments, we study stationary multivariate diffusion processes with a sparsely structured drift. Our main result gives a characterization of the conditional independence relations that hold in a stationary distribution. The result draws on a graphical representation of the drift structure and pertains to conditional independence relations that hold generally as a consequence of the drift's sparsity pattern.

math.ST

Hyperplane Representations of Interventional Characteristic Imset Polytopes

Characteristic imsets are 0/1-vectors representing directed acyclic graphs whose edges represent direct cause-effect relations between jointly distributed random variables. A characteristic imset (CIM) polytope is the convex hull of a collection of characteristic imsets. CIM polytopes arise as feasible regions of a linear programming approach to the problem of causal disovery, which aims to infer a cause-effect structure from data. Linear optimization methods typically require a hyperplane representation of the feasible region, which has proven difficult to compute for CIM polytopes despite continued efforts. We solve this problem for CIM polytopes that are the convex hull of imsets associated to DAGs whose underlying graph of adjacencies is a tree. Our methods use the theory of toric fiber products as well as the novel notion of interventional CIM polytopes. Our solution is obtained as a corollary of a more general result for interventional CIM polytopes. The identified hyperplanes are applied to yield a linear optimization-based causal discovery algorithm for learning polytree causal networks from a combination of observational and interventional data.

math.CO

Faithlessness in Gaussian graphical models

The implication problem for conditional independence (CI) asks whether the fact that a probability distribution obeys a given finite set of CI relations implies that a further CI statement also holds in this distribution. This problem has a long and fascinating history, cumulating in positive results about implications now known as the semigraphoid axioms as well as impossibility results about a general finite characterization of CI implications. Motivated by violation of faithfulness assumptions in causal discovery, we study the implication problem in the special setting where the CI relations are obtained from a directed acyclic graphical (DAG) model along with one additional CI statement. Focusing on the Gaussian case, we give a complete characterization of when such an implication is graphical by using algebraic techniques. Moreover, prompted by the relevance of strong faithfulness in statistical guarantees for causal discovery algorithms, we give a graphical solution for an approximate CI implication problem, in which we ask whether small values of one additional partial correlation entail small values for yet a further partial correlation.

math.ST

The Pfaffian Structure of CFN Phylogenetic Networks

Algebraic techniques in phylogenetics have historically been successful at proving identifiability results and have also led to novel reconstruction algorithms. In this paper, we study the ideal of phylogenetic invariants of the Cavender-Farris-Neyman (CFN) model on a phylogenetic network with the goal of providing a description of the invariants which is useful for network inference. It was previously shown that to characterize the invariants of any level-1 network, it suffices to understand all sunlet networks, which are those consisting of a single cycle with a leaf adjacent to each cycle vertex. We show that the parameterization of an affine open patch of the CFN sunlet model, which intersects the probability simplex, factors through the space of skew-symmetric matrices via Pfaffians. We then show that this affine patch is isomorphic to a determinantal variety and give an explicit Gr{\"o}bner basis for the associated ideal, which involves only $\binom{n}{2}$ coordinates rather than $2^{n}$. Lastly, we show that sunlet networks with at least 6 leaves are identifiable using only these polynomials and run extensive simulations, which show that these polynomials can be used to accurately infer the correct network from DNA sequence data.

math.AG

Computing Implicitizations of Multi-Graded Polynomial Maps

In this paper, we focus on computing the kernel of a map of polynomial rings $\varphi$. This core problem in symbolic computation is known as implicitization. While there are extremely effective Gr\"obner basis methods used to solve this problem, these methods can become infeasible as the number of variables increases. In the case when the map $\varphi$ is multigraded, we consider an alternative approach. We demonstrate how to quickly compute a matrix of maximal rank for which $\varphi$ has a positive multigrading. Then in each graded component we compute the minimal generators of the kernel in that multidegree with linear algebra. We have implemented our techniques in Macaulay2 and show that our implementation can compute many generators of low degree in examples where Gr\"obner techniques have failed. This includes several examples coming from phylogenetics where even a complete list of quadrics and cubics were unknown. When the multigrading refines total degree, our algorithm is \emph{embarassingly parallel} and a fully parallelized version of our algorithm will be forthcoming in OSCAR.

math.AG

Identifiability of Homoscedastic Linear Structural Equation Models using Algebraic Matroids

We consider structural equation models (SEMs), in which every variable is a function of a subset of the other variables and a stochastic error. Each such SEM is naturally associated with a directed graph describing the relationships between variables. When the errors are homoscedastic, recent work has proposed methods for inferring the graph from observational data under the assumption that the graph is acyclic (i.e., the SEM is recursive). In this work, we study the setting of homoscedastic errors but allow the graph to be cyclic (i.e., the SEM to be non-recursive). Using an algebraic approach that compares matroids derived from the parameterizations of the models, we derive sufficient conditions for when two simple directed graphs generate different distributions generically. Based on these conditions, we exhibit subclasses of graphs that allow for directed cycles, yet are generically identifiable. We also conjecture a strengthening of our graphical criterion which can be used to distinguish many more non-complete graphs.

math.CO

Identifiability of the Rooted Tree Parameter under the Cavender-Farris-Neyman Model with a Molecular Clock

Identifiability of the discrete tree parameter is a key property for phylogenetic models since it is necessary for statistically consistent estimation of the tree from sequence data. Algebraic methods have proven to be very effective at showing that tree and network parameters of phylogenetic models are identifiable, especially when the underlying models are group-based. However, since group-based models are time-reversible, only the unrooted tree topology is identifiable and the location of the root is not. In this note we show that the rooted tree parameter of the Cavender-Farris-Neyman Model with a Molecular Clock is generically identifiable by using the invariants of the model which were characterized by Coons and Sullivant.

q-bio.PE

Toric Fiber Products in Geometric Modeling

An important challenge in Geometric Modeling is to classify polytopes with rational linear precision. Equivalently, in Algebraic Statistics one is interested in classifying scaled toric varieties, also known as discrete exponential families, for which the maximum likelihood estimator can be written in closed form as a rational function of the data (rational MLE). The toric fiber product (TFP) of statistical models is an operation to iteratively construct new models with rational MLE from lower dimensional ones. In this paper we introduce TFPs to the Geometric Modeling setting to construct polytopes with rational linear precision and give explicit formulae for their blending functions. A special case of the TFP is taking the Cartesian product of two polytopes and their blending functions. The Horn matrix of a statistical model with rational MLE is a key player in both Geometric Modeling and Algebraic Statistics; it proved to be fruitful providing a characterisation of those polytopes having the more restrictive property of strict linear precision. We give an explicit description of the Horn matrix of a TFP.

math.AG