SearcharxivSearch

arXiv subjects

Elizabeth Gross

Publications and source records attributed to Elizabeth Gross.

At least 19 recordsLinked to original sources

Semialgebraic Conditions for Identifying Triangles in Phylogenetic Networks

An important consideration for a model-based method of phylogenetic network inference is the identifiability of the network parameter of the model. A recurring theme in previous works exploring this issue is that it is often difficult to identify the orientation of edges in a triangle of the network. In fact, it has been shown that for some models it is impossible to determine the orientation of triangle edges utilizing the standard algebraic technique of phylogenetic invariants. In this work, we consider one such model with a Jukes-Cantor site-substitution process and no coalescence. We give a complete semialgebraic description of three, 3-leaf Jukes-Cantor phylogenetic network models with embedded triangles. By describing these base cases, we resolve several questions about the identifiability of networks with embedded triangles. We show that for any pair of models, the intersection and set differences of the models are full-dimensional regions of the space of site-pattern probability distributions. Thus, despite being algebraically indistinguishable, these network models are not identical, nor are they identifiable (or generically identifiable). Our results also yield a straightforward biological interpretation--that the signal from a hybridization event may be immediately detectable but decays over time until it is impossible to identify the orientation of edges in the triangle of a network.

q-bio.PE

Singular Learning Theory for Factor Analysis

Watanabe's singular learning theory provides a framework for asymptotic analysis of Bayesian model selection for statistical models with singularities, where traditional statistical regularity assumptions fail. Learning coefficients, also known as real log canonical thresholds, play a central role in singular learning, as they govern the asymptotic behavior of Bayesian marginal likelihood integrals in settings where the Laplace approximations used for regular statistical models are not applicable. Learning coefficients are algebraic invariants that quantify the geometric complexity of a model and reveal how the singular structure impacts the model's generalization properties. In this paper, we apply algebraic methods to study the learning coefficients of factor analysis models, which are widely used latent variable models for continuously distributed data. Our main result provides exact formulas for learning coefficients of factor analysis models. Moreover, we study the singularity types of specific factor analysis models in detail.

math.ST

Methodological considerations for semialgebraic hypothesis testing with incomplete U-statistics

Recently, Sturma, Drton, and Leung proposed a general-purpose stochastic method for hypothesis testing in models defined by polynomial equality and inequality constraints. Notably, the method remains theoretically valid even near irregular points, such as singularities and boundaries, where traditional testing approaches often break down. In this paper, we evaluate its practical performance on a collection of biologically motivated models from phylogenetics. While the method performs remarkably well across different settings, we catalogue a number of issues that should be considered for effective application.

q-bio.PE

Routing functions for parameter space decomposition to describe stability landscapes of ecological models

Changes in environmental or system parameters often drive major biological transitions, including ecosystem collapse, disease outbreaks, and tumor development. Analyzing the stability of steady states in dynamical systems provides critical insight into these transitions. This paper introduces an algebraic framework for analyzing the stability landscapes of ecological models defined by systems of first-order autonomous ordinary differential equations with polynomial or rational rate functions. Using tools from real algebraic geometry, we characterize parameter regions associated with steady-state feasibility and stability via three key boundaries: singular, stability (Routh-Hurwitz), and coordinate boundaries. With these boundaries in mind, we employ routing functions to compute the connected components of parameter space in which the number and type of stable steady states remain constant, revealing the stability landscape of these ecological models. As case studies, we revisit the classical Levins-Culver competition-colonization model and a recent model of coral-bacteria symbioses. In the latter, our method uncovers complex stability regimes, including regions supporting limit cycles, that are inaccessible via traditional techniques. These results demonstrate the potential of our approach to inform ecological theory and intervention strategies in systems with nonlinear interactions and multiple stable states.

q-bio.PE

Maximum likelihood degree of the $\beta$-stochastic blockmodel

Log-linear exponential random graph models are a specific class of statistical network models that have a log-linear representation. This class includes many stochastic blockmodel variants. In this paper, we focus on $\beta$-stochastic blockmodels, which combine the $\beta$-model with a stochastic blockmodel. Here, using recent results by Almendra-Hern\'{a}ndez, De Loera, and Petrovi\'{c}, which describe a Markov basis for $\beta$-stochastic block model, we give a closed form formula for the maximum likelihood degree of a $\beta$-stochastic blockmodel. The maximum likelihood degree is the number of complex solutions to the likelihood equations. In the case of the $\beta$-stochastic blockmodel, the maximum likelihood degree factors into a product of Eulerian numbers.

math.ST

Group-Based Phylogenetic Models on 3-Sunlet Networks

Phylogenetic networks describe the evolution of a set of taxa for which reticulate events have occurred at some point in their evolutionary history. Of particular interest is when the evolutionary history between a set of just three taxa has a reticulate event. In molecular phylogenetics, substitution models can model the process of evolution at the genetic level, and the case of three taxa with a reticulate event can be modelled using a substitution model on a mixed graph called a 3-sunlet. We investigate a class of substitution models called group-based phylogenetic models on 3-sunlet networks. In particular, we investigate the discrete geometry of the parameter space and how this relates to the dimension of the phylogenetic variety associated to the model. This enables us to give a dimension formula for this variety for general group-based models when the order of the group is odd.

q-bio.PE

New directions in algebraic statistics: Three challenges from 2023

In the last quarter of a century, algebraic statistics has established itself as an expanding field which uses multilinear algebra, commutative algebra, computational algebra, geometry, and combinatorics to tackle problems in mathematical statistics. These developments have found applications in a growing number of areas, including biology, neuroscience, economics, and social sciences. Naturally, new connections continue to be made with other areas of mathematics and statistics. This paper outlines three such connections: to statistical models used in educational testing, to a classification problem for a family of nonparametric regression models, and to phase transition phenomena under uniform sampling of contingency tables. We illustrate the motivating problems, each of which is for algebraic statistics a new direction, and demonstrate an enhancement of related methodologies.

math.ST

Absolute concentration robustness: Algebra and geometry

Motivated by the question of how biological systems maintain homeostasis in changing environments, Shinar and Feinberg introduced in 2010 the concept of absolute concentration robustness (ACR). A biochemical system exhibits ACR in some species if the steady-state value of that species does not depend on initial conditions. Thus, a system with ACR can maintain a constant level of one species even as the environment changes. Despite a great deal of interest in ACR in recent years, the following basic question remains open: How can we determine quickly whether a given biochemical system has ACR? Although various approaches to this problem have been proposed, we show that they are incomplete. Accordingly, we present new methods for deciding ACR, which harness computational algebra. We illustrate our results on several biochemical signaling networks.

math.AG

The Pfaffian Structure of CFN Phylogenetic Networks

Algebraic techniques in phylogenetics have historically been successful at proving identifiability results and have also led to novel reconstruction algorithms. In this paper, we study the ideal of phylogenetic invariants of the Cavender-Farris-Neyman (CFN) model on a phylogenetic network with the goal of providing a description of the invariants which is useful for network inference. It was previously shown that to characterize the invariants of any level-1 network, it suffices to understand all sunlet networks, which are those consisting of a single cycle with a leaf adjacent to each cycle vertex. We show that the parameterization of an affine open patch of the CFN sunlet model, which intersects the probability simplex, factors through the space of skew-symmetric matrices via Pfaffians. We then show that this affine patch is isomorphic to a determinantal variety and give an explicit Gr{\"o}bner basis for the associated ideal, which involves only $\binom{n}{2}$ coordinates rather than $2^{n}$. Lastly, we show that sunlet networks with at least 6 leaves are identifiable using only these polynomials and run extensive simulations, which show that these polynomials can be used to accurately infer the correct network from DNA sequence data.

math.AG

Dimensions of Level-1 Group-Based Phylogenetic Networks

Phylogenetic networks represent evolutionary histories of sets of taxa where horizontal evolution or hybridization has occurred. Placing a Markov model of evolution on a phylogenetic network gives a model that is particularly amenable to algebraic study by representing it as an algebraic variety. In this paper, we give a formula for the dimension of the variety corresponding to a triangle-free level-1 phylogenetic network under a group-based evolutionary model. On our way to this, we give a dimension formula for codimension zero toric fiber products. We conclude by illustrating applications to identifiability.

q-bio.PE

Mixed volumes of networks with binomial steady-states

The steady-state degree of a chemical reaction network is the number of complex steady-states for generic rate constants and initial conditions. One way to bound the steady-state degree is through the mixed volume of the steady-state system or an equivalent system. In this work, we show that for partionable binomial networks, whose resulting steady-state systems are given by a set of binomials and a set of linear (not necessarily binomial) conservation equations, computing the mixed volume is equivalent to finding the volume of a single mixed cell that is the translate of a parallelotope. We then turn our attention to identifying cycles with binomial steady-state ideals. To this end, we give a coloring condition on directed cycles that guarantees the network has a binomial steady-state ideal. We highlight both of these theorems using a class of networks referred to as species-overlapping networks and give a formula for the mixed volume of these networks.

math.CO

Statistical learning with phylogenetic network invariants

Phylogenetic networks provide a means of describing the evolutionary history of sets of species believed to have undergone hybridization or gene flow during their evolution. The mutation process for a set of such species can be modeled as a Markov process on a phylogenetic network. Previous work has shown that a site-pattern probability distributions from a Jukes-Cantor phylogenetic network model must satisfy certain algebraic invariants. As a corollary, aspects of the phylogenetic network are theoretically identifiable from site-pattern frequencies. In practice, because of the probabilistic nature of sequence evolution, the phylogenetic network invariants will rarely be satisfied, even for data generated under the model. Thus, using network invariants for inferring phylogenetic networks requires some means of interpreting the residuals, or deviations from zero, when observed site-pattern frequencies are substituted into the invariants. In this work, we propose a method of utilizing invariant residuals and support vector machines to infer 4-leaf level-one phylogenetic networks, from which larger networks can be reconstructed. Given data for a set of species, the support vector machine is first trained on model data to learn the patterns of residuals corresponding to different network structures to classify the network that produced the data. We demonstrate the performance of our method on simulated data from the specified model and primate data.

q-bio.PE

One-connection rule for structural equation models

Linear structural equation models are multivariate statistical models encoded by mixed graphs. In particular, the set of covariance matrices for distributions belonging to a linear structural equation model for a fixed mixed graph $G=(V, D,B)$ is parameterized by a rational function with parameters for each vertex and edge in $G$. This rational parametrization naturally allows for the study of these models from an algebraic and combinatorial point of view. Indeed, this point of view has led to a collection of results in the literature, mainly focusing on questions related to identifiability and determining relationships between covariances (i.e., finding polynomials in the Gaussian vanishing ideal). So far, a large proportion of these results has focused on the case when $D$, the directed part of the mixed graph $G$, is acyclic. This is due to the fact that in the acyclic case, the parametrization becomes polynomial and there is a description of the entries of the covariance matrices in terms of a finite sum. We move beyond the acyclic case and give a closed form expression for the entries of the covariance matrices in terms of the one-connections in a graph obtained from $D$ through some small operations. This closed form expression then allows us to show that if $G$ is simple, then the parametrization map is generically finite-to-one. Finally, having a closed form expression for the covariance matrices allows for the development of an algorithm for systematically exploring possible polynomials in the Gaussian vanishing ideal.

math.ST

Broken Bracelets and Kostant's Partition Function

Inspired by the work of Amdeberhan, Can, and Moll on broken necklaces, we define a broken bracelet as a linear arrangement of marked and unmarked vertices and introduce a generalization called $n$-stars, which is a collection of $n$ broken bracelets whose final (unmarked) vertices are identified. Through these combinatorial objects, we provide a new framework for the study of Kostant's partition function, which counts the number of ways to express a vector as a nonnegative integer linear combination of the positive roots of a Lie algebra. Our main result establishes that (up to reflection) the number of broken bracelets with a fixed number of unmarked vertices with nonconsecutive marked vertices gives an upper bound for the value of Kostant's partition function for multiples of the highest root of a Lie algebra of type $A$. We connect this work to multiplex juggling sequences, as studied by Benedetti, Hanusa, Harris, Morales, and Simpson, by providing a correspondence to an equivalence relation on $n$-stars.

math.CO

Identifiability of linear compartmental tree models and a general formula for input-output equations

A foundational question in the theory of linear compartmental models is how to assess whether a model is structurally identifiable -- that is, whether parameter values can be inferred from noiseless data -- directly from the combinatorics of the model. Our main result completely answers this question for models (with one input and one output) in which the underlying graph is a bidirectional tree; moreover, identifiability of such models can be verified visually}. Models of this structure include two families of models often appearing in biological applications: catenary and mammillary models. Our analysis of such models is enabled by two supporting results, which are significant in their own right. One result gives the first general formula for the coefficients of input-output equations (certain equations that can be used to determine identifiability) that allows for input and output to be in distinct compartments}. In another supporting result, we prove that identifiability is preserved when a model is enlarged and altered in specific ways involving adding a new compartment with a bidirected edge to an existing compartment.

math.DS

What are higher-order networks?

Network-based modeling of complex systems and data using the language of graphs has become an essential topic across a range of different disciplines. Arguably, this graph-based perspective derives its success from the relative simplicity of graphs: A graph consists of nothing more than a set of vertices and a set of edges, describing relationships between pairs of such vertices. This simple combinatorial structure makes graphs interpretable and flexible modeling tools. The simplicity of graphs as system models, however, has been scrutinized in the literature recently. Specifically, it has been argued from a variety of different angles that there is a need for higher-order networks, which go beyond the paradigm of modeling pairwise relationships, as encapsulated by graphs. In this survey article we take stock of these recent developments. Our goals are to clarify (i) what higher-order networks are, (ii) why these are interesting objects of study, and (iii) how they can be used in applications.

cs.SI

Goodness of fit for log-linear ERGMs

Many popular models from the networks literature can be viewed through a common lens of contingency tables on network dyads, resulting in \emph{log-linear ERGMs}: exponential family models for random graphs whose sufficient statistics are linear on the dyads. We propose a new model in this family, the \emph{$p_1$-SBM}, which combines node and group effects common in network formation mechanisms. In particular, it is a generalization of several well-known ERGMs including the stochastic blockmodel for undirected graphs with known block assignment, the degree-corrected version of it, and the directed $p_1$ model without group structure. We frame the problem of testing model fit for the log-linear ERGM class through an exact conditional test whose $p$-value can be approximated efficiently in networks of both small and moderately large sizes. The sampling methods we build rely on a dynamic adaptation of Markov bases. We use quick estimation algorithms adapted from the contingency table literature and effective sampling methods rooted in graph theory and algebraic statistics. The performance and scalability of the method is demonstrated on two data sets from biology: the connectome of \emph{C. elegans} and the interactome of \emph{Arabidopsis thaliana}. These two networks -- a network and a protein-protein interaction network -- have been popular examples in the network science literature. Our work provides a model-based approach to studying them.

stat.ME

When do two networks have the same steady-state ideal?

Chemical reaction networks are often used to model and understand biological processes such as cell signaling. Under the framework of chemical reaction network theory, a process is modeled with a directed graph and a choice of kinetics, which together give rise to a dynamical system. Under the assumption of mass action kinetics, the dynamical system is polynomial. In this paper, we consider the ideals generated by the these polynomials, which are called steady-state ideals. Steady-state ideals appear in multiple contexts within the chemical reaction network literature, however they have yet to be systematically studied. To begin such a study, we ask and partially answer the following question: when do two reaction networks give rise to the same steady-state ideal? In particular, our main results describe three operations on the reaction graph that preserve the steady-state ideal. Furthermore, since the motivation for this work is the classification of steady-state ideals, monomials play a primary role. To this end, combinatorial conditions are given to identify monomials in a steady-state ideal, and we give a sufficient condition for a steady-state ideal to be monomial.

math.CO