SearcharxivSearch

arXiv subjects

Bryson Kagy

Publications and source records attributed to Bryson Kagy.

9 recordsLinked to original sources

Identifiability of phylogenetic networks and quintet concordance factors

Several statistical methods of phylogenetic network inference and testing for non-tree-like relationships are based on assessing genomic data through quartet Concordance Factors, the frequencies of 4-taxon topological relationships on gene trees. While such an approach obviates making several undesirable modeling assumptions, it also results in non-identifiability issues for network roots and for small cycles. In this work, an algorithm and accompanying Macaulay2 implementation are provided for computing $n$-tet Concordance Factors on any phylogenetic network. We employ this algorithm on quintet Concordance Factors, summarizing 5-taxon gene trees, to explore identifiability of level-1 networks under the Network Multispecies Coalescent model. We show some additional network features become identifiable that are not through quartets. As identifiability is a necessary prerequisite to inference by any method, this lays a foundation for future inference work.

q-bio.PE

Common Real Secants to Pairs of Real Twisted Cubic Curves

It is well established that a general pair of twisted cubic curves in complex projective space has ten common secant lines. As an initial investigation, we show that the monodromy group of the ten common secant lines over the complex numbers is the full symmetric group demonstrating that the common secant lines have no special structure over the complex numbers. We then investigate a novel question in real algebraic geometry: describe the possible collections of ten common secant lines to a pair of real projective twisted cubic curves. In addition to distinguishing between real and nonreal secant lines, we introduce a refinement of this classification which takes intersection points into account yielding totally real, partially real, and minimally real secant lines. Using computational algebraic geometry as well as combinatorics, we show that for each $k$ between 0 and 10, there exist pairs of real twisted cubic curves with exactly $k$ common totally real secant lines. We also obtain examples of real twisted cubics whose sets of common real secants cover a wide range of possibilities within our admissible classification of common real secant lines.

math.AG

The poset of maximal tubings of the cycle graph is a lattice

The poset of maximal tubings of a graph generalizes several well-known and remarkable partial orders. Notable examples include the weak Bruhat order and the Tamari lattice, posets of maximal tubings for the complete graph and the path graph, respectively. It is an open problem to characterize graphs for which the poset of maximal tubings is a lattice. In this paper, we prove that the poset of maximal tubings for the cycle graph is a lattice, and moreover that it is semidistributive and congruence uniform. As main tools, we characterize all order relations in the poset, and introduce a useful map from maximal tubings of the cycle graph to maximal tubings of the path graph.

math.CO

Identifiability of Large Phylogenetic Mixtures for Many Phylogenetic Model Structures

Identifiability of phylogenetic models is a necessary condition to ensure that the model parameters can be uniquely determined from data. Mixture models are phylogenetic models where the probability distributions in the model are convex combinations of distributions in simpler phylogenetic models. Mixture models are used to model heterogeneity in the substitution process in DNA sequences. While many basic phylogenetic models are known to be identifiable, mixture models in generality have only been shown to be identifiable in certain cases. We expand the main theorem of [Rhodes, Sullivant 2012] to prove identifiability of mixture models in equivariant phylogenetic models, specifically the Jukes-Cantor, Kimura 2-parameter model, Kimura 3-parameter model and the Strand Symmetric model.

q-bio.PE

Methodological considerations for semialgebraic hypothesis testing with incomplete U-statistics

Recently, Sturma, Drton, and Leung proposed a general-purpose stochastic method for hypothesis testing in models defined by polynomial equality and inequality constraints. Notably, the method remains theoretically valid even near irregular points, such as singularities and boundaries, where traditional testing approaches often break down. In this paper, we evaluate its practical performance on a collection of biologically motivated models from phylogenetics. While the method performs remarkably well across different settings, we catalogue a number of issues that should be considered for effective application.

q-bio.PE

New directions in algebraic statistics: Three challenges from 2023

In the last quarter of a century, algebraic statistics has established itself as an expanding field which uses multilinear algebra, commutative algebra, computational algebra, geometry, and combinatorics to tackle problems in mathematical statistics. These developments have found applications in a growing number of areas, including biology, neuroscience, economics, and social sciences. Naturally, new connections continue to be made with other areas of mathematics and statistics. This paper outlines three such connections: to statistical models used in educational testing, to a classification problem for a family of nonparametric regression models, and to phase transition phenomena under uniform sampling of contingency tables. We illustrate the motivating problems, each of which is for algebraic statistics a new direction, and demonstrate an enhancement of related methodologies.

math.ST

Equidistant Circular Split Networks

Phylogenetic networks are generalizations of trees that allow for the modeling of non-tree like evolutionary processes. Split networks give a useful way to construct networks with intuitive distance structures induced from the associated split graph. We explore the polyhedral geometry of distance matrices built from circular split systems which have the added property of being equidistant. We give a characterization of the facet defining inequalities and the extreme rays of the cone of distances that arises from an equidistant network associated to any circular split network. We also explain a connection to the Chan-Robbins-Yuen polytope from geometric combinatorics.

math.CO

Phase Transition in the One-bit Johnson-Lindenstrauss Lemma

The Johnson-Lindenstrauss Lemma (J-L Lemma) is a cornerstone of dimension reduction techniques. We study it in the one-bit context, namely we consider the unit sphere $ \mathbb S ^{N-1}$, with normalized geodesic metric, and map a finite set $ \mathbf{X} \subset \mathbb{S}^{N-1}$ into the Hamming cube $\mathbb{H}_m = \{0,1\}^m$, with normalized Hamming metric. We find that for $ 0< δ<1$, and $m>\frac{\ln n}{2δ^2}$ there is a $δ$-RIP from $\mathbf{X}$ into $\mathbb{H}_m$. This is surprising as the value of $ m$ is virtually identical to best known bound linear J-L Lemma. In both the linear and one-bit case, the maps are randomly constructed. We show that the probability of $B_m$ being a $δ$-RIP satisfies a phase transition. It passes from probability of nearly zero to nearly one with a very small change in $m$. Our proof relies on delicate properties of Bernoulli random variables.

math.FA

An analysis of a fair division protocol for drawing legislative districts

Landau, Reid, and Yershov [A Fair Division Solution to the Problem of Redistricting, \textit{Social Choice and Welfare}, 2008] propose a protocol for drawing legislative districts based on a two player fair division process, where each player is entitled to draw the districts for a portion of the state. We call this the \textit{LRY protocol}. Landau and Su [Fair Division and Redistricting, arXiv:1402.0862, 2014] propose a measure of the fairness of a state's districts called the \textit{geometric target}. In this paper we prove that the number of districts a party can win under the LRY protocol can be at most two fewer than their geometric target, assuming no geometric constraints on the districts, and provide examples to prove this bound is tight. We also show that if the LRY protocol is applied on a state with geometric constraints, the result can be arbitrarily far from the geometric target.

math.CO