SearcharxivSearch

arXiv subjects

Hexuan Liu

Publications and source records attributed to Hexuan Liu.

10 recordsLinked to original sources

A Combinatorial Framework for the Pons-Batle Identity: Young Tableaux, Lattice Paths, and Limit Laws

Tree-child networks are an important class of phylogenetic network used to model reticulate evolutionary processes. These networks have attracted increasing attention from researchers with interests in both combinatorics and algorithms. A fundamental open problem posed by Pons and Batle asks whether the number $TC_{n,k}$ of bicombining tree-child networks with $n$ leaves and $k$ reticulation nodes equals the number of certain constrained words, now called Pons-Batle words. In this paper, we confirm the conjecture for tree-child networks with a bounded number of reticulation nodes. Our approach is combinatorial and analytic. We introduce families of Young tableaux with walls and holes and construct explicit bijections with Pons-Batle words, yielding a direct combinatorial explanation of the identities. These tableaux encode structural features of the underlying networks, including the placement of reticulation nodes. By projecting them to decorated Dyck paths, we obtain algebraic generating functions with differential operators encoding step weights, leading to explicit recurrence relations and closed-form formulas for $TC_{n,k}$. Beyond finite verification for moderate $k$, the framework reveals an underlying probabilistic structure. For $k=1$, natural structural parameters, such as the position and value of distinguished cells, converge, after rescaling, to $\mathrm{Beta}(2,1)$, $\mathrm{Beta}(1,2)$, and Uniform (i.e., $\mathrm{Beta}(1,1)$) distributions. These limit laws arise from a coalescence of singularities at the dominant square-root singularity, producing a non-analytic transition in the local expansion. Overall, our results provide both combinatorial insight and a unified analytic perspective on the asymptotic behavior of tree-child networks, showing how algebraic generating functions with interacting singularities systematically produce Beta limit laws.

math.CO

A characterization of terminal planar networks by forbidden structures

The class of terminal planar networks was recently introduced from a biological perspective in relation to the visualization of phylogenetic networks, and its connection to upward planar networks has been established. We provide a Kuratowski-type theorem that characterizes terminal planar networks by a finite set of forbidden structures, defined via six families of 0/1-labeled graphs. Another characterization based on planarity of supergraphs yields linear-time algorithms for testing terminal planarity and for computing such planar drawings. We describe an application that is potentially relevant in broader, non-phylogenetic settings. We also discuss a connection of our main result to an open problem on the forbidden structures of single-source upward planar networks.

math.CO

Asymptotic Enumeration of Subclasses of Level-$2$ Phylogenetic Networks

This paper studies the enumeration of seven subclasses of level-$2$ phylogenetic networks under various planarity and structural constraints, including terminal planar, tree-child, and galled networks. We derive their exponential generating functions, recurrence relations, and asymptotic formulas. Specifically, we show that the number of networks of size $n$ in each class follows: \[ N_n \sim c \cdot n^{n-1} \cdot \gamma^n, \] where $c$ is a class-specific constant and $\gamma$ is the corresponding growth rate. Our results reveal that being terminal planar can significantly reduce the growth rate of general level-2 networks, but has only a minor effect on the growth rates of tree-child and galled level-2 networks. Notably, the growth rate of 3.83 for level-$2$ terminal planar galled tree-child networks is remarkably close to the rate of 2.94 for level-$1$ networks.

math.CO

Enumerative and Distributional Results for $d$-combining Tree-Child Networks

Tree-child networks are one of the most prominent network classes for modeling evolutionary processes which contain reticulation events. Several recent studies have addressed counting questions for bicombining tree-child networks in which every reticulation node has exactly two parents. We extend these studies to $d$-combining tree-child networks where every reticulation node has now $d\geq 2$ parents, and we study one-component as well as general tree-child networks. For the number of one-component networks, we derive an exact formula from which asymptotic results follow that contain a stretched exponential for $d=2$, yet not for $d \geq 3$. For general networks, we find a novel encoding by words which leads to a recurrence for their numbers. From this recurrence, we derive asymptotic results which show the appearance of a stretched exponential for all $d \geq 2$. Moreover, we also give results on the distribution of shape parameters (e.g., number of reticulation nodes, Sackin index) of a network which is drawn uniformly at random from the set of all tree-child networks with the same number of leaves. We show phase transitions depending on $d$, leading to normal, Bessel, Poisson, and degenerate distributions. Some of our results are new even in the bicombining case.

math.CO

Limit Theorems for Patterns in Ranked Tree-Child Networks

We prove limit laws for the number of occurrences of a pattern on the fringe of a ranked tree-child network which is picked uniformly at random. Our results extend the limit law for cherries proved by Bienvenu et al. (2022). For patterns of height $1$ and $2$, we show that they either occur frequently (mean is asymptotically linear and limit law is normal) or sporadically (mean is asymptotically constant and limit law is Poisson) or not all (mean tends to $0$ and limit law is degenerate). We expect that these are the only possible limit laws for any fringe pattern.

math.PR

Enumeration of $d$-combining Tree-Child Networks

Tree-child networks are one of the most prominent network classes for modeling evolutionary processes which contain reticulation events. Several recent studies have addressed counting questions for {\it bicombining tree-child networks} which are tree-child networks with every reticulation node having exactly two parents. In this paper, we extend these studies to {\it $d$-combining tree-child networks} where every reticulation node has now $d\geq 2$ parents. Moreover, we also give results and conjectures on the distributional behavior of the number of reticulation nodes of a network which is drawn uniformly at random from the set of all tree-child networks with the same number of leaves.

math.CO

A Short Note on the Exact Counting of Tree-Child Networks

Tree-child networks are an important network class which are used in phylogenetics to model reticulate evolution. In a recent paper, Pons and Batle (2021) conjectured a relation between tree-child networks and certain words. In this short note, we prove their conjecture for the (important) class of one-component tree-child networks.

q-bio.PE

Analysis of Truncated Orthogonal Iteration for Sparse Eigenvector Problems

A wide range of problems in computational science and engineering require estimation of sparse eigenvectors for high dimensional systems. Here, we propose two variants of the Truncated Orthogonal Iteration to compute multiple leading eigenvectors with sparsity constraints simultaneously. We establish numerical convergence results for the proposed algorithms using a perturbation framework, and extend our analysis to other existing alternatives for sparse eigenvector estimation. We then apply our algorithms to solve the sparse principle component analysis problem for a wide range of test datasets, from simple simulations to real-world datasets including MNIST, sea surface temperature and 20 newsgroups. In all these cases, we show that the new methods get state of the art results quickly and with minimal parameter tuning.

math.NA

On the Convergence Rate of Variants of the Conjugate Gradient Algorithm in Finite Precision Arithmetic

We consider three mathematically equivalent variants of the conjugate gradient (CG) algorithm and how they perform in finite precision arithmetic. It was shown in [{\em Behavior of slightly perturbed Lanczos and conjugate-gradient recurrences}, Lin.~Alg.~Appl., 113 (1989), pp.~7-63] that under certain conditions the convergence of a slightly perturbed CG computation is like that of exact CG for a matrix with many eigenvalues distributed throughout tiny intervals about the eigenvalues of the given matrix, the size of the intervals being determined by how closely these conditions are satisfied. We determine to what extent each of these variants satisfies the desired conditions, using a set of test problems and show that there is significant correlation between how well these conditions are satisfied and how well the finite precision computation converges before reaching its ultimately attainable accuracy. We show that for problems where the width of the intervals containing the eigenvalues of the associated exact CG matrix makes a significant difference in the behavior of exact CG, the different CG variants behave differently in finite precision arithmetic. For problems where the interval width makes little difference or where the convergence of exact CG is essentially governed by the upper bound based on the square root of the condition number of the matrix, the different CG variants converge similarly in finite precision arithmetic until the ultimate level of accuracy is achieved, although this ultimate level of accuracy may be different for the different variants. This points to the need for testing new CG variants on problems that are especially sensitive to rounding errors.

math.NA

Functional Connectomics from Data: Probabilistic Graphical Models for Neuronal Network of C. elegans

We propose a data-driven approach to represent neuronal network dynamics as a Probabilistic Graphical Model (PGM). Our approach learns the PGM structure by employing dimension reduction to network response dynamics evoked by stimuli applied to each neuron separately. The outcome model captures how stimuli propagate through the network and thus represents functional dependencies between neurons, i.e., functional connectome. The benefit of using a PGM as the functional connectome is that posterior inference can be done efficiently and circumvent the complexities in direct inference of response pathways in dynamic neuronal networks. In particular, posterior inference reveals the relations between known stimuli and downstream neurons or allows to query which stimuli are associated with downstream neurons. For validation and as an example for our approach we apply our methodology to a model of Caenorhabiditis elegans nervous system which structure and dynamics are well-studied. From its dynamical model we collect time series of the network response and use singular value decomposition to obtain a low-dimensional projection of the time series data. We then extract dominant patterns in each data matrix to get pairwise dependency information and create a graphical model for the full somatic nervous system. The PGM enables us to obtain and verify underlying neuronal pathways dominant for known behavioral scenarios and to detect possible pathways for novel scenarios.

q-bio.NC