SearcharxivSearch

arXiv subjects

Peter F. Stadler

Publications and source records attributed to Peter F. Stadler.

At least 19 recordsLinked to original sources

Best Matches in Phylogenetic Networks

Best match graphs (BMGs) were introduced in mathematical phylogenetics to describe the concept of closest relatives for related genes (leaves of rooted tree) in different organisms (defining leaf colors). We generalize this concept here to leaf-colored rooted networks, where least common ancestors are in general neither unique nor comparable. We characterize BMGs of rooted networks as those vertex-colored digraphs that are properly colored and satisfy an easy-to-check condition that we call the sicor-in-hub property. BMGs can be recognized in linear time and an explaining network can be constructed in quadratic time. Analogous results are obtained for reciprocal best match graphs (RBMGs), where an edge $\{x,y\}$ corresponds to pairs of vertices with different color that are mutually closest relatives.

q-bio.PE

Mutational dynamics of very short tandem repeats

Short tandem repeats (STRs) are low-entropy regions in the genome, consisting of a short (1-6 bp) unit that is consecutively repeated multiple times. They are known for high mutational instability, due to so-called stutter-mutations, in which the number of units in the run increases or descreases. In particular, STRs with repeat unit length of 1-2 bp are prone to mutate even within several cell divisions. The extremely rapid accumulation of variation makes them interesting phylogenetic markers for retrospective single-cell lineage reconstruction. Here we model their mutational dynamics at the level of individual repeat unit motif and then aggregate length variations over many STR loci with the aim of obtaining a very fast ``molecular clock''. We calibrate our model based on several datasets with known lineage structure prepared from cultured cells. We also verify the error margins around our estimated parameters for the model using simulated data. We find that the mutational dynamics of STRs are reasonably consistent for a given cell-line, but vary among different ones. This suggests that the dynamics are not entirely explained by mutations in caretaker genes, rather, various other factors play a role --- possibly tissue origin and differentiation state. Further data and research is necessary to assess their relative effects.

q-bio.PE

Rule-Based Gillespie Simulation of Chemical Systems

The MØD computational framework implements rule-based generative chemistries as explicit transformations of graphs representing chemical structural formulae. Here, we expand MØD by a stochastic simulation module that simulates the time evolution of species concentrations using Gillespie's well-known stochastic simulation algorithm (SSA). This module distinguishes itself among competing implementations of rule-based stochastic simulation engines by its flexible network expansion mechanism and its functionality for defining custom reaction rate functions. It enables direct sampling from actual reactions instead of rules. We present methodology and implementation details followed by examples which demonstrate the capabilities of the stochastic simulation engine.

q-bio.MN

On the realizability of abstract reaction networks with real molecules and reactions

Abstract reaction networks appear not only as models of chemical reactions but also as models of complex systems, with applications in areas such as ecology and epidemiology, and as one of several alternative paradigms for non-standard computation. It is therefore of interest to determine whether an abstract reaction network can be realized by a concrete set of molecules and plausible chemical reaction mechanisms. It is known that a reaction network has a realization in terms of chemical graphs (i.e., Lewis structures) if and only if it is conservative. Here we consider the problem of assigning a set M of known molecules to a set X of abstract entities in a reaction network (X,R) such that each reaction satisfies mass balance and adheres to one of an allowed set of chemical reaction mechanisms. We show that this problem is NP-complete. Nevertheless, it can be solved in practice using a backtrack-and-prune algorithm inspired by the VF2 family of algorithms originally designed for the subgraph isomorphism problem.

q-bio.MN

Bridging two theoretical frameworks of autocatalysis: RAF sets and stoichiometric autocatalysis

Autocatalysis lies at the heart of many (bio)chemical processes and is key to processes leading up to the origin of life. Two seemingly very different formalisms have emerged that define autocatalysis. Kauffman introduced collective autocatalysis to describe systems of molecules that mutually catalyze each other's formation, emphasizing the self-sustaining character of autocatalytic systems. This view is mathematically formalized in the theory of Reflexively Autocatalytic and Food-generated sets (RAF). In parallel, stoichiometric autocatalysis emerged from the theory of Chemical Reaction Networks (CRN), focusing on the net-productive, self-amplifying character of autocatalytic subnetworks. These two frameworks have coexisted independently in the literature, since RAF theory considers each reaction as explicitly catalyzed, while the CRN approach often excludes explicitly catalyzed reactions altogether. Nevertheless, both frameworks describe reaction networks and thus admit a common mathematical representation in terms of stoichiometric matrices. We highlight this connection and show that the two formalisms are less disparate than they might appear. To illustrate this point we prove that, under mild and general conditions, any RAF is stoichiometrically autocatalytic.

q-bio.MN

Enumeration of Autocatalytic Subsystems in Large Chemical Reaction Networks

Autocatalysis is an important feature of metabolic networks, contributing crucially to the self-maintenance of organisms. Autocatalytic subsystems of chemical reaction networks (CRNs) are characterized in terms of algebraic conditions on submatrices of the stoichiometric matrix. Here, we derive sufficient conditions for subgraphs supporting irreducible autocatalytic systems in the bipartite Kőnig representation of the CRN. On this basis, we develop an efficient algorithm to enumerate autocatalytic subnetworks and, as a special case, autocatalytic cores, i.e., minimal autocatalytic subnetworks, in full-size metabolic networks. The same algorithmic approach can also be used to determine autocatalytic cores only. As a showcase application, we provide a complete analysis of autocatalysis in the core metabolism of E. coli and enumerate irreducible autocatalytic subsystems of limited size in full-fledged metabolic networks of E. coli, human erythrocytes, and Methanosarcina barkeri (Archea). The mathematical and algorithmic results are accompanied by software enabling the routine analysis of autocatalysis in large CRNs.

q-bio.MN

Global Least Common Ancestor (LCA) Networks

Directed acyclic graphs (DAGs) are fundamental structures used across many scientific fields. A key concept in DAGs is the least common ancestor (LCA), which plays a crucial role in understanding hierarchical relationships. Surprisingly little attention has been given to DAGs that admit a unique LCA for every subset of their vertices. Here, we characterize such global lca-DAGs and provide multiple structural and combinatorial characterizations. We show that global lca-DAGs have a close connection to join semi-lattices and establish a connection to forbidden topological minors. In addition, we introduce a constructive approach to generating global lca-DAGs and demonstrate that they can be recognized in polynomial time. We investigate their relationship to clustering systems and other set systems derived from the underlying DAGs.

math.CO

Smooth Graphs

The notion of smoothness was introduced originally in the context of step systems on connected graphs. Smoothness turns out to be a very general property of metrics defined by a five-point condition. Restricted to graphs, it is closely related to the convexity of point-shadows. We show that smoothness is preserved by isometric subgraphs, both Cartesian and strong graph products, and gated amalgams. As a consequence, median graphs and many of their generalizations are smooth. We also show that l1-graphs are smooth. On the other hand, an induced K2,3 or K1,1,3 is incompatible with smoothness. Finally, we characterize smooth graphs among the Ptolemaic graphs as precisely the K1,1,3-free Ptolemaic graphs.

math.CO

Autocatalytic Cores in Reaction Networks with Explicit Catalysis

Autocatalytic cores are minimal units in reaction networks (RNs) responsible for the emergence of autocatalysis. In the absence of explicit catalysis, i.e., when an entity appears both as reactant and product in the same reaction, they are known to be encoded by square submatrices of the stoichiometric matrix whose columns can be reordered as an irreducible child-selection (CS) matrix with negative diagonal and nonnegative off-diagonal (Metzler matrix). In the bipartite Koenig graph representing the RN, these CS matrices can be identified by fluffles, i.e., strong blocks with an identical number of entity and reaction vertices that have out- and in-degree 1, respectively. Here, we adapt the concepts derived for autocatalytic cores to RNs with explicitly catalyzed reactions, which emerge as digons, i.e., elementary circuits in the Koenig graph of length 2. In this setting, we confirm that an inspection of the stoichiometric matrix alone is inconclusive concerning the presence and number of autocatalytic cores, requiring a more delicate algebraic analysis. Nevertheless, this generalization preserves both the graph and the matrix representation as fluffles and irreducible Metzler CS matrices, respectively, although the diagonal is no longer necessarily strictly negative. We introduce the notion of hard autocatalytic cores, i.e. those that do not yield other autocatalytic cores upon inclusion of all reverse reactions. Finally, we consider the case of unit stoichiometries and show that each autocatalytic core can be constructed as the superposition of at most 2 elementary circuits. In particular, autocatalytic cores involving explicitly catalyzed reactions always contain a spanning subgraph consisting of a single elementary circuit together with a simple entity-to-reaction chord. Moreover, we identify the essentially unique example for which at least two circuits are required.

math.CO

A Short Note on Relevant Cuts

The set of relevant cuts in a graph is the union of all minimum weight bases of the cut space. A cut is relevant if and only if it is the a minimum weight cut between two distinct vertices. Moreover, we give a characterization in terms of Picard-Queyranne Directed Acyclic Graphs that can be used to accelerate the enumeration of the relevant cuts. Finally, we perform an experimental evaluation by comparing with state-of-the-art algorithms.

math.CO

SynRXN: An Open Benchmark and Curated Dataset for Computational Reaction Modeling

We present SynRXN, a unified benchmarking framework and open-data resource for computer-aided synthesis planning (CASP). SynRXN decomposes end-to-end synthesis planning into five task families, covering reaction rebalancing, atom-to-atom mapping, reaction classification, reaction property prediction, and synthesis route design. Curated, provenance-tracked reaction corpora are assembled from heterogeneous public sources into a harmonized representation and packaged as versioned datasets for each task family, with explicit source metadata, licence tags, and machine-readable manifests that record checksums, and row counts. For every task, SynRXN provides transparent splitting functions that generate leakage-aware train, validation, and test partitions, together with standardized evaluation workflows and metric suites tailored to classification, regression, and structured prediction settings. For sensitive benchmarking, we combine public training and validation data with held-out gold-standard test sets, and contamination-prone tasks such as reaction rebalancing and atom-to-atom mapping are distributed only as evaluation sets and are explicitly not intended for model training. Scripted build recipes enable bitwise-reproducible regeneration of all corpora across machines and over time, and the entire resource is released under permissive open licences to support reuse and extension. By removing dataset heterogeneity and packaging transparent, reusable evaluation scaffolding, SynRXN enables fair longitudinal comparison of CASP methods, supports rigorous ablations and stress tests along the full reaction-informatics pipeline, and lowers the barrier for practitioners who seek robust and comparable performance estimates for real-world synthesis planning workloads.

cs.LG

Forbidden configurations and dominating bicliques in undirected 2-quasi best match graphs

2-quasi best match graphs (2-qBMGs) are directed graphs that capture a notion of close relatedness in phylogenetics. Here, we investigate the undirected underlying graph of a 2-qBMG (un-2qBMG) and show that they contain neither a path $P_l$ nor a cycle $C_l$ of length $l\geq 6$ as an induced subgraph. This property guarantees the existence of specific vertex decompositions with dominating bicliques that provide further insights into their structure.

math.CO

BiRNe: Symbolic bifurcation analysis of reaction networks with Python

Computer algebra methods for analyzing reaction networks often rely on the assumption of mass-action kinetics, which transform the governing ODEs into polynomial systems amenable to techniques such as Gröbner basis computation and related algebraic tools. However, these methods face significant computational complexity, limiting their applicability to relatively small networks involving only a handful of species. In contrast, building on recent theoretical advances, we introduce here \textsc{BiRNe} (BIfurcations in Reaction NEtworks) Python module, which relies on a symbolic approach designed to detect bifurcations in larger reaction networks (up to 10-20 species, depending on the network's connectivity) equipped with parameter-rich kinetics. This class includes enzymatic kinetics such as Michaelis--Menten, ligand-binding kinetics like Hill functions, and generalized mass-action kinetics. For a given network, the current algorithm identifies all minimal autocatalytic subnetworks and fully characterizes the presence of bifurcations associated with zero eigenvalues, thus determining whether the network admits multistationarity. It also detects oscillatory bifurcations arising from positive-feedback structures, capturing a significant class of possible oscillations.

q-bio.MN

Characterizations of undirected 2-quasi best match graphs

Bipartite best match graphs (BMG) and their generalizations arise in mathematical phylogenetics as combinatorial models describing evolutionary relationships among related genes in a pair of species. In this work, we characterize the class of \emph{undirected 2-quasi-BMGs} (un2qBMGs), which form a proper subclass of the $P_6$-free chordal bipartite graphs. We show that un2qBMGs are exactly the class of bipartite graphs free of $P_6$, $C_6$, and the eight-vertex Sunlet$_4$ graph. Equivalently, a bipartite graph $G$ is un2qBMG if and only if every connected induced subgraph contains a ``heart-vertex'' which is adjacent to all the vertices of the opposite color. We further provide a $O(|V(G)|^3)$ algorithm for the recognition of un2qBMGs that, in the affirmative case, constructs a labeled rooted tree that ``explains'' $G$. Finally, since un2qBMGs coincide with the $(P_6,C_6)$-free bi-cographs, they can also be recognized in linear time.

math.CO

Prime Implicant Explanations for Reaction Feasibility Prediction

Machine learning models that predict the feasibility of chemical reactions have become central to automated synthesis planning. Despite their predictive success, these models often lack transparency and interpretability. We introduce a novel formulation of prime implicant explanations--also known as minimally sufficient reasons--tailored to this domain, and propose an algorithm for computing such explanations in small-scale reaction prediction tasks. Preliminary experiments demonstrate that our notion of prime implicant explanations conservatively captures the ground truth explanations. That is, such explanations often contain redundant bonds and atoms but consistently capture the molecular attributes that are essential for predicting reaction feasibility.

cs.LG

Stoichiometric recipes for periodic oscillations in reaction networks

Oscillatory chemical reactions are functional components in a variety of biological contexts. In chemistry, the construction and identification of even rudimentary oscillators remain elusive and lack a general framework. Using parameter-rich kinetics - a methodology enabling the disentanglement of parametric dependencies from structural analysis - we investigate the stoichiometry of chemical oscillators. We introduce the concept of oscillatory cores: minimal subnetworks that guarantee the potential for oscillations in any reaction network containing them. These cores fall into two classes, depending on whether they involve positive or negative feedback. In particular, the latter class unveils a family of oscillators - yet to be synthesized - that require a minimum number of reaction steps to exhibit oscillations, a phenomenon we refer to as the principle of length. We identify several mechanisms through which catalysis promotes oscillations: (I) furnishing instability (e.g. autocatalysis), (II) lifting dependencies, (III) lowering length thresholds. Notwithstanding this mechanistic ubiquity, we show that oscillators can also be realized without employing any catalysis. Our results highlight branches of chemistry where oscillators are likely to arise by chance, suggest new strategies for their design, and point to novel classes of oscillators yet to be realized experimentally.

q-bio.MN

Assembly in Directed Hypergraphs

Assembly theory has received considerable attention in the recent past. Here we analyze the formal framework of this model and show that assembly pathways coincide with certain minimal hyperpaths in B-hypergraphs. This makes it possible to generalize the notion of assembly to general chemical reaction systems and to make explicit the connection to rule based models of chemistry, in particular DPO graph rewriting. We observe, furthermore, that assembly theory is closely related to retrosynthetic analysis in chemistry. The assembly index fits seamlessly into a large family of cost measures for directed hyperpath problems that also encompasses cost functions used in computational synthesis planning. This allows to devise a generic approach to compute complexity measures derived from minimal hyperpaths in rule-derived directed hypergraphs using integer linear programming.

cs.DM

Pathway Realisability in Chemical Networks

The exploration of pathways and alternative pathways that have a specific function is of interest in numerous chemical contexts. A framework for specifying and searching for pathways has previously been developed, but a focus on which of the many pathway solutions are realisable, or can be made realisable, is missing. Realisable here means that there actually exists some sequencing of the reactions of the pathway that will execute the pathway. We present a method for analysing the realisability of pathways based on the reachability question in Petri nets. For realisable pathways, our method also provides a certificate encoding an order of the reactions which realises the pathway. We present two extended notions of realisability of pathways, one of which is related to the concept of network catalysts. We exemplify our findings on the pentose phosphate pathway. Furthermore, we discuss the relevance of our concepts for elucidating the choices often implicitly made when depicting pathways. Lastly, we lay the foundation for the mathematical theory of realisability.

q-bio.MN