SearcharxivSearch

arXiv subjects

Christian M. Reidys

Publications and source records attributed to Christian M. Reidys.

At least 19 recordsLinked to original sources

Divide and Predict: An Architecture for Input Space Partitioning and Enhanced Accuracy

In this article the authors develop an intrinsic measure for quantifying heterogeneity in training data for supervised learning. This measure is the variance of a random variable which factors through the influences of pairs of training points. The variance is shown to capture data heterogeneity and can thus be used to assess if a sample is a mixture of distributions. The authors prove that the data itself contains key information that supports a partitioning into blocks. Several proof of concept studies are provided that quantify the connection between variance and heterogeneity for EMNIST image data and synthetic data. The authors establish that variance is maximal for equal mixes of distributions, and detail how variance-based data purification followed by conventional training over blocks can lead to significant increases in test accuracy.

cs.LG

Exact enumeration of RNA secondary structures by helices and loops

Enumerative studies of RNA secondary structures were initiated four decades ago by Waterman and his coworkers. Since then, RNA secondary structures have been explored according to many different structural characteristics, for instance, helices, components and loops by Hofacker, Schuster and Stadler, orders by Nebel, saturated structures by Clote, the $5^{\prime}$-$3^{\prime}$ end distance by Clote, Ponty and Steyaert, and the rainbow spectrum by Li and Reidys. However, the majority of the contributions are asymptotic results, and it is harder to derive explicit formulas. In this paper, we obtain exact formulas counting RNA secondary structures with a given number of helices as well as a given joint size distribution of helices and loops, while some related asymptotic results due to Hofacker, Schuster and Stadler have been known for about twenty years. Our approach is combinatorial, analyzing a recent bijection between RNA secondary structures and plane trees discovered by the first author and proposing a variation of Chen's bijective approach of counting trees by forests of simple trees.

math.CO

A computational framework for weighted simplicial homology

We provide a bottom up construction of torsion generators for weighted homology of a weighted complex over a discrete valuation ring $R=\mathbb{F}[[π]]$. This is achieved by starting from a basis for classical homology of the $n$-th skeleton for the underlying complex with coefficients in the residue field $\mathbb{F}$ and then lifting it to a basis for the weighted homology with coefficients in the ring $R$. Using the latter, a bijection is established between $n+1$ and $n$ dimensional simplices whose weight ratios provide the exponents of the $π$-monomials that generate each torsion summand in the structure theorem of the weighted homology modules over $R$. We present algorithms that subsume the torsion computation by reducing it to normalization over the residue field of $R$, and describe a Python package we implemented that takes advantage of this reduction and performs the computation efficiently.

math.AT

On Weighted Simplicial Homology

We develop a framework for computing the homology of weighted simplicial complexes with coefficients in a discrete valuation ring. A weighted simplicial complex, $(X,v)$, introduced by Dawson [Cah. Topol. Géom. Différ. Catég. 31 (1990), pp. 229--243], is a simplicial complex, $X$, together with an integer-valued function, $v$, assigning weights to simplices, such that the weight of any of faces are monotonously increasing. In addition, weighted homology, $H_n^v(X)$, features a new boundary operator, $\partial_n^v$. In difference to Dawson, our approach is centered at a natural homomorphism $θ$ of weighted chain complexes. The key object is $H^v_{n}(X/θ)$, the weighted homology of a quotient of chain complexes induced by $θ$, appearing in a long exact sequence linking weighted homologies with different weights. We shall construct bases for the kernel and image of the weighted boundary map, identifying $n$-simplices as either $κ_n$- or $μ_n$-vertices. Long exact sequences of weighted homology groups and the bases, allow us to prove a structure theorem for the weighted simplicial homology with coefficients in a ring of formal power series $R=\mathbb{F}[[π]]$, where $\mathbb{F}$ is a field. Relative to simplicial homology new torsion arises and we shall show that the torsion modules are connected to a pairing between distinguished $κ_n$ and $μ_{n+1}$ simplices.

math.AT

The energy-spectrum of bicompatible sequences

Background: Genotype-phenotype maps provide a meaningful filtration of sequence space and RNA secondary structures are particular such phenotypes. Compatible sequences i.e.~sequences that satisfy the base pairing constraints of a given RNA structure play an important role in the context of neutral networks and inverse folding. Sequences satisfying the constraints of two structures simultaneously are called bicompatible and phenotypic change, induced by erroneously replicating populations of RNA sequences, is closely connected to bicompatibility. Furthermore, bicompatible sequences are relevant for riboswitch sequences, beacons of evolution, realizing two distinct phenotypes. Results: We present a full loop energy model Boltzmann sampler of bicompatible sequences for pairs of structures. The novel dynamic programming algorithm is based on a topological framework encapsulating the relations between loops. We utilize our sequence sampler to study the energy spectra and density of bicompatible sequences, the rankings of the structures and key properties for evolutionary transitions. Conclusion: Our analysis of riboswitch sequences shows that key properties of bicompatible sequences depend on the particular pair of structures. While there always exist bicompatible sequences for random structure pairs, they are less suited to facilitate transitions. We show that native riboswitch sequences exhibit a distinct signature with regards to the ranking of their two phenotypes relative to the minimum free energy, suggesting a new criterion for identifying native sequences and sequences subjected to evolutionary pressure.

q-bio.BM

On an enhancement of RNA probing data using Information Theory

Identifying the secondary structure of an RNA is crucial for understanding its diverse regulatory functions. This paper focuses on how to enhance target identification in a Boltzmann ensemble of structures via chemical probing data. We employ an information-theoretic approach to solve the problem, via considering a variant of the Rényi-Ulam game. Our framework is centered around the ensemble tree, a hierarchical bi-partition of the input ensemble, that is constructed by recursively querying about whether or not a base pair of maximum information entropy is contained in the target. These queries are answered via relating local with global probing data, employing the modularity in RNA secondary structures. We present that leaves of the tree are comprised of sub-samples exhibiting a distinguished structure with high probability. In particular, for a Boltzmann ensemble incorporating probing data, which is well established in the literature, the probability of our framework correctly identifying the target in the leaf is greater than $90\%$.

q-bio.BM

Loop Homology of Bi-secondary Structures II

In this paper we further describe the features of the topological space $K(R)$ obtained from the loop nerve of $R$, for $R=(S,T)$ a bi-secondary structure. We will first identify certain distinct combinatorial structures in the arc diagram of $R$ which we will call crossing components. The main theorem of this paper shows that the total number of these crossing components equals the rank of $H_2(R)$, the second homology group of the loop nerve.

math.CO

D-chain tomography of networks: a new structure spectrum and an application to the SIR process

The analysis of the dynamics on complex networks is closely connected to structural features of the networks. Features like, for instance, graph-cores and node degrees have been studied ubiquitously. Here we introduce the D-spectrum of a network, a novel new framework that is based on a collection of nested chains of subgraphs within the network. Graph-cores and node degrees are merely from two particular such chains of the D-spectrum. Each chain gives rise to a ranking of nodes and, for a fixed node, the collection of these ranks provides us with the D-spectrum of the node. Besides a node deletion algorithm, we discover a connection between the D-spectrum of a network and some fixed points of certain graph dynamical systems (MC systems) on the network. Using the D-spectrum we identify nodes of similar spreading power in the susceptible-infectious-recovered (SIR) model on a collection of real world networks as a quick application. We then discuss our results and conclude that D-spectra represent a meaningful augmentation of graph-cores and node degrees.

math.CO

Loop Homology of Bi-secondary Structures

In this paper we compute the loop homology of bi-secondary structures. Bi-secondary structures were introduced by Haslinger and Stadler and are pairs of RNA secondary structures, i.e. diagrams having non-crossing arcs in the upper half-plane. A bi-secondary structure is represented by drawing its respective secondary structures in the upper and lower half-plane. An RNA secondary structure has a loop decomposition, where a loop corresponds to a boundary component, regarding the secondary structure as an orientable fatgraph. The loop-decomposition of secondary structures facilitates the computation of its free energy and any two loops intersect either trivially or in exactly two vertices. In bi-secondary structures the intersection of loops is more complex and is of importance in current algorithmic work in bio-informatics and evolutionary optimization. We shall construct a simplicial complex capturing the intersections of loops and compute its homology. We prove that only the zeroth and second homology groups are nontrivial and furthermore show that the second homology group is free. Finally, we provide evidence that the generators of the second homology group have a bio-physical interpretation: they correspond to pairs of mutually exclusive substructures.

math.GN

The block spectrum of RNA pseudoknot structures

In this paper we analyze the length-spectrum of blocks in $γ$-structures. $γ$-structures are a class of RNA pseudoknot structures that plays a key role in the context of polynomial time RNA folding. A $γ$-structure is constructed by nesting and concatenating specific building components having topological genus at most $γ$. A block is a substructure enclosed by crossing maximal arcs with respect to the partial order induced by nesting. We show that, in uniformly generated $γ$-structures, there is a significant gap in this length-spectrum, i.e., there asymptotically almost surely exists a unique longest block of length at least $n-O(n^{1/2})$ and that with high probability any other block has finite length. For fixed $γ$, we prove that the length of the longest block converges to a discrete limit law, and that the distribution of short blocks of given length tends to a negative binomial distribution in the limit of long sequences. We refine this analysis to the length spectrum of blocks of specific pseudoknot types, such as H-type and kissing hairpins. Our results generalize the rainbow spectrum on secondary structures by the first and third authors and are being put into context with the structural prediction of long non-coding RNAs.

math.CO

From unicellular fatgraphs to trees

In this paper we study the minimum number of reversals needed to transform a unicellular fatgraph into a tree. We consider reversals acting on boundary components, having the natural interpretation as gluing, slicing or half-flipping of vertices. Our main result is an expression for the minimum number of reversals needed to transform a unicellular fatgraph to a plane tree. The expression involves the Euler genus of the fatgraph and an additional parameter, which counts the number of certain orientable blocks in the decomposition of the fatgraph. In the process we derive a constructive proof of how to decompose non-orientable, irreducible, unicellular fatgraphs into smaller fatgraphs of the same type or trivial fatgraphs, consisting of a single ribbon. We furthermore provide a detailed analysis how reversals affect the component-structure of the underlying fatgraphs. Our results generalize the Hannenhalli-Pevzner formula for the reversal distance of signed permutations.

math.CO

The rainbow-spectrum of RNA secondary structures

In this paper we analyze the length-spectrum of rainbows in RNA secondary structures. A rainbow in a secondary structure is a maximal arc with respect to the partial order induced by nesting. We show that there is a significant gap in this length-spectrum. We shall prove that there asymptotically almost surely exists a unique longest rainbow of length at least $n-O(n^{1/2})$ and that with high probability any other rainbow has finite length. We show that the distribution of the length of the longest rainbow converges to a discrete limit law and that, for finite $k$, the distribution of rainbows of length $k$, becomes for large $n$ a negative binomial distribution. We then put the results of this paper into context, comparing the analytical results with those observed in RNA minimum free energy structures, biological RNA structures and relate our findings to the sparsification of folding algorithms.

math.CO

Garden-of-Eden states and fixed points of monotone dynamical systems

In this paper we analyze Garden-of-Eden (GoE) states and fixed points of monotone, sequential dynamical systems (SDS). For any monotone SDS and fixed update schedule, we identify a particular set of states, each state being either a GoE state or reaching a fixed point, while both determining if a state is a GoE state and finding out all fixed points are generally hard. As a result, we show that the maximum size of their limit cycles is strictly less than ${n\choose \lfloor n/2 \rfloor}$. We connect these results to the Knaster-Tarski theorem and the LYM inequality. Finally, we establish that there exist monotone, parallel dynamical systems (PDS) that cannot be expressed as monotone SDS, despite the fact that the converse is always true.

math.CO

Linear sequential dynamical systems, incidence algebras, and Möbius functions

A sequential dynamical system (SDS) consists of a graph, a set of local functions and an update schedule. A linear sequential dynamical system is an SDS whose local functions are linear. In this paper, we derive an explicit closed formula for any linear SDS as a synchronous dynamical system. We also show constructively, that any synchronous linear system can be expressed as a linear SDS, i.e. it can be written as a product of linear local functions. Furthermore, we study the connection between linear SDS and the incidence algebras of partially ordered sets (posets). Specifically, we show that the Möbius function of any poset can be computed via an SDS, whose graph is induced by the Hasse diagram of the poset. Finally, we prove a cut theorem for the Möbius functions of posets with respect to certain chain decompositions.

math.CO

Genetic robustness of let-7 miRNA sequence-structure pairs

Genetic robustness, the preservation of evolved phenotypes against genotypic mutations, is one of the central concepts in evolution. In recent years a large body of work has focused on the origins, mechanisms, and consequences of robustness in a wide range of biological systems. In particular, research on ncRNAs studied the ability of sequences to maintain folded structures against single-point mutations. In these studies, the structure is merely a reference. However, recent work revealed evidence that structure itself contributes to the genetic robustness of ncRNAs. We follow this line of thought and consider sequence-structure pairs as the unit of evolution and introduce the spectrum of inverse folding rates (IFR-spectrum) as a measurement of genetic robustness. Our analysis of the miRNA let-7 family captures key features of structure-modulated evolution and facilitates the study of robustness against multiple-point mutations.

q-bio.PE

An efficient dual sampling algorithm with Hamming distance filtration

Recently, a framework considering RNA sequences and their RNA secondary structures as pairs, led to some information-theoretic perspectives on how the semantics encoded in RNA sequences can be inferred. In this context, the pairing arises naturally from the energy model of RNA secondary structures. Fixing the sequence in the pairing produces the RNA energy landscape, whose partition function was discovered by McCaskill. Dually, fixing the structure induces the energy landscape of sequences. The latter has been considered for designing more efficient inverse folding algorithms. We present here the Hamming distance filtered, dual partition function, together with a Boltzmann sampler using novel dynamic programming routines for the loop-based energy model. The time complexity of the algorithm is $O(h^2n)$, where $h,n$ are Hamming distance and sequence length, respectively, reducing the time complexity of samplers, reported in the literature by $O(n^2)$. We then present two applications, the first being in the context of the evolution of natural sequence-structure pairs of microRNAs and the second constructing neutral paths. The former studies the inverse fold rate (IFR) of sequence-structure pairs, filtered by Hamming distance, observing that such pairs evolve towards higher levels of robustness, i.e.,~increasing IFR. The latter is an algorithm that constructs neutral paths: given two sequences in a neutral network, we employ the sampler in order to construct short paths connecting them, consisting of sequences all contained in the neutral network.

q-bio.BM

On a lower bound for sorting signed permutations by reversals

Computing the reversal distances of signed permutations is an important topic in Bioinformatics. Recently, a new lower bound for the reversal distance was obtained via the plane permutation framework. This lower bound appears different from the existing lower bound obtained by Bafna and Pevzner through breakpoint graphs. In this paper, we prove that the two lower bounds are equal. Moreover, we confirm a related conjecture on skew-symmetric plane permutations, which can be restated as follows: let $p=(0,-1,-2,\ldots -n,n,n-1,\ldots 1)$ and let $$ \tilde{s}=(0,a_1,a_2,\ldots a_n,-a_n,-a_{n-1},\ldots -a_1) $$ be any long cycle on the set $\{-n,-n+1,\ldots 0,1,\ldots n\}$. Then, $n$ and $a_n$ are always in the same cycle of the product $p\tilde{s}$. Furthermore, we show the new lower bound via plane permutations can be interpreted as the topological genera of orientable surfaces associated to signed permutations.

math.CO

New formulas counting one-face maps and Chapuy's recursion

In this paper, we begin with the Lehman-Walsh formula counting one-face maps and construct two involutions on pairs of permutations to obtain a new formula for the number $A(n,g)$ of one-face maps of genus $g$. Our new formula is in the form of a convolution of the Stirling numbers of the first kind which immediately implies a formula for the generating function $A_n(x)=\sum_{g\geq 0}A(n,g)x^{n+1-2g}$ other than the well-known Harer-Zagier formula. By reformulating our expression for $A_n(x)$ in terms of the backward shift operator $E: f(x)\rightarrow f(x-1)$ and proving a property satisfied by polynomials of the form $p(E)f(x)$, we easily establish the recursion obtained by Chapuy for $A(n,g)$. Moreover, we give a simple combinatorial interpretation for the Harer-Zagier recurrence.

math.CO