SearcharxivSearch

arXiv subjects

Keith Y. Patarroyo

Publications and source records attributed to Keith Y. Patarroyo.

5 recordsLinked to original sources

Assembly Addition Chains

In this paper we extend the notion of Addition Chains over Z+ to a general set S. We explain how the algebraic structure of Assembly Multi-Magma over the pairs (S,BB proper subset of S) allows to define the concept of Addition Chain over S, called Assembly Addition Chains of S with Building Blocks BB. Analogously to the Z+ case, we introduce the concept of Optimal Assembly Addition Chains over S and prove lower and upper bounds for their lengths, similar to the bounds found by Schonhage for the Z+ case. In the general case the unit 1 is in set Z+ is replaced by the subset BB and the mentioned bounds for the length of an Optimal Assembly Addition Chain of O is in set S are defined in terms of the size of O (i.e. the number of Building Blocks required to construct O). The main examples of S that we consider through this papers are (i) j-Strings (Strings with an alphabeth of j letters), (ii) Colored Connected Graphs and (iii) Colored Polyominoes.

math.CO

Rapid Exploration of Assembly Chemical Space of Molecular Graphs

Quantifying how hard it is to build a molecular graph matters for biosignature detection, chemical complexity, and cheminformatics. We present an exact, scalable algorithm to compute the molecular assembly index (MA) which prioritizes the largest duplicate subgraphs, represents fragmentation with an 'assembly state' array of edge-lists, reuses states via hashing/DAGs, and prunes the search using a dynamic-programming branch-and-bound guided by a conditional-addition-chain lower bound. For organic molecules in the greater than 500 Da range our approach is up to six orders of magnitude faster than prior methods and yields exact MAs where previous algorithms would have timed out. We compute MAs to convergence for ~300k COCONUT natural products with <50 bonds, profiling time and memory scaling. Finally, we exploit the speed of our algorithm to calculate joint assembly spaces and introduce the Joint Assembly Overlap (JAO), a Jaccard-like metric that emphasizes global scaffold reuse and show that the JAO yields substantially different rankings from Tanimoto similarity with ECFP fingerprints and MCS (e.g. in steroids 270-380/Da and short peptides), accounting for substructural similarity beyond local environments. Together, these advances turn the molecular assembly index into a practical tool for large-scale exploration of chemical space.

cs.DS

A digression on Hermite polynomials

Orthogonal polynomials are of fundamental importance in many fields of mathematics and science, therefore the study of a particular family is always relevant. In this manuscript, we present a survey of some general results of the Hermite polynomials and show a few of their applications in the connection problem of polynomials, probability theory and the combinatorics of a simple graph. Most of the content presented here is well known, except for a few sections where we add our own work to the subject, nevertheless, the text is meant to be a self-contained personal exposition.

math.NA

Mean conservation for density estimation via diffusion using the finite element method

We propose boundary conditions for the diffusion equation that maintain the initial mean and the total mass of a discrete data sample in the density estimation process. A complete study of this framework with numerical experiments using the finite element method is presented for the one dimensional diffusion equation, some possible applications of this results are presented as well. We also comment on a similar methodology for the two-dimensional diffusion equation for future applications in two-dimensional domains.

math.NA

Pronunciation recognition of English phonemes /\textipa{@}/, /æ/, /\textipa{A}:/ and /\textipa{2}/ using Formants and Mel Frequency Cepstral Coefficients

The Vocal Joystick Vowel Corpus, by Washington University, was used to study monophthongs pronounced by native English speakers. The objective of this study was to quantitatively measure the extent at which speech recognition methods can distinguish between similar sounding vowels. In particular, the phonemes /\textipa{@}/, /æ/, /\textipa{A}:/ and /\textipa{2}/ were analysed. 748 sound files from the corpus were used and subjected to Linear Predictive Coding (LPC) to compute their formants, and to Mel Frequency Cepstral Coefficients (MFCC) algorithm, to compute the cepstral coefficients. A Decision Tree Classifier was used to build a predictive model that learnt the patterns of the two first formants measured in the data set, as well as the patterns of the 13 cepstral coefficients. An accuracy of 70\% was achieved using formants for the mentioned phonemes. For the MFCC analysis an accuracy of 52 \% was achieved and an accuracy of 71\% when /\textipa{@}/ was ignored. The results obtained show that the studied algorithms are far from mimicking the ability of distinguishing subtle differences in sounds like human hearing does.

cs.CL