Searcharxiv⌕ Search

arXiv subjects

Donald E. K. Martin

Publications and source records attributed to Donald E. K. Martin.

3 recordsLinked to original sources

Large Sample Properties of Higher Order Markov Models

We study large-sample properties of higher-order Markov chains on a finite alphabet $Σ$ when the order $m_n$ is allowed to grow with the sequence length $n$. By embedding the process into a first-order chain on $Σ^{m_n}$ and exploiting return-time decompositions, we establish a central limit theorem for additive functionals $\sum_{t}\! g_n(Y_t^{(n)})$ under natural ergodicity and sparsity conditions. The normalization involves the stationary return time to a suitably chosen state and accommodates triangular arrays with $m_n\!\to\!\infty$ and $m_n/n\!\to\!0$. We further illustrate the assumptions in a binary variable length Markov chain (VLMC), deriving explicit lower bounds on stationary masses that yield a concrete growth regime (e.g., $m_n\log m_n/n \to 0$) ensuring the CLT. These results provide asymptotic foundations for inference in sparse/partitioned higher-order models; including VLMCs and sparse Markov models (SMMs) where the effective dimensionality grows with the sample size.

math.ST↗

A Coverage Criterion for Spaced Seeds and its Applications to Support Vector Machine String Kernels and k-Mer Distances

Spaced seeds have been recently shown to not only detect more alignments, but also to give a more accurate measure of phylogenetic distances (Boden et al., 2013, Horwege et al., 2014, Leimeister et al., 2014), and to provide a lower misclassification rate when used with Support Vector Machines (SVMs) (On-odera and Shibuya, 2013), We confirm by independent experiments these two results, and propose in this article to use a coverage criterion (Benson and Mak, 2008, Martin, 2013, Martin and No{é}, 2014), to measure the seed efficiency in both cases in order to design better seed patterns. We show first how this coverage criterion can be directly measured by a full automaton-based approach. We then illustrate how this criterion performs when compared with two other criteria frequently used, namely the single-hit and multiple-hit criteria, through correlation coefficients with the correct classification/the true distance. At the end, for alignment-free distances, we propose an extension by adopting the coverage criterion, show how it performs, and indicate how it can be efficiently computed.

q-bio.QM↗

Distributions associated with general runs and patterns in hidden Markov models

This paper gives a method for computing distributions associated with patterns in the state sequence of a hidden Markov model, conditional on observing all or part of the observation sequence. Probabilities are computed for very general classes of patterns (competing patterns and generalized later patterns), and thus, the theory includes as special cases results for a large class of problems that have wide application. The unobserved state sequence is assumed to be Markovian with a general order of dependence. An auxiliary Markov chain is associated with the state sequence and is used to simplify the computations. Two examples are given to illustrate the use of the methodology. Whereas the first application is more to illustrate the basic steps in applying the theory, the second is a more detailed application to DNA sequences, and shows that the methods can be adapted to include restrictions related to biological knowledge.

stat.ME↗