SearcharxivSearch

arXiv subjects

Kiyoshi Asai

Publications and source records attributed to Kiyoshi Asai.

3 recordsLinked to original sources

From Derivatives to Exact Sequence Substitution Effects in Dynamic Programming for Biological Sequence Analysis

Background: Dynamic programming in biological sequence analysis computes probabilities or partition functions by summing over exponentially many latent paths, alignments, derivation trees, or RNA secondary structures. Their backward and outside quantities are used model-specifically, but the relation between differential sensitivities and exact finite sequence changes is rarely stated in a common framework. Methods: We represent hidden Markov models, affine-gap alignment ensembles, stochastic context-free grammars, and RNA secondary-structure ensembles as sum--product dynamic programs, defining backward and outside quantities as adjoints of forward or inside variables and sequence changes as finite replacements of sequence-dependent local factors. Results: Posterior item marginals are normalized inside--outside products, local-event posteriors additionally include the local factor and child inside terms, and expected feature counts are logarithmic derivatives of the partition function. For HMMs, ordinary SCFGs, and single-position substitutions in affine-gap alignment, the partition function is multi-affine in position-specific factor groups, so a one-site change is recovered exactly from first-derivative coefficients and multisite changes from mixed derivatives. In nearest-neighbor RNA models a substitution alters overlapping loop, stacking, and multiloop factors and boundary contexts, so exact mutation effects instead require context-dependent inside--outside recombination, as in the Rchange algorithm. Numerical experiments reproduce brute-force recomputation to machine precision. Conclusions: The framework identifies when derivatives give exact finite sequence effects and when broader recombination is required, providing a unified basis for posterior marginals, expected counts, parameter sensitivity, mutation analysis, and sequence design.

q-bio.QM

Generalized Centroid Estimators in Bioinformatics

In a number of estimation problems in bioinformatics, accuracy measures of the target problem are usually given, and it is important to design estimators that are suitable to those accuracy measures. However, there is often a discrepancy between an employed estimator and a given accuracy measure of the problem. In this study, we introduce a general class of efficient estimators for estimation problems on high-dimensional binary spaces, which representmany fundamental problems in bioinformatics. Theoretical analysis reveals that the proposed estimators generally fit with commonly-used accuracy measures (e.g. sensitivity, PPV, MCC and F-score) as well as it can be computed efficiently in many cases, and cover a wide range of problems in bioinformatics from the viewpoint of the principle of maximum expected accuracy (MEA). It is also shown that some important algorithms in bioinformatics can be interpreted in a unified manner. Not only the concept presented in this paper gives a useful framework to design MEA-based estimators but also it is highly extendable and sheds new light on many problems in bioinformatics.

q-bio.QM

Approximating Incomplete Kernel Matrices by the em Algorithm

In biological data, it is often the case that observed data are available only for a subset of samples. When a kernel matrix is derived from such data, we have to leave the entries for unavailable samples as missing. In this paper, we make use of a parametric model of kernel matrices, and estimate missing entries by fitting the model to existing entries. The parametric model is created as a set of spectral variants of a complete kernel matrix derived from another information source. For model fitting, we adopt the em algorithm based on the information geometry of positive definite matrices. We will report promising results on bacteria clustering experiments using two marker sequences: 16S and gyrB.

cs.LG