arXiv · 1511.04944
NASCUP: Nucleic Acid Sequence Classification by Universal Probability
Abstract
Motivated by the need for fast and accurate classification of unlabeled nucleotide sequences on a large scale, we developed NASCUP, a new classification method that captures statistical structures of nucleotide sequences by compact context-tree models and universal probability from information theory. NASCUP achieved BLAST-like classification accuracy consistently for several large-scale databases in orders-of-magnitude reduced runtime, and was applied to other bioinformatics tasks such as outlier detection and synthetic sequence generation.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Sunyoung Kwon, Gyuwan Kim, Byunghan Lee, Jongsik Chun, Sungroh Yoon, Young-Han Kim. 2018-11-29. NASCUP: Nucleic Acid Sequence Classification by Universal Probability. https://arxiv.org/abs/1511.04944
Cite the original work for its findings. Save a collection to share your selection of sources.