arXiv · 1910.04640
E2FM: an encrypted and compressed full-text index for collections of genomic sequences
Abstract
Next Generation Sequencing (NGS) platforms and, more generally, high-throughput technologies are giving rise to an exponential growth in the size of nucleotide sequence databases. Moreover, many emerging applications of nucleotide datasets -- as those related to personalized medicine -- require the compliance with regulations about the storage and processing of sensitive data. We have designed and carefully engineered E2FM-index, a new full-text index in minute space which was optimized for compressing and encrypting nucleotide sequence collections in FASTA format and for performing fast pattern-search queries. E2FM-index allows to build self-indexes which occupy till to 1/20 of the storage required by the input FASTA file, thus permitting to save about 95% of storage when indexing collections of highly similar sequences; moreover, it can exactly search the built indexes for patterns in times ranging from few milliseconds to a few hundreds milliseconds, depending on pattern length. Supplementary material and supporting datasets are available through Bioinformatics Online and https://figshare.com/s/6246ee9c1bd730a8bf6e.
Explore related subjects
Keep this discovery
Ferdinando Montecuollo, Giovannni Schmid, Roberto Tagliaferri. 2019-10-10. E2FM: an encrypted and compressed full-text index for collections of genomic sequences. https://doi.org/10.1093/bioinformatics%2Fbtx313
Cite the original work for its findings. Save a collection to share your selection of sources.