SearcharxivSearch

arXiv subjects

German Tischler

Publications and source records attributed to German Tischler.

3 recordsLinked to original sources

Faster Average Case Low Memory Semi-External Construction of the Burrows-Wheeler Transform

The Burrows Wheeler transform has applications in data compression as well as full text indexing. Despite its important applications and various existing algorithmic approaches the construction of the transform for large data sets is still challenging. In this paper we present a new semi external memory algorithm for constructing the Burrows Wheeler transform. It is capable of constructing the transform for an input text of length $n$ over a finite alphabet in time $O(n\log^2\log n)$ on average, if sufficient internal memory is available to hold a fixed fraction of the input text. In the worst case the run-time is $O(n\log n \log\log n)$. The amount of space used by the algorithm in external memory is $O(n)$ bits. Based on the serial version we also present a shared memory parallel algorithm running in time $O(\frac{n}{p}\max\{\log^2\log n+\log p\})$ on average when $p$ processors are available.

cs.DS

Low Space External Memory Construction of the Succinct Permuted Longest Common Prefix Array

The longest common prefix (LCP) array is a versatile auxiliary data structure in indexed string matching. It can be used to speed up searching using the suffix array (SA) and provides an implicit representation of the topology of an underlying suffix tree. The LCP array of a string of length $n$ can be represented as an array of length $n$ words, or, in the presence of the SA, as a bit vector of $2n$ bits plus asymptotically negligible support data structures. External memory construction algorithms for the LCP array have been proposed, but those proposed so far have a space requirement of $O(n)$ words (i.e. $O(n \log n)$ bits) in external memory. This space requirement is in some practical cases prohibitively expensive. We present an external memory algorithm for constructing the $2n$ bit version of the LCP array which uses $O(n \log σ)$ bits of additional space in external memory when given a (compressed) BWT with alphabet size $σ$ and a sampled inverse suffix array at sampling rate $O(\log n)$. This is often a significant space gain in practice where $σ$ is usually much smaller than $n$ or even constant. We also consider the case of computing succinct LCP arrays for circular strings.

cs.DS

biobambam: tools for read pair collation based algorithms on BAM files

Sequence alignment data is often ordered by coordinate (id of the reference sequence plus position on the sequence where the fragment was mapped) when stored in BAM files, as this simplifies the extraction of variants between the mapped data and the reference or of variants within the mapped data. In this order paired reads are usually separated in the file, which complicates some other applications like duplicate marking or conversion to the FastQ format which require to access the full information of the pairs. In this paper we introduce biobambam, an API for efficient BAM file reading supporting the efficient collation of alignments by read name without performing a complete resorting of the input file and some tools based on this API performing tasks like marking duplicate reads and conversion to the FastQ format. In comparison with previous approaches to problems involving the collation of alignments by read name like the BAM to FastQ or duplication marking utilities in the Picard suite the approach of biobambam can often perform an equivalent task more efficiently in terms of the required main memory and run-time.

q-bio.GN