Searcharxiv⌕ Search

arXiv subjects

Ahsan Sanaullah

Publications and source records attributed to Ahsan Sanaullah.

2 recordsLinked to original sources

Optimal-Time Move Structure Construction

The move structure represents a permutation $π$ of $[0,n)$ by partitioning the domain into $O(r)$ disjoint, contiguously permuted intervals, with $r$ being the minimum number of such intervals. This data structure occupies $O(r)$ words of space and enables $O(1)$-time computation of $π(i)$ given the interval that contains $i$. For permutations where $r \ll n$, this provides an efficient, compressed representation for navigation. While existing best $O(r)$-space construction approaches require $O(r\log r)$-time, we present an optimal $O(r)$-time and space construction algorithm. This is achieved by replacing balanced search trees with pointer-based lists by introducing a bidirectional strategy that synchronizes construction of the structures for $π$ and its inverse $π^{-1}$ in a single, unified pass. By applying this algorithm, we achieve the first optimal $O(n)$-time construction of the longest common prefix (LCP) array from a run-length-encoded Burrows-Wheeler transform (RLBWT) of $r$ runs in $O(r)$ working space. Empirical evaluation on pangenome-scale data confirms that our move structure construction algorithm is consistently faster than the previous best, achieving speedups of up to $\sim 2\times$ with comparable memory usage.

cs.DS↗

An Efficient Data Structure and Algorithm for Long-Match Query in Run-Length Compressed BWT

In this paper, we describe a new type of match between a pattern and a text that aren't necessarily maximal in the query, but still contain useful matching information: locally maximal exact matches (LEMs). There are usually a large amount of LEMs, so we only consider those above some length threshold $\mathcal{L}$. These are referred to as long LEMs. The purpose of long LEMs is to capture substring matches between a query and a text that are not necessarily maximal in the pattern but still long enough to be important. Therefore efficient long LEMs finding algorithms are desired for these datasets. However, these datasets are too large to query on traditional string indexes. Fortunately, these datasets are very repetitive. Recently, compressed string indexes that take advantage of the redundancy in the data but retain efficient querying capability have been proposed as a solution. We therefore give an efficient algorithm for computing all the long LEMs of a query and a text in a BWT runs compressed string index. We describe an $O(m+occ)$ expected time algorithm that relies on an $O(r)$ words space string index for outputting all long LEMs of a pattern with respect to a text given the matching statistics of the pattern with respect to the text. Here $m$ is the length of the query, $occ$ is the number of long LEMs outputted, and $r$ is the number of runs in the BWT of the text. The $O(r)$ space string index we describe relies on an adaptation of the move data structure by Nishimoto and Tabei. We are able to support $LCP[i]$ queries in constant time given $SA[i]$. In other words, we answer $PLCP[i]$ queries in constant time. Long LEMs may provide useful similarity information between a pattern and a text that MEMs may ignore. This information is particularly useful in pangenome and biobank scale haplotype panel contexts.

cs.DS↗