arXiv · 2601.07053
Random Access in DNA Storage: Algorithms, Constructions, and Bounds
Abstract
As DNA data storage advances toward practical deployment, minimizing sequencing coverage depth is critical for reducing operational costs and retrieval latency. We study the random access problem of recovering a specific information strand from a DNA-based storage system. In this setting, $k$ information strands are encoded into $n$ strands using a generator matrix $G$, and each sequencing read returns one encoded strand sampled uniformly at random with replacement. We derive an exact formula for the expected number of samples required to recover a specific information strand, yielding an $O(n)$-time algorithm for fixed field size $q$ and dimension $k$. We further obtain explicit formulas for the average and maximum expected number of samples, enabling an efficient search for optimal generator matrices for small parameters. We present new constructions that improve the best-known upper bounds from $0.8815k$ to $0.8811k$ for $k=3$, and from $0.8637k$ to $0.8629k$ for $k=4$, for sufficiently large $q$. We also establish a tighter lower bound on the expected number of samples, which in particular proves the optimality of the simple parity code when $n=k+1$ over any field size $q$. Finally, for the non-random access setting, we derive new lower bounds and constructions that characterize the asymptotic behavior of the expected number of samples required to recover all information strands.
Explore related subjects
Keep this discovery
Chen Wang, Eitan Yaakobi. 2026-01-11. Random Access in DNA Storage: Algorithms, Constructions, and Bounds. https://arxiv.org/abs/2601.07053
Cite the original work for its findings. Save a collection to share your selection of sources.