SearcharxivSearch

arXiv subjects

Omer Sabary

Publications and source records attributed to Omer Sabary.

10 recordsLinked to original sources

Error characterization and error correction approaches in combinatorial DNA-based storage

Data storage in DNA has recently emerged as a promising archival solution, offering space-efficient and long-lasting digital storage. Combinatorial DNA encoding enhances this potential by increasing the logical density through combinations of DNA shortmers, where each sequence position is represented by a set of predefined short DNA fragments, allowing more data to be encoded using fewer synthesis cycles. However, this method introduces unique synthesis and sequencing errors. In this study, we characterize errors in combinatorial DNA-based storage systems. We reveal that asymmetric combinatorial erasure errors, defined as the omission of a single shortmer from the set defining the combinatorial letter, are a prevalent error type, particularly in large-scale systems where read coverage is limited. In two previously published datasets, we observed a high frequency of erasure errors, where missing sequences obstruct the reconstruction of combinatorial letters. We conducted a large-scale experimental proof-of-concept and confirmed that erasure errors become increasingly prominent with reduced sequencing depth: below 50 reads per sequence, their frequency sharply increased. We developed an asymmetric error-correcting code for these errors, utilizing tensor-product codes to integrate standard erasure and substitution-correcting codes (such as Reed-Solomon (RS) codes) with asymmetric Varshamov-Tenengolts codes. We validated its performance in simulations and in a second large-scale experiment directly comparing it with the more straightforward 2D RS scheme. Our method consistently outperformed 2D RS, particularly in erasure-dominated scenarios, and demonstrated superior decoding accuracy under low coverage conditions where 2D RS struggled to decode the data. Our findings demonstrate the importance of error correction schemes tailored to the asymmetric nature of errors in combinatorial DNA.

cs.IT

More on Codes for Combinatorial Composite DNA

In this paper, we focus on constructions of unique-decodable/list-decodable on the recently studied $(t,e)$-composite-asymmetric error-correcting codes ($(t,e)$-CAECCs). Let $X$ be an $m\times n$ binary matrix, in which each row has Hamming weight $w$. When at most $t$ rows of $X$ suffer from errors and in each of these erroneous rows, there are at most $e$ $1 \to 0$ errors, we say that a $(t,e)$-composite-asymmetric-error occurs in $X$. For general $m,n,w,t,e$, we propose new constructions of $(t,e)$-CAECCs with redundancy at most $(t-1)\log(m)+O(1)$, where $O(1)$ is a number independent of the code-length $m$. In particular, this gives a class of $(2,e)$-CAECCs that are optimal in terms of their redundancy. %(in terms of redundancy, regarded as a function of the number of rows $m$) s. When $m$ is a prime power, the redundancy can be further reduced to $(t-1)\log(m)-O(\log(m))$. To further increase the size of these codes, we introduce a combinatorial object called a weak $B_e$-sets. When $e=w$, we show an efficient way to encode/decode our codes. At last, we investigate how much we can gain if we relax the requirement of uniquely decoding to list-decoding. It is shown that when the list size is $t!$ or an exponential function of $t$, there are list-decodable $(t,e)$-CAECCs with constant redundancy. When the list size is two, we show that there are list-decodable $(3,2)$-CAECCs with redundancy $\log(m)+O(1)$.

cs.IT

On The Decoding Error Weight of One or Two Deletion Channels

This paper tackles two problems that fall under the study of coding for insertions and deletions. These problems are motivated by several applications, among them is reconstructing strands in DNA-based storage systems. Under this paradigm, a word is transmitted over some fixed number of identical independent channels and the goal of the decoder is to output the transmitted word or some close approximation of it. The first part of the paper studies optimal decoding for a special case of the deletion channel, referred by the $k$-deletion channel, which deletes exactly $k$ symbols of the transmitted word uniformly at random. In this part, the goal is to understand how an optimal decoder operates in order to minimize the expected normalized distance. A full characterization of an efficient optimal decoder for this setup, referred to as the maximum likelihood* (ML*) decoder, is given for a channel that deletes one or two symbols. The second part of this paper studies the deletion channel that deletes a symbol with some fixed probability $p$, while focusing on two instances of this channel. Since operating the maximum likelihood (ML) decoder, in this case, is computationally unfeasible, we study a slightly degraded version of this decoder for two channels and study its expected normalized distance. We observe that the dominant error patterns are deletions in the same run or errors resulting from alternating sequences. Based on these observations, we derive lower bounds on the expected normalized distance of the degraded ML decoder for any transmitted $q$-ary sequence of length $n$ and any deletion probability $p$. We further show that as the word length approaches infinity and the channel's deletion probability $p$ approaches zero, these bounds converge to approximately $\frac{3q - 1}{q - 1} p^2$. These theoretical results are verified by corresponding simulations.

cs.IT

Conditional Entropies of k-Deletion/Insertion Channels

The channel output entropy of a transmitted sequence is the entropy of the possible channel outputs and similarly the channel input entropy of a received sequence is the entropy of all possible transmitted sequences. The goal of this work is to study these entropy values for the k-deletion, k-insertion channels, where exactly k symbols are deleted, inserted in the transmitted sequence, respectively. If all possible sequences are transmitted with the same probability then studying the input and output entropies is equivalent. For both the 1-deletion and 1-insertion channels, it is proved that among all sequences with a fixed number of runs, the input entropy is minimized for sequences with a skewed distribution of their run lengths and it is maximized for sequences with a balanced distribution of their run lengths. Among our results, we establish a conjecture by Atashpendar et al. which claims that for the 1-deletion channel, the input entropy is maximized by the alternating sequences over all binary sequences. This conjecture is also verified for the 2-deletion channel, where it is proved that constant sequences with a single run minimize the input entropy.

cs.IT

Error-Correcting Codes for Combinatorial Composite DNA

Data storage in DNA is developing as a possible solution for archival digital data. Recently, to further increase the potential capacity of DNA-based data storage systems, the combinatorial composite DNA synthesis method was suggested. This approach extends the DNA alphabet by harnessing short DNA fragment reagents, known as shortmers. The shortmers are building blocks of the alphabet symbols, consisting of a fixed number of shortmers. Thus, when information is read, it is possible that one of the shortmers that forms part of the composition of a symbol is missing and therefore the symbol cannot be determined. In this paper, we model this type of error as a type of asymmetric error and propose code constructions that can correct such errors in this setup. We also provide a lower bound on the redundancy of such error-correcting codes and give an explicit encoder and decoder pair for our construction. Our suggested error model is also supported by an analysis of data from actual experiments that produced DNA according to the combinatorial scheme. Lastly, we also provide a statistical evaluation of the probability of observing such error events, as a function of read depth.

cs.IT

Coding for Composite DNA to Correct Substitutions, Strand Losses, and Deletions

Composite DNA is a recent method to increase the base alphabet size in DNA-based data storage.This paper models synthesizing and sequencing of composite DNA and introduces coding techniques to correct substitutions, losses of entire strands, and symbol deletion errors. Non-asymptotic upper bounds on the size of codes with $t$ occurrences of these error types are derived. Explicit constructions are presented which can achieve the bounds.

cs.IT

Deep DNA Storage: Scalable and Robust DNA Storage via Coding Theory and Deep Learning

DNA-based storage is an emerging technology that enables digital information to be archived in DNA molecules. This method enjoys major advantages over magnetic and optical storage solutions such as exceptional information density, enhanced data durability, and negligible power consumption to maintain data integrity. To access the data, an information retrieval process is employed, where some of the main bottlenecks are the scalability and accuracy, which have a natural tradeoff between the two. Here we show a modular and holistic approach that combines Deep Neural Networks (DNN) trained on simulated data, Tensor-Product (TP) based Error-Correcting Codes (ECC), and a safety margin mechanism into a single coherent pipeline. We demonstrated our solution on 3.1MB of information using two different sequencing technologies. Our work improves upon the current leading solutions by up to x3200 increase in speed, 40% improvement in accuracy, and offers a code rate of 1.6 bits per base in a high noise regime. In a broader sense, our work shows a viable path to commercial DNA storage solutions hindered by current information retrieval processes.

cs.IT

Cover Your Bases: How to Minimize the Sequencing Coverage in DNA Storage Systems

Although the expenses associated with DNA sequencing have been rapidly decreasing, the current cost of sequencing information stands at roughly $120/GB, which is dramatically more expensive than reading from existing archival storage solutions today. In this work, we aim to reduce not only the cost but also the latency of DNA storage by initiating the study of the DNA coverage depth problem, which aims to reduce the required number of reads to retrieve information from the storage system. Under this framework, our main goal is to understand the effect of error-correcting codes and retrieval algorithms on the required sequencing coverage depth. We establish that the expected number of reads that are required for information retrieval is minimized when the channel follows a uniform distribution. We also derive upper and lower bounds on the probability distribution of this number of required reads and provide a comprehensive upper and lower bound on its expected value. We further prove that for a noiseless channel and uniform distribution, MDS codes are optimal in terms of minimizing the expected number of reads. Additionally, we study the DNA coverage depth problem under the random-access setup, in which the user aims to retrieve just a specific information unit from the entire DNA storage system. We prove that the expected retrieval time is at least k for [n,k] MDS codes as well as for other families of codes. Furthermore, we present explicit code constructions that achieve expected retrieval times below k and evaluate their performance through analytical methods and simulations. Lastly, we provide lower bounds on the maximum expected retrieval time. Our findings offer valuable insights for reducing the cost and latency of DNA storage.

cs.DM

The Input and Output Entropies of the $k$-Deletion/Insertion Channel

The channel output entropy of a transmitted word is the entropy of the possible channel outputs and similarly, the input entropy of a received word is the entropy of all possible transmitted words. The goal of this work is to study these entropy values for the k-deletion, k-insertion channel, where exactly k symbols are deleted, and inserted in the transmitted word, respectively. If all possible words are transmitted with the same probability then studying the input and output entropies is equivalent. For both the 1-insertion and 1-deletion channels, it is proved that among all words with a fixed number of runs, the input entropy is minimized for words with a skewed distribution of their run lengths and it is maximized for words with a balanced distribution of their run lengths. Among our results, we establish a conjecture by Atashpendar et al. which claims that for the binary 1-deletion, the input entropy is maximized for the alternating words. This conjecture is also verified for the 2-deletion channel, where it is proved that constant words with a single run minimize the input entropy.

cs.IT

The Error Probability of Maximum-Likelihood Decoding over Two Deletion Channels

This paper studies the problem of reconstructing a word given several of its noisy copies. This setup is motivated by several applications, among them is reconstructing strands in DNA-based storage systems. Under this paradigm, a word is transmitted over some fixed number of identical independent channels and the goal of the decoder is to output the transmitted word or some close approximation. The main focus of this paper is the case of two deletion channels and studying the error probability of the maximum-likelihood (ML) decoder under this setup. First, it is discussed how the ML decoder operates. Then, we observe that the dominant error patterns are deletions in the same run or errors resulting from alternating sequences. Based on these observations, it is derived that the error probability of the ML decoder is roughly $\frac{3q-1}{q-1}p^2$, when the transmitted word is any $q$-ary sequence and $p$ is the channel's deletion probability. We also study the cases when the transmitted word belongs to the Varshamov Tenengolts (VT) code or the shifted VT code. Lastly, the insertion channel is studied as well. These theoretical results are verified by corresponding simulations.

cs.IT