SearcharxivSearch

arXiv subjects

Eitan Yaakobi

Publications and source records attributed to Eitan Yaakobi.

At least 19 recordsLinked to original sources

Efficient Synthesis for Two-Dimensional Strand Arrays with Row Constraints

In large-scale array-based DNA synthesis, optical and chemical coupling between nearby sites can limit simultaneous activations. Motivated by this constraint, we study strands synthesized according to a fixed global synthesis sequence, with at most one strand per row advancing in each cycle. We focus on the fundamental case of two strands in a single row and analyze the expected completion time of row-constrained synthesis. We introduce the laggard-first (LF) policy, a simple rule that always advances the strand with fewer synthesized symbols when a conflict arises, and establish that it is asymptotically optimal among online policies without look-ahead. In the binary case, one-symbol look-ahead strictly improves on the no-look-ahead bound. We further show that even complete advance knowledge does not eliminate the scheduling loss, as even a globally optimal schedule incurs an unavoidable expected overhead that grows linearly with the strand length. Finally, we complement these scheduling results with a dynamic programming algorithm for computing an optimal offline synthesis order and a constant-redundancy binary coding scheme that yields a deterministic worst-case synthesis time guarantee.

cs.IT

Correcting Tail Deletions in Rank Modulated Composite Encoding for Data Storage in DNA

We study the combination of two recent coding approaches, in the context of DNA based data storage. Composite DNA alphabets leverage properties of the DNA synthesis and sequencing process. A composite symbol does not represent a single nucleotide, but rather a designed mixture of DNA nucleotides. Using the high multiplicity that is intrinsic to synthesis and sequencing a composite symbol consists of frequencies in the mixture. Rank modulation codes use permutations to represent information. Combining the two, we construct encoding that uses permutations of nucleotide frequencies rather than the exact frequency values. Codes for this approach were addressed in previous work, under Kendall's tau distances. In this work we study deletion and insertion codes. We present bounds and constructions of efficient codes defined over partial permutations.

cs.IT

Random Access in DNA Storage: Algorithms, Constructions, and Bounds

As DNA data storage advances toward practical deployment, minimizing sequencing coverage depth is critical for reducing operational costs and retrieval latency. We study the random access problem of recovering a specific information strand from a DNA-based storage system. In this setting, $k$ information strands are encoded into $n$ strands using a generator matrix $G$, and each sequencing read returns one encoded strand sampled uniformly at random with replacement. We derive an exact formula for the expected number of samples required to recover a specific information strand, yielding an $O(n)$-time algorithm for fixed field size $q$ and dimension $k$. We further obtain explicit formulas for the average and maximum expected number of samples, enabling an efficient search for optimal generator matrices for small parameters. We present new constructions that improve the best-known upper bounds from $0.8815k$ to $0.8811k$ for $k=3$, and from $0.8637k$ to $0.8629k$ for $k=4$, for sufficiently large $q$. We also establish a tighter lower bound on the expected number of samples, which in particular proves the optimality of the simple parity code when $n=k+1$ over any field size $q$. Finally, for the non-random access setting, we derive new lower bounds and constructions that characterize the asymptotic behavior of the expected number of samples required to recover all information strands.

cs.IT

Error characterization and error correction approaches in combinatorial DNA-based storage

Data storage in DNA has recently emerged as a promising archival solution, offering space-efficient and long-lasting digital storage. Combinatorial DNA encoding enhances this potential by increasing the logical density through combinations of DNA shortmers, where each sequence position is represented by a set of predefined short DNA fragments, allowing more data to be encoded using fewer synthesis cycles. However, this method introduces unique synthesis and sequencing errors. In this study, we characterize errors in combinatorial DNA-based storage systems. We reveal that asymmetric combinatorial erasure errors, defined as the omission of a single shortmer from the set defining the combinatorial letter, are a prevalent error type, particularly in large-scale systems where read coverage is limited. In two previously published datasets, we observed a high frequency of erasure errors, where missing sequences obstruct the reconstruction of combinatorial letters. We conducted a large-scale experimental proof-of-concept and confirmed that erasure errors become increasingly prominent with reduced sequencing depth: below 50 reads per sequence, their frequency sharply increased. We developed an asymmetric error-correcting code for these errors, utilizing tensor-product codes to integrate standard erasure and substitution-correcting codes (such as Reed-Solomon (RS) codes) with asymmetric Varshamov-Tenengolts codes. We validated its performance in simulations and in a second large-scale experiment directly comparing it with the more straightforward 2D RS scheme. Our method consistently outperformed 2D RS, particularly in erasure-dominated scenarios, and demonstrated superior decoding accuracy under low coverage conditions where 2D RS struggled to decode the data. Our findings demonstrate the importance of error correction schemes tailored to the asymmetric nature of errors in combinatorial DNA.

cs.IT

More on Codes for Combinatorial Composite DNA

In this paper, we focus on constructions of unique-decodable/list-decodable on the recently studied $(t,e)$-composite-asymmetric error-correcting codes ($(t,e)$-CAECCs). Let $X$ be an $m\times n$ binary matrix, in which each row has Hamming weight $w$. When at most $t$ rows of $X$ suffer from errors and in each of these erroneous rows, there are at most $e$ $1 \to 0$ errors, we say that a $(t,e)$-composite-asymmetric-error occurs in $X$. For general $m,n,w,t,e$, we propose new constructions of $(t,e)$-CAECCs with redundancy at most $(t-1)\log(m)+O(1)$, where $O(1)$ is a number independent of the code-length $m$. In particular, this gives a class of $(2,e)$-CAECCs that are optimal in terms of their redundancy. %(in terms of redundancy, regarded as a function of the number of rows $m$) s. When $m$ is a prime power, the redundancy can be further reduced to $(t-1)\log(m)-O(\log(m))$. To further increase the size of these codes, we introduce a combinatorial object called a weak $B_e$-sets. When $e=w$, we show an efficient way to encode/decode our codes. At last, we investigate how much we can gain if we relax the requirement of uniquely decoding to list-decoding. It is shown that when the list size is $t!$ or an exponential function of $t$, there are list-decodable $(t,e)$-CAECCs with constant redundancy. When the list size is two, we show that there are list-decodable $(3,2)$-CAECCs with redundancy $\log(m)+O(1)$.

cs.IT

Bounds on Multiple $b$-Burst Deletion-Correcting Codes

Motivated by their applications in DNA-based storage systems, codes capable of correcting consecutive deletions have attracted significant attention. An important class of such codes consists of those that can correct multiple consecutive deletion errors, commonly referred to as multiple $b$-burst deletion-correcting codes. In this paper, we investigate the fundamental limits of multiple $b$-burst deletion-correcting codes. Specifically, we first characterize several structural properties of the associated deletion balls. Then, leveraging these properties, we derive several upper bounds and a combinatorial lower bound on the maximum size of such codes. As a consequence, our bounds improve upon the previously known results for general parameter regimes and are shown to be asymptotically optimal for certain cases.

cs.IT

Sequence Reconstruction for Substitution Channel: New Sufficient Conditions and Algorithms

In the sequence reconstruction problem, a codeword $\x$ is transmitted through several identical channels where each channel produces a noisy read of $\x$, and the problem is to analyze how to uniquely reconstruct $\x$ based on these noisy reads. Levenshtein has studied the minimum number of reads which guarantees unique reconstruction of $\x$, which is one sufficient condition for unique reconstruction. In this paper, we move on to a different perspective and propose a new framework for unique reconstruction. Our new sufficient condition for unique reconstruction takes both the number of reads and the distances among the reads into consideration. We offer both theoretical analysis and corresponding efficient reconstruction algorithms for our reconstruction framework.

cs.IT

Rank Modulated Composite Encoding for Data Storage in DNA

This paper studies two problems that are motivated by combining two novel approaches, namely DNA composite and rank modulation. The recent approach of composite DNA takes advantage of the DNA synthesis property which generates a huge number of copies for every synthesized strand. Under this paradigm, every composite symbols does not store a single nucleotide but a mixture of the four DNA nucleotides. Instead of considering all the possible composite symbols we are interested only in the rank of the motifs in the symbol. The first problem in this paper addresses the capacity of a channel that uses such symbols, while in the second, bounds and construction of such codes are studied.

cs.IT

Random Access Expectation in DNA Storage and Fountain Codes

Motivated by DNA data storage, we study the expected number of coded symbols drawn from a linear code until a desired information symbol can be decoded - the random access expectation. We focus on generator matrices with a type of symmetry, conjectured in prior work to be optimal, which we call fully symmetric. We point out an equivalence between binary fully symmetric codes and LT codes. Using this observation, we analyze the random access expectation of binary fully symmetric codes under a peeling decoder, in the large blocklength limit. Under these assumptions, the random access expectation, normalized by the number of information symbols, is at least $π/4 \approx 0.7854$, while a value of $\approx 0.7869$ is achievable.

cs.IT

Serving Every Symbol: All-Symbol PIR and Batch Codes

A $t$-all-symbol PIR code and a $t$-all-symbol batch code of dimension $k$ consist of $n$ servers storing linear combinations of $k$ information symbols with the following recovery property: any symbol stored by a server can be recovered from $t$ pairwise disjoint subsets of servers. In the batch setting, we further require that any multiset of size $t$ of stored symbols can be recovered from~$t$ disjoint subsets of servers. This framework unifies and extends several well-known code families, including one-step majority-logic decodable codes, (functional) PIR codes, and (functional) batch codes. In this paper, we determine the minimum code length for some small values of $k$ and $t$, characterize structural properties of codes attaining this optimum, and derive bounds that show the trade-offs between length, dimension, minimum distance, and $t$. In addition, we study MDS codes and the simplex code, demonstrating how these classical families fit within our framework, and establish new cases of an open conjecture from \cite{YAAKOBI2020} concerning the minimal $t$ for which the simplex code is a $t$-functional batch code.

cs.IT

Error-Correcting Codes for the Sum Channel

We introduce the sum channel, a new channel model motivated by applications in distributed storage and DNA data storage. In the error-free case, it takes as input an $\ell$-row binary matrix and outputs an $(\ell+1)$-row matrix whose first $\ell$ rows equal the input and whose last row is their parity (sum) row. We construct a two-deletion-correcting code with redundancy $2\lceil\log_2\log_2 n\rceil + O(\ell^2)$ for $\ell$-row inputs. When $\ell=2$, we establish an upper bound of $\lceil\log_2\log_2 n\rceil + O(1)$, implying that our redundancy is optimal up to a factor of 2. We also present a code correcting a single substitution with $\lceil \log_2(\ell+1)\rceil$ redundant bits and prove that it is within one bit of optimality.

cs.IT

The DNA Coverage Depth Problem: Duality, Weight Distributions, and Applications

The coverage depth problem in DNA data storage is about computing the expected number of reads needed to recover all encoded strands. Given a generator matrix of a linear code, this quantity equals the expected number of randomly drawn columns required to obtain full rank. While MDS codes are optimal when they exist, i.e., over large fields, practical scenarios may rely on structured code families defined over small fields. In this work, we develop combinatorial tools to solve the DNA coverage depth problem for various linear codes, based on duality arguments and the notion of extended weight enumerator. Using these methods, we derive closed formulas for the simplex, Hamming, ternary Golay, extended ternary Golay, and first-order Reed-Muller codes. The centerpiece of this paper is a general expression for the coverage depth of a linear code in terms of the weight distributions of its higher-field extensions.

cs.IT

Expected Recovery Time in DNA-based Distributed Storage Systems

We initiate the study of DNA-based distributed storage systems, where information is encoded across multiple DNA data storage containers to achieve robustness against container failures. In this setting, data are distributed over $M$ containers, and the objective is to guarantee that the contents of any failed container can be reliably reconstructed from the surviving ones. Unlike classical distributed storage systems, DNA data storage containers are fundamentally constrained by sequencing technology, since each read operation yields the content of a uniformly random sampled strand from the container. Within this framework, we consider several erasure-correcting codes and analyze the expected recovery time of the data stored in a failed container. Our results are obtained by analyzing generalized versions of the classical Coupon Collector's Problem, which may be of independent interest.

cs.IT

Analyzing Collection Strategies: A Computational Perspective on the Coupon Collector Problem

The Coupon Collector Problem (CCP) is a well-known combinatorial problem that seeks to estimate the number of random draws required to complete a collection of $n$ distinct coupon types. Various generalizations of this problem have been applied in numerous engineering domains. However, practical applications are often hindered by the computational challenges associated with deriving numerical results for moments and distributions. In this work, we present three algorithms for solving the most general form of the CCP, where coupons are collected under any arbitrary drawing probability, with the objective of obtaining $t$ copies of a subset of $k$ coupons from a total of $n$. The First algorithm provides the base model to compute the expectation, variance, and the second moment of the collection process. The second algorithm utilizes the construction of the base model and computes the same values in polynomial time with respect to $n$ under the uniform drawing distribution, and the third algorithm extends to any general drawing distribution. All algorithms leverage Markov models specifically designed to address computational challenges, ensuring exact computation of the expectation and variance of the collection process. Their implementation uses a dynamic programming approach that follows from the Markov models framework, and their time complexity is analyzed accordingly.

cs.DS

Reconstructing Reed-Solomon Codes from Multiple Noisy Channel Outputs

The sequence reconstruction problem, introduced by Levenshtein in 2001, considers a communication setting in which a sender transmits a codeword and the receiver observes K independent noisy versions of this codeword. In this work, we study the problem of efficient reconstruction when each of the $K$ outputs is corrupted by a $q$-ary discrete memoryless symmetric (DMS) substitution channel with substitution probability $p$. Focusing on Reed-Solomon (RS) codes, we adapt the Koetter-Vardy soft-decision decoding algorithm to obtain an efficient reconstruction algorithm. For sufficiently large blocklength and alphabet size, we derive an explicit rate threshold, depending only on $(p, K)$, such that the transmitted codeword can be reconstructed with arbitrarily small probability of error whenever the code rate $R$ lies below this threshold.

cs.IT

Error-Correcting Codes for Labeled DNA Sequences

Labeling of DNA molecules is a fundamental technique for DNA visualization and analysis. This process was mathematically modeled in [1], where the received sequence indicates the positions of the used labels. In this work, we develop error correcting codes for labeled DNA sequences, establishing bounds and constructing explicit systematic encoders for single substitution, insertion, and deletion errors. We focus on two cases: (1) using the complete set of length-two labels and (2) using the minimal set of length-two labels that ensures the recovery of DNA sequences from their labeling for 'almost' all DNA sequences.

cs.IT

Coding for Ordered Composite DNA Sequences

To increase the information capacity of DNA storage, composite DNA letters were introduced. We propose a novel channel model for composite DNA in which composite sequences are decomposed into ordered standard non-composite sequences. The model is designed to handle any alphabet size and composite resolution parameter. We study the problem of reconstructing composite sequences of arbitrary resolution over the binary alphabet under substitution errors. We define two families of error-correcting codes and provide lower and upper bounds on their cardinality. In addition, we analyze the case in which a single deletion error occurs in the channel and present a systematic code construction for this setting. Finally, we briefly discuss the channel's capacity, which remains an open problem.

cs.IT

General Coverage Models: Structure, Monotonicity, and Shotgun Sequencing

We study coverage processes in which each draw reveals a subset of $[n]$, and the goal is to determine the expected number of draws until all items are seen at least once. A classical example is the Coupon Collector's Problem, where each draw reveals exactly one item. Motivated by shotgun DNA sequencing, we introduce a model where each draw is a contiguous window of fixed length, in both cyclic and non-cyclic variants. We develop a unifying combinatorial tool that shifts the task of finding coverage time from probability, to a counting problem over families of subsets of $[n]$ that together contain all items, enabling exact calculation. Using this result, we obtain exact expressions for the window models. We then leverage past results on a continuous analogue of the cyclic window model to analyze the asymptotic behavior of both models. We further study what we call uniform $\ell$-regular models, where every draw has size $\ell$ and every item appears in the same number of admissible draws. We compare these to the batch sampling model, in which all $\ell$-subsets are drawn uniformly at random and present upper and lower bounds, which were also obtained independently by Berend and Sher. We conjecture, and prove for special cases, that this model maximizes the coverage time among all uniform $\ell$-regular models. Finally, we prove a universal upper bound on the entire class of uniform $\ell$-regular models, which illuminates the fact that many sampling models share the same leading asymptotic order, while potentially differing significantly in lower-order terms.

cs.IT