SearcharxivSearch

arXiv subjects

Ohad Elishco

Publications and source records attributed to Ohad Elishco.

At least 19 recordsLinked to original sources

More on Codes for Combinatorial Composite DNA

In this paper, we focus on constructions of unique-decodable/list-decodable on the recently studied $(t,e)$-composite-asymmetric error-correcting codes ($(t,e)$-CAECCs). Let $X$ be an $m\times n$ binary matrix, in which each row has Hamming weight $w$. When at most $t$ rows of $X$ suffer from errors and in each of these erroneous rows, there are at most $e$ $1 \to 0$ errors, we say that a $(t,e)$-composite-asymmetric-error occurs in $X$. For general $m,n,w,t,e$, we propose new constructions of $(t,e)$-CAECCs with redundancy at most $(t-1)\log(m)+O(1)$, where $O(1)$ is a number independent of the code-length $m$. In particular, this gives a class of $(2,e)$-CAECCs that are optimal in terms of their redundancy. %(in terms of redundancy, regarded as a function of the number of rows $m$) s. When $m$ is a prime power, the redundancy can be further reduced to $(t-1)\log(m)-O(\log(m))$. To further increase the size of these codes, we introduce a combinatorial object called a weak $B_e$-sets. When $e=w$, we show an efficient way to encode/decode our codes. At last, we investigate how much we can gain if we relax the requirement of uniquely decoding to list-decoding. It is shown that when the list size is $t!$ or an exponential function of $t$, there are list-decodable $(t,e)$-CAECCs with constant redundancy. When the list size is two, we show that there are list-decodable $(3,2)$-CAECCs with redundancy $\log(m)+O(1)$.

cs.IT

Quantum Codes with Transversal $CCZ$ Gates and Sublinear $Z$-Stabilizers

We construct asymmetric quantum CSS codes with transversal \(CCZ\) gates from algebraic expander codes \cite{KT26}. For every fixed \(m\ge 3\), our growing-alphabet codes have length \(N\), dimension \(Θ(N)\), and distances \[ d_X=Θ(N), \qquad d_Z=Θ(N^{1/m}). \] Moreover, the \(Z\)-stabilizer space has an explicit generating set of weight \(O(N^{1/m})\). We build on the algebraic puncturing framework of Golowich and Guruswami \cite{GG24}, which turns classical codes with the required Schur-product and distance conditions into CSS codes with transversal \(CCZ\). However, applying the framework directly to the algebraic expander codes runs into their small dual distance, and therefore produces only sublinear dimension. Our main technical step is a refined puncturing theorem in which the global dual-distance assumption is replaced by a condition only on the selected puncturing set. We also reduce the alphabet to a fixed prime field using a projective-multiplicity version of multiplication-friendly codes. The resulting fixed-prime-field CSS code triples, of length \(n\), still have transversal \(CCZ\) gates. Their dimension is \(Θ(n/(\log n)^4)\), with distances \[ d_X=Ω\!\left(\frac{n}{(\log n)^4}\right), \qquad d_Z=Ω\!\left(\frac{n^{1/m}}{(\log n)^{4/m}}\right), \] and the \(Z\)-stabilizer generating set remains sublinear.

cs.IT

Central Limit Theorem for Mutation Systems

DNA-based storage has emerged as a promising alternative to traditional data storage methods, offering unmatched advantages in data density, longevity, and sustainability. Two main approaches have developed: in-vitro storage, where information is synthesized in controlled environments, and in-vivo storage, where data is embedded within an organism's DNA for enhanced confidentiality and protection. While in-vivo DNA storage provides unique advantages, it faces significant challenges from mutations, including duplications, deletions, and substitutions, which cause sequence evolution over time. Thus, in-vivo systems experience continuous sequence alterations that increase length and change composition, making error correction particularly challenging. We study the asymptotic behavior of mutation systems, which model the probabilistic evolution of sequences over a finite alphabet, and are central to the analysis of in-vivo DNA-based data storage. Building upon prior works that established the limit of empirical $k$-tuple frequencies, we characterize the stochastic fluctuations around these values by establishing a Central Limit Theorem (CLT). Our approach leverages the spectral properties of the $k$-substitution matrix to project the centered count vectors, allowing us to approximate the system via a martingale difference sequence, and then verifying the classical martingale CLT conditions. In addition, we explicitly derive the limiting covariance matrix.

cs.IT

Coding for Ordered Composite DNA Sequences

To increase the information capacity of DNA storage, composite DNA letters were introduced. We propose a novel channel model for composite DNA in which composite sequences are decomposed into ordered standard non-composite sequences. The model is designed to handle any alphabet size and composite resolution parameter. We study the problem of reconstructing composite sequences of arbitrary resolution over the binary alphabet under substitution errors. We define two families of error-correcting codes and provide lower and upper bounds on their cardinality. In addition, we analyze the case in which a single deletion error occurs in the channel and present a systematic code construction for this setting. Finally, we briefly discuss the channel's capacity, which remains an open problem.

cs.IT

On Multidimensional 2-Weight-Limited Burst-Correcting Codes

We consider multidimensional codes capable of correcting a burst error of weight at most $2$. When two positions are in error, the burst limits their relative position. We study three such limitations: the $L_\infty$ distance between the positions is bounded, the $L_1$ distance between the positions is bounded, or the two positions are on an axis-parallel line with bounded distance between them. In all cases we provide explicit code constructions, and compare their excess redundancy to a lower bound we prove.

cs.IT

Codes Correcting Two Bursts of Exactly $b$ Deletions

In this paper, we investigate codes designed to correct two bursts of deletions, where each burst has a length of exactly $b$, where $b>1$. The previous best construction, achieved through the syndrome compression technique, had a redundancy of at most $7\log n+O\left(\log n/\log\log n\right)$ bits. In contrast, our work introduces a novel approach for constructing $q$-ary codes that attain a redundancy of at most $5\log n+O(\log\log n)$ bits for all $b>1$ and $q\ge2$. Additionally, for the case where $b=1$, we present a new construction of $q$-ary two-deletion correcting codes with a redundancy of $5\log n+O(\log\log n)$ bits, for all $q>2$.

cs.IT

Making it to First: The Random Access Problem in DNA Storage

In this paper, we study the Random Access Problem in DNA storage, which addresses the challenge of retrieving a specific information strand from a DNA-based storage system. In this framework, the data is represented by $k$ information strands which represent the data and are encoded into $n$ strands using a linear code. Then, each sequencing read returns one encoded strand which is chosen uniformly at random. The goal under this paradigm is to design codes that minimize the expected number of reads required to recover an arbitrary information strand. We fully solve the case when $k=2$, showing that the best possible code attains a random access expectation of $1+\frac{2}{\sqrt{2}+1}\approx 0.914\cdot 2$ for $q$ large enough. Moreover, we generalize a construction from~\cite{GMZ24}, specifically to $k=3$, for any value of $k$. Our construction uses $B_{k-1}$ sequences over $\mathbb{Z}_{q-1}$, that always exist over large finite fields. We show that for every $k\geq 4$, this generalized construction outperforms all previous constructions in terms of reducing the random access expectation.

cs.IT

On the Long-Term behavior of $k$-tuples Frequencies in Mutation Systems

In response to the evolving landscape of data storage, researchers have increasingly explored non-traditional platforms, with DNA-based storage emerging as a cutting-edge solution. Our work is motivated by the potential of in-vivo DNA storage, known for its capacity to store vast amounts of information efficiently and confidentially within an organism's native DNA. While promising, in-vivo DNA storage faces challenges, including susceptibility to errors introduced by mutations. To understand the long-term behavior of such mutation systems, we investigate the frequency of $k$-tuples after multiple mutation applications. Drawing inspiration from related works, we generalize results from the study of mutation systems, particularly focusing on the frequency of $k$-tuples. In this work, we provide a broad analysis through the construction of a specialized matrix and the identification of its eigenvectors. In the context of substitution and duplication systems, we leverage previous results on almost sure convergence, equating the expected frequency to the limiting frequency. Moreover, we demonstrate convergence in probability under certain assumptions.

cs.IT

Bounds and Constructions for Generalized Batch Codes

Private information retrieval (PIR) codes and batch codes are two important types of codes that are designed for coded distributed storage systems and private information retrieval protocols. These codes have been the focus of much attention in recent years, as they enable efficient and secure storage and retrieval of data in distributed systems. In this paper, we introduce a new class of codes called \emph{$(s,t)$-batch codes}. These codes are a type of storage codes that can handle any multi-set of $t$ requests, comprised of $s$ distinct information symbols. Importantly, PIR codes and batch codes are special cases of $(s,t)$-batch codes. The main goal of this paper is to explore the relationship between the number of redundancy symbols and the $(s,t)$-batch code property. Specifically, we establish a lower bound on the number of redundancy symbols required and present several constructions of $(s,t)$-batch codes. Furthermore, we extend this property to the case where each request is a linear combination of information symbols, which we refer to as \emph{functional $(s,t)$-batch codes}. Specifically, we demonstrate that simplex codes are asymptotically optimal functional $(s,t)$-batch codes, in terms of the number of redundancy symbols required, under certain parameter regime.

cs.IT

Storage codes and recoverable systems on lines and grids

A storage code is an assignment of symbols to the vertices of a connected graph $G(V,E)$ with the property that the value of each vertex is a function of the values of its neighbors, or more generally, of a certain neighborhood of the vertex in $G$. In this work we introduce a new construction method of storage codes, enabling one to construct new codes from known ones via an interleaving procedure driven by resolvable designs. We also study storage codes on $\mathbb Z$ and ${\mathbb Z}^2$ (lines and grids), finding closed-form expressions for the capacity of several one and two-dimensional systems depending on their recovery set, using connections between storage codes, graphs, anticodes, and difference-avoiding sets.

cs.IT

Codes Over Absorption Channels

In this paper, we present a novel communication channel, called the absorption channel, inspired by information transmission in neurons. Our motivation comes from in-vivo nano-machines, emerging medical applications, and brain-machine interfaces that communicate over the nervous system. Another motivation comes from viewing our model as a specific deletion channel, which may provide a new perspective and ideas to study the general deletion channel. For any given finite alphabet, we give codes that can correct absorption errors. For the binary alphabet, the problem is relatively trivial and we can apply binary (multiple-) deletion correcting codes. For single-absorption error, we prove that the Varshamov-Tenengolts codes can provide a near-optimal code in our setting. When the alphabet size $q$ is at least $3$, we first construct a single-absorption correcting code whose redundancy is at most $3\log_q(n)+O(1)$. Then, based on this code and ideas introduced in \cite{Gabrys2022IT}, we give a second construction of single-absorption correcting codes with redundancy $\log_q(n)+12\log_q\log_q(n)+O(1)$, which is optimal up to an $O\left(\log_q\log_q(n)\right)$. Finally, we apply the syndrome compression technique with pre-coding to obtain a subcode of the single-absorption correcting code. This subcode can combat multiple-absorption errors and has low redundancy. For each setup, efficient encoders and decoders are provided.

cs.IT

Binary $t_1$-Deletion-$t_2$-Insertion-Burst Correcting Codes and Codes Correcting a Burst of Deletions

We first give a construction of binary $t_1$-deletion-$t_2$-insertion-burst correcting codes with redundancy at most $\log(n)+(t_1-t_2-1)\log\log(n)+O(1)$, where $t_1\ge 2t_2$. Then we give an improved construction of binary codes capable of correcting a burst of $4$ non-consecutive deletions, whose redundancy is reduced from $7\log(n)+2\log\log(n)+O(1)$ to $4\log(n)+6\log\log(n)+O(1)$. Lastly, by connecting non-binary $b$-burst-deletion correcting codes with binary $2b$-deletion-$b$-insertion-burst correcting codes, we give a new construction of non-binary $b$-burst-deletion correcting codes with redundancy at most $\log(n)+(b-1)\log\log(n)+O(1)$. This construction is different from previous results.

cs.IT

Reconstruction of a Single String from a Part of its Composition Multiset

Motivated by applications in polymer-based data storage, we study the problem of reconstructing a string from part of its composition multiset. We give a full description of the structure of the strings that cannot be uniquely reconstructed (up to reversal) from their multiset of all of their prefix-suffix compositions. Leveraging this description, we prove that for all $n\ge 6$, there exists a string of length $n$ that cannot be uniquely reconstructed up to reversal. Moreover, for all $n\ge 6$, we explicitly construct the set consisting of all length $n$ strings that can be uniquely reconstructed up to reversal. As a by product, we obtain that any binary string can be constructed using Dyck strings and Catalan-Bertrand strings. For any given string $\bm{s}$, we provide a method to explicitly construct the set of all strings with the same prefix-suffix composition multiset as $\bm{s}$, as well as a formula for the size of this set. As an application, we construct a composition code of maximal size. Furthermore, we construct two classes of composition codes which can respectively correct composition missing errors and mass reducing substitution errors. In addition, we raise two new problems: reconstructing a string from its composition multiset when at most a constant number of substring compositions are lost; reconstructing a string when only given its compositions of substrings of length at most $r$. For each of these setups, we give suitable codes under some conditions.

cs.IT

Optimal Reference for DNA Synthesis

In the recent years, DNA has emerged as a potentially viable storage technology. DNA synthesis, which refers to the task of writing the data into DNA, is perhaps the most costly part of existing storage systems. Accordingly, this high cost and low throughput limits the practical use in available DNA synthesis technologies. It has been found that the homopolymer run (i.e., the repetition of the same nucleotide) is a major factor affecting the synthesis and sequencing errors. Quite recently, [26] studied the role of batch optimization in reducing the cost of large scale DNA synthesis, for a given pool $\mathcal{S}$ of random quaternary strings of fixed length. Among other things, it was shown that the asymptotic cost savings of batch optimization are significantly greater when the strings in $\mathcal{S}$ contain repeats of the same character (homopolymer run of length one), as compared to the case where strings are unconstrained. Following the lead of [26], in this paper, we take a step forward towards the theoretical understanding of DNA synthesis, and study the homopolymer run of length $k\geq1$. Specifically, we are given a set of DNA strands $\mathcal{S}$, randomly drawn from a natural Markovian distribution modeling a general homopolymer run length constraint, that we wish to synthesize. For this problem, we prove that for any $k\geq 1$, the optimal reference strand, minimizing the cost of DNA synthesis is, perhaps surprisingly, the periodic sequence $\overline{\mathsf{ACGT}}$. It turns out that tackling the homopolymer constraint of length $k\geq2$ is a challenging problem; our main technical contribution is the representation of the DNA synthesis process as a certain constrained system, for which string techniques can be applied.

cs.IT

Recoverable Systems

Motivated by the established notion of storage codes, we consider sets of infinite sequences over a finite alphabet such that every $k$-tuple of consecutive entries is uniquely recoverable from its $l$-neighborhood in the sequence. We address the problem of finding the maximum growth rate of the set, which we term capacity, as well as constructions of explicit families that approach the optimal rate. The techniques that we employ rely on the connection of this problem with constrained systems. In the second part of the paper we consider a modification of the problem wherein the entries in the sequence are viewed as random variables over a finite alphabet that follow some joint distribution, and the recovery condition requires that the Shannon entropy of the $k$-tuple conditioned on its $l$-neighborhood be bounded above by some $ε>0.$ We study properties of measures on infinite sequences that maximize the metric entropy under the recoverability condition. Drawing on tools from ergodic theory, we prove some properties of entropy-maximizing measures. We also suggest a procedure of constructing an $ε$-recoverable measure from a corresponding deterministic system.

cs.IT

Repeat-Free Codes

In this paper we consider the problem of encoding data into \textit{repeat-free} sequences in which sequences are imposed to contain any $k$-tuple at most once (for predefined $k$). First, the capacity of the repeat-free constraint are calculated. Then, an efficient algorithm, which uses two bits of redundancy, is presented to encode length-$n$ sequences for $k=2+2\log (n)$. This algorithm is then improved to support any value of $k$ of the form $k=a\log (n)$, for $1<a$, while its redundancy is $o(n)$. We also calculate the capacity of repeat-free sequences when combined with local constraints which are given by a constrained system, and the capacity of multi-dimensional repeat-free codes.

cs.IT

Capacity of dynamical storage systems

We introduce a dynamical model of node repair in distributed storage systems wherein the storage nodes are subjected to failures according to independent Poisson processes. The main parameter that we study is the time-average capacity of the network in the scenario where a fixed subset of the nodes support a higher repair bandwidth than the other nodes. The sequence of node failures generates random permutations of the nodes in the encoded block, and we model the state of the network as a Markov random walk on permutations of $n$ elements. As our main result we show that the capacity of the network can be increased compared to the static (worst-case) model of the storage system, while maintaining the same (average) repair bandwidth, and we derive estimates of the increase. We also quantify the capacity increase in the case that the repair center has information about the sequence of the recently failed storage nodes.

cs.IT

The Capacity of Some Pólya String Models

We study random string-duplication systems, which we call Pólya string models. These are motivated by DNA storage in living organisms, and certain random mutation processes that affect their genome. Unlike previous works that study the combinatorial capacity of string-duplication systems, or various string statistics, this work provides exact capacity or bounds on it, for several probabilistic models. In particular, we study the capacity of noisy string-duplication systems, including the tandem-duplication, end-duplication, and interspersed-duplication systems. Interesting connections are drawn between some systems and the signature of random permutations, as well as to the beta distribution common in population genetics.

cs.IT