SearcharxivSearch

arXiv subjects

Zuo Ye

Publications and source records attributed to Zuo Ye.

13 recordsLinked to original sources

More on Codes for Combinatorial Composite DNA

In this paper, we focus on constructions of unique-decodable/list-decodable on the recently studied $(t,e)$-composite-asymmetric error-correcting codes ($(t,e)$-CAECCs). Let $X$ be an $m\times n$ binary matrix, in which each row has Hamming weight $w$. When at most $t$ rows of $X$ suffer from errors and in each of these erroneous rows, there are at most $e$ $1 \to 0$ errors, we say that a $(t,e)$-composite-asymmetric-error occurs in $X$. For general $m,n,w,t,e$, we propose new constructions of $(t,e)$-CAECCs with redundancy at most $(t-1)\log(m)+O(1)$, where $O(1)$ is a number independent of the code-length $m$. In particular, this gives a class of $(2,e)$-CAECCs that are optimal in terms of their redundancy. %(in terms of redundancy, regarded as a function of the number of rows $m$) s. When $m$ is a prime power, the redundancy can be further reduced to $(t-1)\log(m)-O(\log(m))$. To further increase the size of these codes, we introduce a combinatorial object called a weak $B_e$-sets. When $e=w$, we show an efficient way to encode/decode our codes. At last, we investigate how much we can gain if we relax the requirement of uniquely decoding to list-decoding. It is shown that when the list size is $t!$ or an exponential function of $t$, there are list-decodable $(t,e)$-CAECCs with constant redundancy. When the list size is two, we show that there are list-decodable $(3,2)$-CAECCs with redundancy $\log(m)+O(1)$.

cs.IT

Deletion-Correcting Codes for the $\ell$-Symbol Read Channel

This paper studies deletion-correcting codes for the $\ell$-symbol read channel, whose noiseless output is the vector of all consecutive $\ell$-mers of a transmitted sequence. This model is motivated by overlapping-read mechanisms arising in nanopore sequencing, racetrack memories with consecutive read heads, and related sequence-labeling problems. We consider an adversarial setting in which a fixed number of $\ell$-mers are deleted from the read vector. Our first contribution is a structural characterization of the effect of such deletions: after a minimum number of $\ell$-mers are inserted to restore consistency, the resulting sequence is obtained from the transmitted sequence by deleting symbols from certain periodic substrings; when $t\le \ell-2$, these deletions correspond to complete minimum periods. Based on this characterization, we introduce check patterns and construct $\ell$-read deletion-correcting codes via power-sum syndromes. For every $\ell\ge2$, we obtain single-deletion correcting codes with redundancy $\log\lfloor (n+2\ell)/(\ell-1)\rfloor$. For $2\le t\le \ell/2$, we construct $q$-ary $\ell$-read $t$-deletion correcting codes with redundancy $t\log n+O(1)$, and for $\ell=2t-1$ with $t\ge3$, we construct codes with redundancy $(2t-1)\log n+O(1)$. We also study the sporadic parameter pairs $(\ell,t)\in\{(2,2),(3,2),(3,3)\}$ and obtain improved constructions, including binary $\ell$-read $2$-deletion correcting codes with redundancy $2\log n+O(1)$ for $\ell=2,3$, a non-binary $3$-read $2$-deletion correcting code with redundancy $3\log n+O(1)$, a binary $3$-read $3$-deletion correcting code with redundancy $5\log n+O(1)$, and a non-binary $3$-read $3$-deletion correcting code with redundancy $7\log n+O(1)$.

cs.IT

Bounds and Constructions of Codes for Ordered Composite DNA Sequences

This paper extends the foundational work of Dollma \emph{et al}. on codes for ordered composite DNA sequences. We consider the general setting with an alphabet of size $q$ and a resolution parameter $k$, moving beyond the binary ($q=2$) case primarily studied previously. We investigate error-correcting codes for substitution errors and deletion errors under several channel models, including $(e_1,\ldots,e_k)$-composite error/deletion, $e$-composite error/deletion, and the newly introduced $t$-$(e_1,\ldots,e_t)$-composite error/deletion model. We first establish equivalence relations among families of composite-error correcting codes (CECCs) and among families of composite-deletion correcting codes (CDCCs). This significantly reduces the number of distinct error-parameter sets that require separate analysis. We then derive novel and general upper bounds on the sizes of CECCs using refined sphere-packing arguments and probabilistic methods. These bounds together cover all values of parameters $q$, $k$, $(e_1,\ldots,e_k)$ and $e$. In contrast, previous bounds were only established for $q=2$ and limited choices of $k$, $(e_1,\ldots,e_k)$ and $e$. For CDCCs, we generalize a known non-asymptotic upper bound for $(1,0,\ldots,0)$-CDCCs and then provide a cleaner asymptotic bound. On the constructive side, for any $q\ge2$, we propose $(1,0,\ldots,0)$-CDCCs, $1$-CDCCs and $t$-$(1,\ldots,1)$-CDCCs with near-optimal redundancies. These codes have efficient and systematic encoders. For substitution errors, we design the first explicit encoding and decoding algorithms for the binary $(1,0,\ldots,0)$-CECC constructed by Dollma \emph{et al}, and extend the approach to general $q$. Furthermore, we give an improved construction of binary $1$-CECCs, a construction of nonbinary $1$-CECCs, and a construction of $t$-$(1,\ldots,1)$-CECCs. These constructions are also systematic.

cs.IT

Correcting Bursty/Localized Deletions: A New Error-Position-Estimation Code

Codes correcting bursts of deletions and localized deletions have garnered significant research interest in recent years. One of the primary objectives is to construct codes with minimal redundancy. Currently, the best known constructions of $q$-ary codes correcting a burst of at most $t$ deletions ($(\le t)$-burst-deletion correcting codes) achieve redundancy $\log n+8\log\log n+o(\log\log n)$ (for any $q$ and $t$) or $\log n+t\log\log n+O(1)$ (for even $q$). For codes correcting single $t$-localized-deletion ($t$-localized-deletion correcting codes), state-of-the-art constructions attain redundancy $\log n+O\parenv{t(\log\log n)^2}$ (for any $q$ and $t$) or $\log n+2t\log\log n+O(1)$ (for even $q$). Here, $n$ denotes the code-length, and $q$ and $t$ are fixed. These codes employ a position-estimation component to approximate error positions, augmented by additional constraints that enable error-correction given the information about error positions. In this work, we select codewords from the set of sequences whose differential sequences are strong-$(\ell,\epsilon)$-locally-balanced. By imposing a VT-type constraint and an $L_1$-weight constraint on the differential sequences of codewords, we construct novel position-estimation codes. When $q\ge 2$ and $t<q$, or $q$ is even and $t<2q$, this approach gives a $q$-ary $(\le t)$-burst-deletion correcting code and a $t$-localized-deletion correcting code with redundancy $\log n+(t-1)\log\log n+O(1)$. In addition to improving previous redundancy, the method is new and our position-estimation codes are simpler than those in previous works. Finally, we give an efficient encoder to encode an arbitrary input sequence into a sequence whose differential sequence is strong-$(\ell,\epsilon)$-locally-balanced. To our knowledge, no prior algorithm for this specific task has been reported.

cs.IT

On the Maximum Size of Codes Under the Damerau-Levenshtein Metric

The Damerau-Levenshtein distance between two sequences is the minimum number of operations (deletions, insertions, substitutions, and adjacent transpositions) required to convert one sequence into another. Notwithstanding a long history of this metric, research on error-correcting codes under this distance has remained limited. Recently, motivated by applications in DNA-based storage systems, Gabrys \textit{et al} and Wang \texit{et al} reinvigorated interest in this metric. In their works, some codes correcting both deletions and adjacent transpositions were constructed. However, theoretical upper bounds on code sizes under this metric have not yet been established. This paper seeks to establish upper bounds for code sizes in the Damerau-Levenshtein metric. Our results show that the code correcting one deletion and asymmetric adjacent transpositions proposed by Wang \textit{et al} achieves optimal redundancy up to an additive constant.

cs.IT

On the Asymptotic Rate of Optimal Codes that Correct Tandem Duplications for Nanopore Sequencing

We study codes that can correct backtracking errors during nanopore sequencing. In this channel, a sequence of length $n$ over an alphabet of size $q$ is being read by a sliding window of length $\ell$, where from each window we obtain only its composition. Backtracking errors cause some windows to repeat, hence manifesting as tandem-duplication errors of length $k$ in the $\ell$-read vector of window compositions. While existing constructions for duplication-correcting codes can be straightforwardly adapted to this model, even resulting in optimal codes, their asymptotic rate is hard to find. In the regime of unbounded number of duplication errors, we either give the exact asymptotic rate of optimal codes, or bounds on it, depending on the values of $k$, $\ell$ and $q$. In the regime of a constant number of duplication errors, $t$, we find the redundancy of optimal codes to be $t\log_q n+O(1)$ when $\ell|k$, and only upper bounded by this quantity otherwise.

cs.IT

Codes Correcting Two Bursts of Exactly $b$ Deletions

In this paper, we investigate codes designed to correct two bursts of deletions, where each burst has a length of exactly $b$, where $b>1$. The previous best construction, achieved through the syndrome compression technique, had a redundancy of at most $7\log n+O\left(\log n/\log\log n\right)$ bits. In contrast, our work introduces a novel approach for constructing $q$-ary codes that attain a redundancy of at most $5\log n+O(\log\log n)$ bits for all $b>1$ and $q\ge2$. Additionally, for the case where $b=1$, we present a new construction of $q$-ary two-deletion correcting codes with a redundancy of $5\log n+O(\log\log n)$ bits, for all $q>2$.

cs.IT

Reconstruction of Sequences Distorted by Two Insertions

Reconstruction codes are generalizations of error-correcting codes that can correct errors by a given number of noisy reads. The study of such codes was initiated by Levenshtein in 2001 and developed recently due to applications in modern storage devices such as racetrack memories and DNA storage. The central problem on this topic is to design codes with redundancy as small as possible for a given number $N$ of noisy reads. In this paper, the minimum redundancy of such codes for binary channels with exactly two insertions is determined asymptotically for all values of $N\ge 5$. Previously, such codes were studied only for channels with single edit errors or two-deletion errors.

cs.IT

Codes Over Absorption Channels

In this paper, we present a novel communication channel, called the absorption channel, inspired by information transmission in neurons. Our motivation comes from in-vivo nano-machines, emerging medical applications, and brain-machine interfaces that communicate over the nervous system. Another motivation comes from viewing our model as a specific deletion channel, which may provide a new perspective and ideas to study the general deletion channel. For any given finite alphabet, we give codes that can correct absorption errors. For the binary alphabet, the problem is relatively trivial and we can apply binary (multiple-) deletion correcting codes. For single-absorption error, we prove that the Varshamov-Tenengolts codes can provide a near-optimal code in our setting. When the alphabet size $q$ is at least $3$, we first construct a single-absorption correcting code whose redundancy is at most $3\log_q(n)+O(1)$. Then, based on this code and ideas introduced in \cite{Gabrys2022IT}, we give a second construction of single-absorption correcting codes with redundancy $\log_q(n)+12\log_q\log_q(n)+O(1)$, which is optimal up to an $O\left(\log_q\log_q(n)\right)$. Finally, we apply the syndrome compression technique with pre-coding to obtain a subcode of the single-absorption correcting code. This subcode can combat multiple-absorption errors and has low redundancy. For each setup, efficient encoders and decoders are provided.

cs.IT

Binary $t_1$-Deletion-$t_2$-Insertion-Burst Correcting Codes and Codes Correcting a Burst of Deletions

We first give a construction of binary $t_1$-deletion-$t_2$-insertion-burst correcting codes with redundancy at most $\log(n)+(t_1-t_2-1)\log\log(n)+O(1)$, where $t_1\ge 2t_2$. Then we give an improved construction of binary codes capable of correcting a burst of $4$ non-consecutive deletions, whose redundancy is reduced from $7\log(n)+2\log\log(n)+O(1)$ to $4\log(n)+6\log\log(n)+O(1)$. Lastly, by connecting non-binary $b$-burst-deletion correcting codes with binary $2b$-deletion-$b$-insertion-burst correcting codes, we give a new construction of non-binary $b$-burst-deletion correcting codes with redundancy at most $\log(n)+(b-1)\log\log(n)+O(1)$. This construction is different from previous results.

cs.IT

Reconstruction of a Single String from a Part of its Composition Multiset

Motivated by applications in polymer-based data storage, we study the problem of reconstructing a string from part of its composition multiset. We give a full description of the structure of the strings that cannot be uniquely reconstructed (up to reversal) from their multiset of all of their prefix-suffix compositions. Leveraging this description, we prove that for all $n\ge 6$, there exists a string of length $n$ that cannot be uniquely reconstructed up to reversal. Moreover, for all $n\ge 6$, we explicitly construct the set consisting of all length $n$ strings that can be uniquely reconstructed up to reversal. As a by product, we obtain that any binary string can be constructed using Dyck strings and Catalan-Bertrand strings. For any given string $\bm{s}$, we provide a method to explicitly construct the set of all strings with the same prefix-suffix composition multiset as $\bm{s}$, as well as a formula for the size of this set. As an application, we construct a composition code of maximal size. Furthermore, we construct two classes of composition codes which can respectively correct composition missing errors and mass reducing substitution errors. In addition, we raise two new problems: reconstructing a string from its composition multiset when at most a constant number of substring compositions are lost; reconstructing a string when only given its compositions of substrings of length at most $r$. For each of these setups, we give suitable codes under some conditions.

cs.IT

Strong quantum nonlocality in $N$-partite systems

A set of multipartite orthogonal quantum states is strongly nonlocal if it is locally irreducible for every bipartition of the subsystems [Phys. Rev. Lett. 122, 040403 (2019)]. Although this property has been shown in three-, four- and five-partite systems, the existence of strongly nonlocal sets in $N$-partite systems remains unknown when $N\geq 6$. In this paper, we successfully show that a strongly nonlocal set of orthogonal entangled states exists in $(\mathbb{C}^d)^{\otimes N}$ for all $N\geq 3$ and $d\geq 2$, which for the first time reveals the strong quantum nonlocality in general $N$-partite systems. For $N=3$ or $4$ and $d\geq 3$, we present a strongly nonlocal set consisting of genuinely entangled states, which has a smaller size than any known strongly nonlocal orthogonal product set. Finally, we connect strong quantum nonlocality with local hiding of information as an application.

quant-ph

Some New Results on Splitter Sets

Splitter sets have been widely studied due to their applications in flash memories, and their close relations with lattice tilings and conflict avoiding codes. In this paper, we give necessary and sufficient conditions for the existence of nonsingular perfect splitter sets, $B[-k_1,k_2](p)$ sets, where $0\le k_{1}\leq k_{2}=4$. Meanwhile, constructions of nonsingular perfect splitter sets are given. When perfect splitter sets do not exist, we present four new constructions of quasi-perfect splitter sets. Finally, we give a connection between nonsingular splitter sets and Cayley graphs, and as a byproduct, a general lower bound on the maximum size of nonsingular splitter sets is given.

cs.IT