SearcharxivSearch

arXiv subjects

Wentu Song

Publications and source records attributed to Wentu Song.

At least 19 recordsLinked to original sources

Analog Error Correcting Codes with Constant Redundancy

We consider analog error-correcting codes (analog ECCs) that are designed to correct/detect outlying errors arising in analog implementations of vector-matrix multiplication. The error-correction/detection capability of an analog ECC can be characterized by its height profile, which is expected to be as small as possible. In this paper, we consider analog ECCs whose parity check matrix has columns of unit Euclidean norm. We first present an upper bound on the height profile of such codes as well as a simple decoder for correcting a single error. We then construct a family of single error-correcting analog ECCs with redundancy three for any code length, which has smaller height profile compared to the known $[n,n-2]$ MDS constructions.

cs.IT

On the Sequence Reconstruction Problem for the Single-Deletion Two-Substitution Channel

The Levenshtein sequence reconstruction problem studies the reconstruction of a transmitted sequence from multiple erroneous copies of it. A fundamental question in this field is to determine the minimum number of erroneous copies required to guarantee correct reconstruction of the original sequence. This problem is equivalent to determining the maximum possible intersection size of two error balls associated with the underlying channel. Existing research on the sequence reconstruction problem has largely focused on channels with a single type of error, such as insertions, deletions, or substitutions alone. However, relatively little is known for channels that involve a mixture of error types, for instance, channels allowing both deletions and substitutions. In this work, we study the sequence reconstruction problem for the single-deletion two-substitution channel, which allows one deletion and at most two substitutions applied to the transmitted sequence. Specifically, we prove that if two $q$-ary length-$n$ sequences have the Hamming distance $d\geq 2$, where $q\geq 2$ is any fixed integer, then the intersection size of their error balls under the single-deletion two-substitution channel is upper bounded by $(q^2-1)n^2-(3q^2+5q-5)n+O_q(1)$, where $O_q(1)$ is a constant independent from $n$ but dependent on $q$. Moreover, we show that this upper bound is tight up to an additive constant.

cs.IT

Sequence Reconstruction for the Single-Deletion Single-Substitution Channel

The central problem in sequence reconstruction is to find the minimum number of distinct channel outputs required to uniquely reconstruct the transmitted sequence. According to Levenshtein's work in 2001, this number is determined by the size of the maximum intersection between the error balls of any two distinct input sequences of the channel. In this work, we study the sequence reconstruction problem for single-deletion single-substitution channel, assuming that the transmitted sequence belongs to a $q$-ary code with minimum Hamming distance at least $2$, where $q\geq 2$ is any fixed integer. Specifically, we prove that for any two $q$-ary sequences of length $n$ and with Hamming distance $d\geq 2$, the size of the intersection of their error balls is upper bounded by $2qn-3q-2-\delta_{q,2}$, where $\delta_{i,j}$ is the Kronecker delta. We also prove the tightness of this bound by constructing two sequences the intersection size of whose error balls achieves this bound.

cs.IT

High Information Density and Low Coverage Data Storage in DNA with Efficient Channel Coding Schemes

DNA-based data storage has been attracting significant attention due to its extremely high data storage density, low power consumption, and long duration compared to conventional data storage media. Despite the recent advancements in DNA data storage technology, significant challenges remain. In particular, various types of errors can occur during the processes of DNA synthesis, storage, and sequencing, including substitution errors, insertion errors, and deletion errors. Furthermore, the entire oligo may be lost. In this work, we report a DNA-based data storage architecture that incorporates efficient channel coding schemes, including different types of error-correcting codes (ECCs) and constrained codes, for both the inner coding and outer coding for the DNA data storage channel. We also carried out large scale experiments to validate our proposed DNA-based data storage architecture. Specifically, 1.61 and 1.69 MB data were encoded into 30,000 oligos each, with information densities of 1.731 and 1.815, respectively. It has been found that the stored information can be fully recovered without any error at average coverages of 4.5 and 6.0, respectively. This experiment achieved the highest net information density and lowest coverage among existing DNA-based data storage experiments (with standard DNA), with data recovery rates and coverage approaching theoretical optima.

cs.IT

New Construction of $q$-ary Codes Correcting a Burst of at most $t$ Deletions

In this paper, for any fixed positive integers $t$ and $q>2$, we construct $q$-ary codes correcting a burst of at most $t$ deletions with redundancy $\log n+8\log\log n+o(\log\log n)+γ_{q,t}$ bits and near-linear encoding/decoding complexity, where $n$ is the message length and $γ_{q,t}$ is a constant that only depends on $q$ and $t$. In previous works there are constructions of such codes with redundancy $\log n+O(\log q\log\log n)$ bits or $\log n+O(t^2\log\log n)+O(t\log q)$. The redundancy of our new construction is independent of $q$ and $t$ in the second term.

cs.IT

Non-binary Two-Deletion Correcting Codes and Burst-Deletion Correcting Codes

In this paper, we construct systematic $q$-ary two-deletion correcting codes and burst-deletion correcting codes, where $q\geq 2$ is an even integer. For two-deletion codes, our construction has redundancy $5\log n+O(\log q\log\log n)$ and has encoding complexity near-linear in $n$, where $n$ is the length of the message sequences. For burst-deletion codes, we first present a construction of binary codes with redundancy $\log n+9\log\log n+γ_t+o(\log\log n)$ bits $(γ_t$ is a constant that depends only on $t)$ and capable of correcting a burst of at most $t$ deletions, which improves the Lenz-Polyanskii Construction (ISIT 2020). Then we give a construction of $q$-ary codes with redundancy $\log n+(8\log q+9)\log\log n+o(\log q\log\log n)+γ_t$ bits and capable of correcting a burst of at most $t$ deletions.

cs.IT

List-decodable Codes for Single-deletion Single-substitution with List-size Two

In this paper, we present an explicit construction of list-decodable codes for single-deletion and single-substitution with list size two and redundancy 3log n+4, where n is the block length of the code. Our construction has lower redundancy than the best known explicit construction by Gabrys et al. (arXiv 2021), whose redundancy is 4log n+O(1).

cs.IT

Dynamic Programming for Sequential Deterministic Quantization of Discrete Memoryless Channels

In this paper, under a general cost function $C$, we present a dynamic programming (DP) method to obtain an optimal sequential deterministic quantizer (SDQ) for $q$-ary input discrete memoryless channel (DMC). The DP method has complexity $O(q (N-M)^2 M)$, where $N$ and $M$ are the alphabet sizes of the DMC output and quantizer output, respectively. Then, starting from the quadrangle inequality, two techniques are applied to reduce the DP method's complexity. One technique makes use of the Shor-Moran-Aggarwal-Wilber-Klawe (SMAWK) algorithm and achieves complexity $O(q (N-M) M)$. The other technique is much easier to be implemented and achieves complexity $O(q (N^2 - M^2))$. We further derive a sufficient condition under which the optimal SDQ is optimal among all quantizers and the two techniques are applicable. This generalizes the results in the literature for binary-input DMC. Next, we show that the cost function of $α$-mutual information ($α$-MI)-maximizing quantizer belongs to the category of $C$. We further prove that under a weaker condition than the sufficient condition we derived, the aforementioned two techniques are applicable to the design of $α$-MI-maximizing quantizer. Finally, we illustrate the particular application of our design method to practical pulse-amplitude modulation systems.

cs.IT

Systematic Single-Deletion Multiple-Substitution Correcting Codes

Recent work by Smagloy et al. (ISIT 2020) shows that the redundancy of a single-deletion $s$-substitution correcting code is asymptotically at least $(s+1)\log n+o(\log n)$, where $n$ is the length of the codes. They also provide a construction of single-deletion and single-substitution codes with redundancy $6\log n+8$. In this paper, we propose a family of systematic single-deletion $s$-substitution correcting codes of length $n$ with asymptotical redundancy at most $(3s+4)\log n+o(\log n)$ and polynomial encoding/decoding complexity, where $s\geq 2$ is a constant. Specifically, the encoding and decoding complexity of the proposed codes are $O(n^{s+3})$ and $O(n^{s+2})$, respectively.

cs.IT

Sequence-Subset Distance and Coding for Error Control in DNA-based Data Storage

The process of DNA-based data storage (DNA storage for short) can be mathematically modelled as a communication channel, termed DNA storage channel, whose inputs and outputs are sets of unordered sequences. To design error correcting codes for DNA storage channel, a new metric, termed the sequence-subset distance, is introduced, which generalizes the Hamming distance to a distance function defined between any two sets of unordered vectors and helps to establish a uniform framework to design error correcting codes for DNA storage channel. We further introduce a family of error correcting codes, referred to as \emph{sequence-subset codes}, for DNA storage and show that the error-correcting ability of such codes is completely determined by their minimum distance. We derive some upper bounds on the size of the sequence-subset codes including a tight bound for a special case, a Singleton-like bound and a Plotkin-like bound. We also propose some constructions, including an optimal construction for that special case, which imply lower bounds on the size of such codes.

cs.IT

Coded Caching with Polynomial Subpacketization

Consider a centralized caching network with a single server and $K$ users. The server has a database of $N$ files with each file being divided into $F$ packets ($F$ is known as subpacketization), and each user owns a local cache that can store $\frac{M}{N}$ fraction of the $N$ files. We construct a family of centralized coded caching schemes with polynomial subpacketization. Specifically, given $M$, $N$ and an integer $n\geq 0$, we construct a family of coded caching schemes for any $(K,M,N)$ caching system with $F=O(K^{n+1})$. More generally, for any $t\in\{1,2,\cdots,K-2\}$ and any integer $n$ such that $0\leq n\leq t$, we construct a coded caching scheme with $\frac{M}{N}=\frac{t}{K}$ and $F\leq K\binom{\left(1-\frac{M}{N}\right)K+n}{n}$.

cs.IT

Some new Constructions of Coded Caching Schemes with Reduced Subpacketization

We study the problem of constructing centralized coded caching schemes with low subpacketization level based on the placement delivery array (PDA) design framework. PDA design is an efficient way to construct centralized coded caching schemes and most existing schemes, including the famous Maddah-Ali-Niesen scheme, can be described using PDA. In this paper, we first prove that constructing a PDA is equivalent to constructing three binary matrices that satisfy certain conditions. From this perspective, we then propose some new constructions of coded caching schemes using PDA design based on projective geometries over finite fields, combinatorial configurations, and t-designs, respectively. Our constructions achieve low subpacketization level (e.g., linear subpacketization) with reasonable rate loss and include several known results as special cases. Finally, we give an approach to construct new coded caching scheme from existing schemes based on direct product of PDAs. Our results enrich the coded caching schemes of low subpacketization level.

cs.IT

Generalized Reed-Solomon Codes with Sparsest and Balanced Generator Matrices

We prove that for any positive integers $n$ and $k$ such that $n\!\geq\! k\!\geq\! 1$, there exists an $[n,k]$ generalized Reed-Solomon (GRS) code that has a sparsest and balanced generator matrix (SBGM) over any finite field of size $q\!\geq\! n\!+\!\lceil\frac{k(k-1)}{n}\rceil$, where sparsest means that each row of the generator matrix has the least possible number of nonzeros, while balanced means that the number of nonzeros in any two columns differ by at most one. Previous work by Dau et al (ISIT'13) showed that there always exists an MDS code that has an SBGM over any finite field of size $q\geq {n-1\choose k-1}$, and Halbawi et al (ISIT'16, ITW'16) showed that there exists a cyclic Reed-Solomon code (i.e., $n=q-1$) with an SBGM for any prime power $q$. Hence, this work extends both of the previous results.

cs.IT

On Sequential Locally Repairable Codes

We consider the locally repairable codes (LRC), aiming at sequential recovering multiple erasures. We define the (n,k,r,t)-SLRC (Sequential Locally Repairable Codes) as an [n,k] linear code where any t'(>= t) erasures can be sequentially recovered, each one by r (2<=r =3 erasures and bounds to evaluate the performance of such codes. We first derive a tight upper bound on the code rate of (n, k, r, t)-SLRC for t=3 and r>=2. We then propose two constructions of binary (n, k, r, t)-SLRCs for general r,t>=2 (Existing constructions are dealing with t<=7 erasures. The first construction generalizes the method of direct product construction. The second construction is based on the resolvable configurations and yields SLRCs for any r>=2 and odd t>=3. For both constructions, the rates are optimal for t in {2,3} and are higher than most of the existing LRC families for arbitrary t>=4.

cs.IT

Binary Locally Repairable Codes ---Sequential Repair for Multiple Erasures

Locally repairable codes (LRC) for distribute storage allow two approaches to locally repair multiple failed nodes: 1) parallel approach, by which each newcomer access a set of $r$ live nodes $(r$ is the repair locality$)$ to download data and recover the lost packet; and 2) sequential approach, by which the newcomers are properly ordered and each newcomer access a set of $r$ other nodes, which can be either a live node or a newcomer ordered before it. An $[n,k]$ linear code with locality $r$ and allows local repair for up to $t$ failed nodes by sequential approach is called an $(n,k,r,t)$-exact locally repairable code (ELRC). In this paper, we present a family of binary codes which is equivalent to the direct product of $m$ copies of the $[r+1,r]$ single-parity-check code. We prove that such codes are $(n,k,r,t)$-ELRC with $n=(r+1)^m,k=r^m$ and $t=2^m-1$, which implies that they permit local repair for up to $2^m-1$ erasures by sequential approach. Our result shows that the sequential approach has much bigger advantage than parallel approach.

cs.IT

Erasure codes with symbol locality and group decodability for distributed storage

We introduce a new family of erasure codes, called group decodable code (GDC), for distributed storage system. Given a set of design parameters {α; β; k; t}, where k is the number of information symbols, each codeword of an (α; β; k; t)-group decodable code is a t-tuple of strings, called buckets, such that each bucket is a string of βsymbols that is a codeword of a [β; α] MDS code (which is encoded from αinformation symbols). Such codes have the following two properties: (P1) Locally Repairable: Each code symbol has locality (α; β-α+ 1). (P2) Group decodable: From each bucket we can decode αinformation symbols. We establish an upper bound of the minimum distance of (α; β; k; t)-group decodable code for any given set of {α; β; k; t}; We also prove that the bound is achievable when the coding field F has size |F| > n-1 \choose k-1.

cs.IT

Locally Repairable Codes with Functional Repair and Multiple Erasure Tolerance

We consider the problem of designing [n; k] linear codes for distributed storage systems (DSS) that satisfy the (r, t)-Local Repair Property, where any t'(<=t) simultaneously failed nodes can be locally repaired, each with locality r. The parameters n, k, r, t are positive integers such that r<k<n and t <= n-k. We consider the functional repair model and the sequential approach for repairing multiple failed nodes. By functional repair, we mean that the packet stored in each newcomer is not necessarily an exact copy of the lost data but a symbol that keep the (r, t)-local repair property. By the sequential approach, we mean that the t' newcomers are ordered in a proper sequence such that each newcomer can be repaired from the live nodes and the newcomers that are ordered before it. Such codes, which we refer to as (n, k, r, t)-functional locally repairable codes (FLRC), are the most general class of LRCs and contain several subclasses of LRCs reported in the literature. In this paper, we aim to optimize the storage overhead (equivalently, the code rate) of FLRCs. We derive a lower bound on the code length n given t belongs to {2,3} and any possible k, r. For t=2, our bound generalizes the rate bound proved in [14]. For t=3, our bound improves the rate bound proved in [10]. We also give some onstructions of exact LRCs for t belongs to {2,3} whose length n achieves the bound of (n, k, r, t)-FLRC, which proves the tightness of our bounds and also implies that there is no gap between the optimal code length of functional LRCs and exact LRCs for certain sets of parameters. Moreover, our constructions are over the binary field, hence are of interest in practice.

cs.IT

Local Codes with Addition Based Repair

We consider the complexities of repair algorithms for locally repairable codes and propose a class of codes that repair single node failures using addition operations only, or codes with addition based repair. We construct two families of codes with addition based repair. The first family attains distance one less than the Singleton-like upper bound, while the second family attains the Singleton-like upper bound.

cs.IT