SearcharxivSearch

arXiv subjects

Hengjia Wei

Publications and source records attributed to Hengjia Wei.

At least 19 recordsLinked to original sources

Constant-List Insertion--Deletion Codes:New Bounds and an Improvement of Levenshtein's Lower Bound

We study codes correcting adversarial insertions and deletions with list size $L$ fixed independently of the block length. We derive new achievable-rate bounds for binary codes and upper bounds over every fixed alphabet of size $q\ge2$, retaining explicit dependence on $L$. We establish a combinatorial reduction that trades $L$ units of insertion budget for one unit of deletion budget in the decoding guarantee, without changing the code or increasing the list size. Consequently, asymptotic bounds for mixed errors with insertion fraction $γ$ and deletion fraction $δ$ follow from insertion-only lower bounds at $γ+Lδ$ and deletion-only upper bounds at $δ+γ/L$. For binary unique decoding, we strictly improve Levenshtein's classical asymptotic rate lower bound for every deletion fraction $0<δ<1/2$ for which the classical rate expression is nonnegative. At $δ=0.1$, the lower bound increases from approximately $0.162009$ to $0.180431$, a relative increase of about $11.37\%$. Our framework also yields insertion and deletion lower bounds for every fixed list size. The existence proofs combine the Lovász local lemma with sampling from words having a specified number of runs, where a run is a maximal block of equal symbols. Generating functions provide refined bounds on the probability that $L+1$ sampled words share an allowed received word. We also derive a Levenshtein-type upper bound by run counting and, separately, a higher-order Elias bound using intersections and unions of the position sets used to embed $L+1$ codewords in a common supersequence. The latter recovers Yasunaga's asymptotic unique-decoding bound at $L=1$ and strictly improves the Haeupler--Shahrasbi--Sudan insertion bound for every fixed $L$ and $0<γ<q-1$. Numerical comparisons quantify the gains and remaining gaps.

cs.IT

Random Algebraic Geometry Codes Approach the Half-Singleton Bound for Insertions and Deletions

In this paper, we study the performance of algebraic geometry (AG) codes against adversarial insertion-deletion (insdel) errors. The half-Singleton bound states that an $[n,k]_q$ linear code can correct at most $n-2k+1$ insdel errors. It was recently proven that random Reed-Solomon codes approach this bound. However, these constructions require the field size $q$ to grow linearly with the code length $n$. We overcome this barrier by extending the probabilistic analysis of general linear insdel codes to AG codes. We demonstrate that curves with many rational points allow for nearly optimal codes over significantly smaller alphabets. We prove the following main asymptotic results: (1) For general smooth complete curves of fixed genus, random AG codes are nearly optimal, that is, they can correct $(1-\varepsilon)n-2k$ insdel errors with high probability over linear-sized fields ($q=Θ(n)$). (2) By utilizing Hermitian curves, we achieve this optimality over sublinear fields of size $q=Θ(n^{2/3})$, breaking the linear field size barrier. (3) Using asymptotically optimal García-Stichtenoth towers, we prove the existence of random AG codes that approach the half-Singleton bound with high probability over fields of size $q=2^{O_R(1/\varepsilon^2)}$, independent of $n$.

cs.IT

The Optimal Asymptotic Rate of Generalized Covering Codes

Let $G_q$ be an alphabet of size $q\geq2$. We determine the optimal asymptotic rate of generalized covering codes $C\subseteq G_q^n$, whose covering centers in $G_q^{t\times n}$ are constrained to the product form $C^t$. For every fixed integer $t\geq1$ and every $ρ\in[0,1]$, we prove that \[ κ_t(ρ,q)= \begin{cases} 1-H_{q^t}(ρ),&0\leqρ<1-q^{-t},\\ 0,&1-q^{-t}\leqρ\leq1, \end{cases} \] where $κ_t(ρ,q)$ denotes the minimum asymptotic rate $n^{-1}\log_q|C|$ among codes whose $t$-th covering radius is at most $ρn$, and $H_{q^t}$ is the $q^t$-ary entropy function. When $q$ is a prime power, we prove that the same formula holds under the additional requirement that $C\leq\mathbb F_q^n$. Thus, both the product-form constraint and linearity are asymptotically cost-free: the resulting rate is the ordinary sphere-covering rate over an alphabet of size $q^t$. This extends the recent $t=2$ result of Elimelech and Schwartz for codes without a linearity constraint and the classical $t=1$ result of Cohen and Frankl for linear codes, thereby resolving both open problems posed by Elimelech and Schwartz. Our proofs are probabilistic and combine tools from information theory and probabilistic combinatorics, including the method of types, Janson's inequality, the second-moment method, and a structured alteration argument. Direct applications of Janson's inequality and the second-moment method are obstructed by highly dependent pairs of candidate error matrices. We overcome this obstruction by restricting the errors to a balanced exact-type class of optimal exponential size. Standard type-class estimates, together with Shearer's inequality, then give the required bounds on the number of error-matrix pairs whose selected rows have a prescribed difference.

cs.IT

Moment-based linear programming bounds for locally recoverable codes

In this paper we derive new Delsarte-type linear programming bounds for $q$-ary $(r,δ)$-locally recoverable codes (LRCs) with three attributes: first, the variable set is comparable in size to that of the classical Delsarte LP; second, our LP exploits the higher-order information forced by the local-distance condition through order \(δ-2\), in the sense that for nondegenerate linear codes, its balanced base part gives exactly the same dimension bound as the symmetrized refined-weight LP of Gruica, Jany, and Ravagnani, while the additional constraints, nonvacuous whenever $δ\ge 3$, give a further strengthening; and third, it applies to general $(r,δ)$-LRCs, linear and nonlinear alike. Extensive computations over binary and ternary alphabets show that the convex-hull LP yields improvements not captured by the previous LP and often sharpens the shortening and generalized Singleton bounds.

cs.IT

Sharp Bounds and New Constructions for Single-Error Detection and Correction in Analog Codes

We study single-error detection and correction for analog codes over $\mathbb{R}$. The key performance measures are the parameters $Γ_1(\mathcal{C})$ and $Γ_2(\mathcal{C})$, which quantify, respectively, the minimum separation required between large outlying errors that must be detected or located and the magnitude of tolerable perturbations. First, we prove that every real linear $[n,k]$ code $\mathcal{C}$ satisfies \[ Γ_1(\mathcal{C})\ge 2\left\lceil\frac{n}{n-k}\right\rceil. \] Moreover, when $k=n-2$, we prove that every real linear $[n,n-2]$ code $\mathcal{C}$ satisfies \[ Γ_2(\mathcal{C})\ge \frac{1}{\sin^2(π/2n)}. \] Together, these two lower bounds settle all four open problems of Roth concerning the optimality of single-error-detecting and single-error-correcting analog codes. The proof of the first bound is based on a double-induction argument, while the proof of the second combines a zonotope-based geometric characterization of $Γ_2(\mathcal{C})$ with a cyclic sine-product inequality. In addition, we construct analog codes with higher fixed redundancy and show that, for every fixed $r\ge 2$, there exists a class of linear $[n,\ge n-r]$ codes over $\mathbb{R}$ such that \[ Γ_2(\mathcal{C})\le O\left(n^{1+\frac{1}{r-1}}\right). \] This gives a new upper bound in the fixed-redundancy regime, which was not covered by previously known constructions.

cs.IT

Elias-type Bounds for Codes in the Symmetric Limited-Magnitude Error Channel

We study perfect error-correcting codes in $\mathbb{Z}^n$ for the symmetric limited-magnitude error channel, where at most $e$ coordinates of an integer vector may be altered by a value whose magnitude is at most $s$. Geometrically, such codes correspond to tilings of $\mathbb{Z}^n$ by the symmetric limited-magnitude error ball $\mathcal{B}(n,e,s,s)$. Given $n$ and $s$, we adapt the geometric ideas underlying the Elias bound for the Hamming metric to the distance $d_s$ tailed for this channel, and derive new necessary conditions on $e$ for the existence of perfect codes / tilings, without assuming any lattice structure. Our main results identify two distinct regimes depending on the error magnitude. For small error magnitudes ($s \in \{1, 2\}$), we prove that if the number of correctable errors does not exceed a certain fraction of $n$, then it is asymptotically bounded by $e = \mathcal{O}(\sqrt{n \log n})$. In contrast, for larger magnitudes ($s \geq 3$), we establish a significantly sharper bound of $e < \sqrt{12.36n}$, which holds without any restriction on $e$ being below a given fraction of $n$. Finally, by extending our method to non-perfect codes, we derive an upper bound on packing density, showing that for codes correcting a linear or $Ω(\sqrt{n})$ number of errors, the density is bounded by a factor inversely proportional to the error magnitude $s$.

cs.IT

On Lattice Tilings of Asymmetric Limited-Magnitude Balls $\cB(n,2,m,m-1)$

Limited-magnitude errors modify a transmitted integer vector in at most $t$ entries, where each entry can increase by at most $\kp$ or decrease by at most $\km$. This channel model is particularly relevant to applications such as flash memories and DNA storage. A perfect code for this channel is equivalent to a tiling of $\Z^n$ by asymmetric limited-magnitude balls $\cB(n,t,\kp,\km)$. In this paper, we focus on the case where $t=2$ and $\km=\kp-1$, and we derive necessary conditions on $m$ and $n$ for the existence of a lattice tiling of $\cB(n,2,m,m-1)$. Specifically, we prove that if such a tiling exists, then either $4\leq m \leq 512$ and $n<7.23m+4$, or $m>512$ and $n<4m$. In particular, for $m=2$ and $m=3$, we show that no lattice tiling of $\cB(n,2,2,1)$ or $\cB(n,2,3,2)$ exists for any $n\geq 3$.

math.CO

Linear Network Coding for Robust Function Computation and Its Applications in Distributed Computing

We investigate linear network coding in the context of robust function computation, where a sink node is tasked with computing a target function of messages generated at multiple source nodes. In a previous work, a new distance measure was introduced to evaluate the error tolerance of a linear network code for function computation, along with a Singleton-like bound for this distance. In this paper, we first present a minimum distance decoder for these linear network codes. We then focus on the sum function and the identity function, showing that in any directed acyclic network there are two classes of linear network codes for these target functions, respectively, that attain the Singleton-like bound. Additionally, we explore the application of these codes in distributed computing and design a distributed gradient coding scheme in a heterogeneous setting, optimizing the trade-off between straggler tolerance, computation cost, and communication cost. This scheme can also defend against Byzantine attacks.

cs.IT

Combinatorial alphabet-dependent bounds for insdel codes

Error-correcting codes resilient to synchronization errors such as insertions and deletions are known as insdel codes. Due to their important applications in DNA storage and computational biology, insdel codes have recently become a focal point of research in coding theory. In this paper, we present several new combinatorial upper and lower bounds on the maximum size of $q$-ary insdel codes. Our main upper bound is a sphere-packing bound obtained by solving a linear programming (LP) problem. It improves upon previous results for cases when the distance $d$ or the alphabet size $q$ is large. Our first lower bound is derived from a connection between insdel codes and matchings in special hypergraphs. This lower bound, together with our upper bound, shows that for fixed block length $n$ and edit distance $d$, when $q$ is sufficiently large, the maximum size of insdel codes is $ \frac{q^{n-\frac{d}{2}+1}}{{n\choose \frac{d}{2}-1}}(1 \pm o(1))$. The second lower bound refines Alon et al.'s recent logarithmic improvement on Levenshtein's GV-type bound and extends its applicability to large $q$ and $d$.

math.CO

Linearized Reed-Solomon Codes with Support-Constrained Generator Matrix and Applications in Multi-Source Network Coding

Linearized Reed-Solomon (LRS) codes are evaluation codes based on skew polynomials. They achieve the Singleton bound in the sum-rank metric and therefore are known as maximum sum-rank distance (MSRD) codes. In this work, we give necessary and sufficient conditions for the existence of MSRD codes with a support-constrained generator matrix. The conditions on the support constraints are identical to those for MDS codes and MRD codes. The required field size for an $[n,k]_{q^m}$ LRS codes with support-constrained generator matrix is $q\geq \ell+1$ and $m\geq \max_{l\in[\ell]}\{k-1+\log_qk, n_l\}$, where $\ell$ is the number of blocks and $n_l$ is the size of the $l$-th block. The special cases of the result coincide with the known results for Reed-Solomon codes and Gabidulin codes. For the support constraints that do not satisfy the necessary conditions, we derive the maximum sum-rank distance of a code whose generator matrix fulfills the constraints. Such a code can be constructed from a subcode of an LRS code with a sufficiently large field size. Moreover, as an application in network coding, the conditions can be used as constraints in an integer programming problem to design distributed LRS codes for a distributed multi-source network.

cs.IT

Multiple-Error-Correcting Codes for Analog Computing on Resistive Crossbars

Error-correcting codes over the real field are studied which can locate outlying computational errors when performing approximate computing of real vector--matrix multiplication on resistive crossbars. Prior work has concentrated on locating a single outlying error and, in this work, several classes of codes are presented which can handle multiple errors. It is first shown that one of the known constructions, which is based on spherical codes, can in fact handle multiple outlying errors. A second family of codes is then presented with $\zeroone$~parity-check matrices which are sparse and disjunct; such matrices have been used in other applications as well, especially in combinatorial group testing. In addition, a certain class of the codes that are obtained through this construction is shown to be efficiently decodable. As part of the study of sparse disjunct matrices, this work also contains improved lower and upper bounds on the maximum Hamming weight of the rows in such matrices.

cs.IT

Reconstruction from Noisy Substrings

This paper studies the problem of encoding messages into sequences which can be uniquely recovered from some noisy observations about their substrings. The observed reads comprise consecutive substrings with some given minimum overlap. This coded reconstruction problem has applications to DNA storage. We consider both single-strand reconstruction codes and multi-strand reconstruction codes, where the message is encoded into a single strand or a set of multiple strands, respectively. Various parameter regimes are studied. New codes are constructed, some of whose rates asymptotically attain the upper bounds.

cs.IT

Perfect Codes Correcting a Single Burst of Limited-Magnitude Errors

Motivated by applications to DNA-storage, flash memory, and magnetic recording, we study perfect burst-correcting codes for the limited-magnitude error channel. These codes are lattices that tile the integer grid with the appropriate error ball. We construct two classes of such perfect codes correcting a single burst of length $2$ for $(1,0)$-limited-magnitude errors, both for cyclic and non-cyclic bursts. We also present a generic construction that requires a primitive element in a finite field with specific properties. We then show that in various parameter regimes such primitive elements exist, and hence, infinitely many perfect burst-correcting codes exist.

cs.IT

Sequence Reconstruction for Limited-Magnitude Errors

Motivated by applications to DNA storage, we study reconstruction and list-reconstruction schemes for integer vectors that suffer from limited-magnitude errors. We characterize the asymptotic size of the intersection of error balls in relation to the code's minimum distance. We also devise efficient reconstruction algorithms for various limited-magnitude error parameter ranges. We then extend these algorithms to the list-reconstruction scheme, and show the trade-off between the asymptotic list size and the number of required channel outputs. These results apply to all codes, without any assumptions on the code structure. Finally, we also study linear reconstruction codes with small intersection, as well as show a connection to list-reconstruction codes for the tandem-duplication channel.

cs.IT

On the Generalized Covering Radii of Reed-Muller Codes

We study generalized covering radii, a fundamental property of linear codes that characterizes the trade-off between storage, latency, and access in linear data-query protocols such as PIR. We prove lower and upper bounds on the generalized covering radii of Reed-Muller codes, as well as finding their exact value in certain extreme cases. With the application to linear data-query protocols in mind, we also construct a covering algorithm that gets as input a set of points in space, and find a corresponding set of codewords from the Reed-Muller code that are jointly not farther away from the input than the upper bound on the generalized covering radius of the code. We prove that the algorithm runs in time that is polynomial in the code parameters.

cs.IT

Improved Coding over Sets for DNA-Based Data Storage

Error-correcting codes over sets, with applications to DNA storage, are studied. The DNA-storage channel receives a set of sequences, and produces a corrupted version of the set, including sequence loss, symbol substitution, symbol insertion/deletion, and limited-magnitude errors in symbols. Various parameter regimes are studied. New bounds on code parameters are provided, which improve upon known bounds. New codes are constructed, at times matching the bounds up to lower-or der terms or small constant factors.

cs.IT

On the Gap between Scalar and Vector Solutions of Generalized Combination Networks

We study scalar-linear and vector-linear solutions of the generalized combination network. We derive new upper and lower bounds on the maximum number of nodes in the middle layer, depending on the network parameters and the alphabet size. These bounds improve and extend the parameter range of known bounds. Using these new bounds we present a lower bound and an upper bound on the gap in the alphabet size between optimal scalar-linear and optimal vector-linear network coding solutions. For a fixed network structure, while varying the number of middle-layer nodes $r$, the asymptotic behavior of the upper and lower bounds shows that the gap is in $Θ(\log(r))$.

cs.IT

On Tilings of Asymmetric Limited-Magnitude Balls

We study whether an asymmetric limited-magnitude ball may tile $\mathbb{Z}^n$. This ball generalizes previously studied shapes: crosses, semi-crosses, and quasi-crosses. Such tilings act as perfect error-correcting codes in a channel which changes a transmitted integer vector in a bounded number of entries by limited-magnitude errors. A construction of lattice tilings based on perfect codes in the Hamming metric is given. Several non-existence results are proved, both for general tilings, and lattice tilings. A complete classification of lattice tilings for two certain cases is proved.

math.CO