SearcharxivSearch

arXiv subjects

Iwan Duursma

Publications and source records attributed to Iwan Duursma.

At least 19 recordsLinked to original sources

Accelerating Polarization via Alphabet Extension

Polarization is an unprecedented coding technique in that it not only achieves channel capacity, but also does so at a faster speed of convergence than any other coding technique. This speed is measured by the ``scaling exponent'' and its importance is three-fold. Firstly, estimating the scaling exponent is challenging and demands a deeper understanding of the dynamics of communication channels. Secondly, scaling exponents serve as a benchmark for different variants of polar codes that helps us select the proper variant for real-life applications. Thirdly, the need to optimize for the scaling exponent sheds light on how to reinforce the design of polar codes. In this paper, we generalize the binary erasure channel (BEC), the simplest communication channel and the protagonist of many coding theory studies, to the ``tetrahedral erasure channel'' (TEC). We then invoke Mori--Tanaka's $2 \times 2$ matrix over GF$(4)$ to construct polar codes over TEC. Our main contribution is showing that the dynamic of TECs converges to an almost--one-parameter family of channels, which then leads to an upper bound of $3.328$ on the scaling exponent. This is the first non-binary matrix whose scaling exponent is upper-bounded. It also polarizes BEC faster than all known binary matrices up to $23 \times 23$ in size. Our result indicates that expanding the alphabet is a more effective and practical alternative to enlarging the matrix in order to achieve faster polarization.

cs.IT

Parity-Checked Strassen Algorithm

To multiply astronomic matrices using parallel workers subject to straggling, we recommend interleaving checksums with some fast matrix multiplication algorithms. Nesting the parity-checked algorithms, we weave a product code flavor protection. Two demonstrative configurations are as follows: (A) $9$ workers multiply two $2\times 2$ matrices; each worker multiplies two linear combinations of entries therein. Then the entry products sent from any $8$ workers suffice to assemble the matrix product. (B) $754$ workers multiply two $9\times 9$ matrices. With empirical frequency $99.8\%$, $729$ workers suffice, wherein $729$ is the complexity of the schoolbook algorithm. In general, we propose probability-wisely favorable configurations whose numbers of workers are close to, if not less than, the thresholds of other codes (e.g., entangled polynomial code and PolyDot code). Our proposed scheme applies recursively, respects worker locality, incurs moderate pre- and post-processes, and extends over small finite fields.

cs.IT

Multilinear Algebra for Minimum Storage Regenerating Codes

An $(n, k, d, α)$-MSR (minimum storage regeneration) code is a set of $n$ nodes used to store a file. For a file of total size $kα$, each node stores $α$ symbols, any $k$ nodes recover the file, and any $d$ nodes can repair any other node via each sending out $α/(d-k+1)$ symbols. In this work, we explore various ways to re-express the infamous product-matrix construction using skew-symmetric matrices, polynomials, symmetric algebras, and exterior algebras. We then introduce a multilinear algebra foundation to produce $\bigl(n, k, \frac{(k-1)t}{t-1}, \binom{k-1}{t-1}\bigr)$-MSR codes for general $t\geq2$. At the $t=2$ end, they include the product-matrix construction as a special case. At the $t=k$ end, we recover determinant codes of mode $m=k$; further restriction to $n=k+1$ makes it identical to the layered code at the MSR point. Our codes' sub-packetization level---$α$---is independent of $n$ and small. It is less than $L^{2.8(d-k+1)}$, where $L$ is Alrabiah--Guruswami's lower bound on $α$. Furthermore, it is less than other MSR codes' $α$ for a subset of practical parameters. We offer hints on how our code repairs multiple failures at once.

cs.IT

Multilinear Algebra for Distributed Storage

An $(n, k, d, α, β, M)$-ERRC (exact-repair regenerating code) is a collection of $n$ nodes used to store a file. For a file of total size $M$, each node stores $α$ symbols, any $k$ nodes recover the file, and any $d$ nodes repair any other node via sending out $β$ symbols. We establish a multilinear algebra foundation to assemble $(n, k, d, α, β, M)$-ERRCs for all meaningful $(n, k, d)$ tuples. Our ERRCs tie the $α/M$-versus-$β/M$ trade-off with cascade codes, the best known construction for this trade-off. We give directions on how these ERRCs repair multiple failures.

cs.IT

Repairing Reed-Solomon Codes With Multiple Erasures

Despite their exceptional error-correcting properties, Reed-Solomon codes have been overlooked in distributed storage applications due to the common belief that they have poor repair bandwidth: A naive repair approach would require the whole file to be reconstructed in order to recover a single erased codeword symbol. In a recent work, Guruswami and Wootters (STOC'16) proposed a single-erasure repair method for Reed-Solomon codes that achieves the optimal repair bandwidth amongst all linear encoding schemes. Their key idea is to recover the erased symbol by collecting a sufficiently large number of its traces, each of which can be constructed from a number of traces of other symbols. We extend the trace collection technique to cope with two and three erasures.

cs.IT

Johnson Graph Codes

We define a Johnson graph code as a subspace of labelings of the vertices in a Johnson graph with the property that labelings are uniquely determined by their restriction to vertex neighborhoods specified by the parameters of the code. We give a construction and main properties for the codes and show their role in the concatenation of layered codes that are used in distributed storage systems. A similar class of codes for the Hamming graph is discussed in an appendix. Codes of the latter type are in general different from affine Reed-Muller codes, but for the special case of the hypercube they agree with binary Reed-Muller codes.

math.CO

Polar Codes' Simplicity, Random Codes' Durability

Over any discrete memoryless channel, we build codes such that: for one, their block error probabilities and code rates scale like random codes'; and for two, their encoding and decoding complexities scale like polar codes'. Quantitatively, for any constants $π,ρ>0$ such that $π+2ρ<1$, we construct a sequence of error correction codes with block length $N$ approaching infinity, block error probability $\exp(-N^π)$, code rate $N^{-ρ}$ less than the Shannon capacity, and encoding and decoding complexity $O(N\log N)$ per code block. The putative codes take uniform $ς$-ary messages for sender's choice of prime $ς$. The putative codes are optimal in the following manner: Should $π+2ρ>1$, no such codes exist for generic channels regardless of alphabet and complexity.

cs.IT

Isometry-Dual Flags of AG Codes

Consider a complete flag $\{0\} = C_0 < C_1 < \cdots < C_n = \mathbb{F}^n$ of one-point AG codes of length $n$ over the finite field $\mathbb{F}$. The codes are defined by evaluating functions with poles at a given point $Q$ in points $P_1,\dots,P_n$ distinct from $Q$. A flag has the isometry-dual property if the given flag and the corresponding dual flag are the same up to isometry. For several curves, including the projective line, Hermitian curves, Suzuki curves, Ree curves, and the Klein curve over the field of eight elements, the maximal flag, obtained by evaluation in all rational points different from the point $Q$, is self-dual. More generally, we ask whether a flag obtained by evaluation in a proper subset of rational points is isometry-dual. In [3] it is shown, for a curve of genus $g$, that a flag of one-point AG codes defined with a subset of $n > 2g+2$ rational points is isometry-dual if and only if the last code $C_n$ in the flag is defined with functions of pole order at most $n+2g-1$. Using a different approach, we extend this characterization to all subsets of size $n \geq 2g+2$. Moreover we show that this is best possible by giving examples of isometry-dual flags with $n=2g+1$ such that $C_n$ is generated by functions of pole order at most $n+2g-2$. We also prove a necessary condition, formulated in terms of maximum sparse ideals of the Weierstrass semigroup of $Q$, under which a flag of punctured one-point AG codes inherits the isometry-dual property from the original unpunctured flag.

cs.IT

Log-logarithmic Time Pruned Polar Coding

A pruned variant of polar coding is proposed for binary erasure channels. For sufficiently small $\varepsilon>0$, we construct a series of capacity achieving codes with block length $N=\varepsilon^{-5}$, code rate $R=\text{Capacity}-\varepsilon$, error probability $P=\varepsilon$, and encoding and decoding time complexity $\text{bC}=O(\log\left|\log\varepsilon\right|)$ per information bit. The given per-bit complexity $\text{bC}$ is log-logarithmic in $N$, in $\text{Capacity}-R$, and in $P$; no known family of codes possesses this property. It is also the second lowest $\text{bC}$ after repeat-accumulate codes and their variants. While random codes and classical polar codes are the only two families of capacity-achieving codes whose $N$, $R$, $P$, and $\text{bC}$ were written down as explicit functions, our construction gives the third family. Then we generalize the result to: Fix a prime $q$ and fix a $q$-ary-input discrete symmetric memoryless channel. For sufficiently small $\varepsilon>0$, we construct a series of capacity achieving codes with block length $N=\varepsilon^{-O(1)}$, code rate $R=\text{Capacity}-\varepsilon$, error probability $P=\varepsilon$, and encoding and decoding time complexity $\text{bC}=O(\log\left|\log\varepsilon\right|)$ per information bit. The later construction gives the fastest family of capacity-achieving codes to date on those channels.

cs.IT

Codes with Locality in the Rank and Subspace Metrics

We extend the notion of locality from the Hamming metric to the rank and subspace metrics. Our main contribution is to construct a class of array codes with locality constraints in the rank metric. Our motivation for constructing such codes stems from designing codes for efficient data recovery from correlated and/or mixed (i.e., complete and partial) failures in distributed storage systems. Specifically, the proposed local rank-metric codes can recover locally from 'crisscross errors and erasures', which affect a limited number of rows and/or columns of the storage system. We also derive a Singleton-like upper bound on the minimum rank distance of (linear) codes with 'rank-locality' constraints. Our proposed construction achieves this bound for a broad range of parameters. The construction builds upon Tamo and Barg's method for constructing locally repairable codes with optimal minimum Hamming distance. Finally, we construct a class of constant-dimension subspace codes (also known as Grassmannian codes) with locality constraints in the subspace metric. The key idea is to show that a Grassmannian code with locality can be easily constructed from a rank-metric code with locality by using the lifting method proposed by Silva et al. We present an application of such codes for distributed storage systems, wherein nodes are connected over a network that can introduce errors and erasures.

cs.IT

Log-logarithmic Time Pruned Polar Coding on Binary Erasure Channels

A pruned variant of polar coding is reinvented for all binary erasure channels. For small $\varepsilon>0$, we construct codes with block length $\varepsilon^{-5}$, code rate $\text{Capacity}-\varepsilon$, error probability $\varepsilon$, and encoding and decoding time complexity $O(N\log|\log\varepsilon|)$ per block, equivalently $O(\log|\log\varepsilon|)$ per information bit (Propositions 5 to 8). This result also follows if one applies systematic polar coding [Arıkan 10.1109/LCOMM.2011.061611.110862] with simplified successive cancelation decoding [Alamdar-Yazdi-Kschischang 10.1109/LCOMM.2011.101811.111480], and then analyzes the performance using [Guruswami-Xia arXiv:1304.4321] or [Mondelli-Hassani-Urbanke arXiv:1501.02444].

cs.IT

Polar-like Codes and Asymptotic Tradeoff among Block Length, Code Rate, and Error Probability

A general framework is proposed that includes polar codes over arbitrary channels with arbitrary kernels. The asymptotic tradeoff among block length $N$, code rate $R$, and error probability $P$ is analyzed. Given a tradeoff between $N,P$ and a tradeoff between $N,R$, we return an interpolating tradeoff among $N,R,P$ (Theorem 5). $\def\Capacity{\text{Capacity}}$Quantitatively, if $P=\exp(-N^{β^*})$ is possible for some $β^*$ and if $R=\Capacity-N^{1/μ^*}$ is possible for some $1/μ^*$, then $(P,R)=(\exp(-N^{β'}),\Capacity-N^{-1/μ'})$ is possible for some pair $(β',1/μ')$ determined by $β^*$, $1/μ^*$, and auxiliary information. In fancy words, an error exponent regime tradeoff plus a scaling exponent regime tradeoff implies a moderate deviations regime tradeoff. The current world records are: [arXiv:1304.4321][arXiv:1501.02444][arXiv:1806.02405] analyzing Arıkan's codes over BEC; [arXiv:1706.02458] analyzing Arıkan's codes over AWGN; and [arXiv:1802.02718][arXiv:1810.04298] analyzing general codes over general channels. An attempt is made to generalize all at once (Section IX). As a corollary, a grafted variant of polar coding almost catches up the code rate and error probability of random codes with complexity slightly larger than $N\log N$ over BEC. In particular, $(P,R)=(\exp(-N^{.33}),\Capacity-N^{-.33})$ is possible (Corollary 10). In fact, all points in this triangle are possible $(β',1/μ')$-pairs. $$ \require{enclose} \def\r{\phantom{\Rule{4em}{1em}{1em}}} \enclose{}\r^\llap{(0,1/2)}_\llap{(0,0)} \enclose{left,bottom,downdiagonalstrike}\r_\rlap{(1,0)} \enclose{}\r $$

cs.IT

Polar Code Moderate Deviation: Recovering the Scaling Exponent

In 2008 Arikan proposed polar coding [arXiv:0807.3917] which we summarize as follows: (a) From the root channel $W$ synthesize recursively a series of channels $W_N^{(1)},\dotsc,W_N^{(N)}$. (b) Select sophisticatedly a subset $A$ of synthetic channels. (c) Transmit information using synthetic channels indexed by $A$ and freeze the remaining synthetic channels. Arikan gives each synthetic channel a score (called the Bhattacharyya parameter) that determines whether it should be selected or frozen. As $N$ grows, a majority of the scores are either very high or very low, i.e., they polarize. By characterizing how fast they polarize, Arikan showed that polar coding is able to produce a series of codes that achieve capacity on symmetric binary-input memoryless channels. In measuring how the scores polarize the relation among block length, gap to capacity, and block error probability are studied. In particular, the error exponent regime fixes the gap to capacity and varies the other two. The scaling exponent regime fixes the block error probability and varies the other two. The moderate deviation regime varies all three factors at once. The latest result [arxiv:1501.02444, Theorem 7] in the moderate deviation regime does not imply the scaling exponent regime as a special case. We give a result that does. (See Corollary 8.)

cs.IT

On the I/O Costs of Some Repair Schemes for Full-Length Reed-Solomon Codes

Network transfer and disk read are the most time consuming operations in the repair process for node failures in erasure-code-based distributed storage systems. Recent developments on Reed-Solomon codes, the most widely used erasure codes in practical storage systems, have shown that efficient repair schemes specifically tailored to these codes can significantly reduce the network bandwidth spent to recover single failures. However, the I/O cost, that is, the number of disk reads performed in these repair schemes remains largely unknown. We take the first step to address this gap in the literature by investigating the I/O costs of some existing repair schemes for full-length Reed-Solomon codes.

cs.IT

Repairing Reed-Solomon Codes With Two Erasures

Despite their exceptional error-correcting properties, Reed-Solomon (RS) codes have been overlooked in distributed storage applications due to the common belief that they have poor repair bandwidth: A naive repair approach would require the whole file to be reconstructed in order to recover a single erased codeword symbol. In a recent work, Guruswami and Wootters (STOC'16) proposed a single-erasure repair method for RS codes that achieves the optimal repair bandwidth amongst all linear encoding schemes. We extend their trace collection technique to cope with two erasures.

cs.IT

Smooth Embeddings for the Suzuki and Ree Curves

The Hermitian, Suzuki and Ree curves form three special families of curves with unique properties. They arise as the Deligne-Lusztig varieties of dimension one and their automorphism groups are the algebraic groups of type 2A2, 2B2 and 2G2, respectively. For the Hermitian and Suzuki curves very ample divisors are known that yield smooth projective embeddings of the curves. In this paper we establish a very ample divisor for the Ree curves. Moreover, for all three families of curves we find a symmetric set of equations for a smooth projective model, in dimensions 2, 4 and 13, respectively. Using the smooth model we determine the unknown nongaps in the Weierstrass semigroup for a rational point on the Ree curve.

math.AG

Using concatenated algebraic geometry codes in channel polarization

Polar codes were introduced by Arikan in 2008 and are the first family of error-correcting codes achieving the symmetric capacity of an arbitrary binary-input discrete memoryless channel under low complexity encoding and using an efficient successive cancellation decoding strategy. Recently, non-binary polar codes have been studied, in which one can use different algebraic geometry codes to achieve better error decoding probability. In this paper, we study the performance of binary polar codes that are obtained from non-binary algebraic geometry codes using concatenation. For binary polar codes (i.e. binary kernels) of a given length $n$, we compare numerically the use of short algebraic geometry codes over large fields versus long algebraic geometry codes over small fields. We find that for each $n$ there is an optimal choice. For binary kernels of size up to $n \leq 1,800$ a concatenated Reed-Solomon code outperforms other choices. For larger kernel sizes concatenated Hermitian codes or Suzuki codes will do better.

cs.IT

Distributed Reed-Solomon Codes for Simple Multiple Access Networks

We consider a simple multiple access network in which a destination node receives information from multiple sources via a set of relay nodes. Each relay node has access to a subset of the sources, and is connected to the destination by a unit capacity link. We also assume that $z$ of the relay nodes are adversarial. We propose a computationally efficient distributed coding scheme and show that it achieves the full capacity region for up to three sources. Specifically, the relay nodes encode in a distributed fashion such that the overall codewords received at the destination are codewords from a single Reed-Solomon code.

cs.IT