SearcharxivSearch

arXiv subjects

Krishna Gopal Benerjee

Publications and source records attributed to Krishna Gopal Benerjee.

16 recordsLinked to original sources

On Algebraic Approaches for DNA Codes with Multiple Constraints

DNA strings and their properties are widely studied since last 20 years due to its applications in DNA computing. In this area, one designs a set of DNA strings (called DNA code) which satisfies certain thermodynamic and combinatorial constraints such as reverse constraint, reverse-complement constraint, $GC$-content constraint and Hamming constraint. However recent applications of DNA codes in DNA data storage resulted in many new constraints on DNA codes such as avoiding tandem repeats constraint (a generalization of non-homopolymer constraint) and avoiding secondary structures constraint. Therefore, in this chapter, we introduce DNA codes with recently developed constraints. In particular, we discuss reverse, reverse-complement, $GC$-content, Hamming, uncorrelated-correlated, thermodynamic, avoiding tandem repeats and avoiding secondary structures constraints. DNA codes are constructed using various approaches such as algebraic, computational, and combinatorial. In particular, in algebraic approaches, one uses a finite ring and a map to construct a DNA code. Most of such approaches does not yield DNA codes with high Hamming distance. In this chapter, we focus on algebraic constructions using maps (usually an isometry on some finite ring) which yields DNA codes with high Hamming distance. We focus on non-cyclic DNA codes. We briefly discuss various metrics such as Gau distance, Non-Homopolymer distance etc. We discuss about algebraic constructions of families of DNA codes that satisfy multiple constraints and/or properties. Further, we also discuss about algebraic bounds on DNA codes with multiple constraints. Finally, we present some open research directions in this area.

cs.IT

ISI-Aware Code Design: A Linear Approach Towards Reliable Molecular Communication

Intersymbol Interference (ISI) is a major bottleneck in Molecular Communication via Diffusion (MCvD), degrading system performance. This paper introduces two families of linear channel codes to mitigate ISI: Zero Pad Zero Start (ZPZS) and Zero Pad (ZP) codes, ensuring that each codeword avoids consecutive bit-1s. The ZPZS and ZP codes are then combined to form a binary ZP code, offering a higher code rate than linear ZP codes and allowing simple decoding via the Majority Location Rule (MLR). Additionally, a Leading One Zero Pad (LOZP) code is proposed, which relaxes zero-padding constraints by prioritizing the placement of bit-1s, achieving a higher rate than ZP. A closed-form expression is derived to compute expected ISI, showing it depends on the average bit-1 density in the codewords. ISI and Bit Error Rate (BER) performance are evaluated under two MCvD channel models: (i) without refresh, where past bits persist longer, and (ii) with refresh, where the channel is cleared after each reception. Results show that the LOZP code performs better in the refresh channel due to initial bit-1 placement, while ZP excels without refresh by reducing average bit-1 density. The asymptotic upper bound on code rate illustrates a trade-off between ISI and rate. Simulations demonstrate that ZP and LOZP codes improve BER by controlling bit-1 positions and density, providing better reliability in ISI-dominated regimes compared to conventional error-correcting codes.

cs.IT

On Designing Novel ISI-Reducing Single Error Correcting Codes in an MCvD System

Intersymbol Interference (ISI) has a detrimental impact on any Molecular Communication via Diffusion (MCvD) system. Also, the receiver noise can severely degrade the MCvD channel performance. However, the channel codes proposed in the literature for the MCvD system have only addressed one of these two challenges independently. In this paper, we have designed single Error Correcting Codes in an MCvD system with channel memory and noise. We have also provided encoding and decoding algorithms for the proposed codes, which are simple to follow despite having a non-linear code construction. Finally, through simulation results, we show that the proposed single ECCs, for given code parameters, perform better than the existing codes in the literature in combating the effect of ISI in the channel and improving the average Bit Error Rate (BER) performance in a noisy channel.

cs.IT

Bounds on Size of Homopolymer Free Codes

For any given alphabet of size $q$, a Homopolymer Free code (HF code) refers to an $(n, M, d)_q$ code of length $n$, size $M$ and minimum Hamming distance $d$, where all the codewords are homopolymer free sequences. For any given alphabet, this work provides upper and lower bounds on the maximum size of any HF code using Sphere Packing bound and Gilbert-Varshamov bound. Further, upper and lower bounds on the maximum size of HF codes for various HF code families are calculated. Also, as a specific case, upper and lower bounds are obtained on the maximum size of homopolymer free DNA codes.

cs.IT

On DNA Codes Over the Non-Chain Ring $\mathbb{Z}_4+u\mathbb{Z}_4+u^2\mathbb{Z}_4$ with $u^3=1$

In this paper, we present a novel design strategy of DNA codes with length $3n$ over the non-chain ring $R=\mathbb{Z}_4+u\mathbb{Z}_4+u^2\mathbb{Z}_4$ with $64$ elements and $u^3=1$, where $n$ denotes the length of a code over $R$. We first study and analyze a distance conserving map defined over the ring $R$ into the length-$3$ DNA sequences. Then, we derive some conditions on the generator matrix of a linear code over $R$, which leads to a DNA code with reversible, reversible-complement, homopolymer $2$-run-length, and $\frac{w}{3n}$-GC-content constraints for integer $w$ ($0\leq w\leq 3n$). Finally, we propose a new construction of DNA codes using Reed-Muller type generator matrices. This allows us to obtain DNA codes with reversible, reversible-complement, homopolymer $2$-run-length, and $\frac{2}{3}$-GC-content constraints.

cs.IT

DNA Codes over the Ring $\mathbb{Z}_4 + w\mathbb{Z}_4$

In this present work, we generalize the study of construction of DNA codes over the rings $\mathcal{R}_θ=\mathbb{Z}_4+w\mathbb{Z}_4$, $w^2 = θ$ for $θ\in \mathbb{Z}_4+w\mathbb{Z}_4$. Rigorous study along with characterization of the ring structures is presented. We extend the Gau map and Gau distance, defined in \cite{DKBG}, over all the $16$ rings $\mathcal{R}_θ$. Furthermore, an isometry between the codes over the rings $\mathcal{R}_θ$ and the analogous DNA codes is established in general. Brief study of dual and self dual codes over the rings is given including the construction of special class of self dual codes that satisfy reverse and reverse-complement constraints. The technical contributions of this paper are twofold. Considering the Generalized Gau distance, Sphere Packing-like bound, GV-like bound, Singleton like bound and Plotkin-like bound are established over the rings $\mathcal{R}_θ$. In addition to this, optimal class of codes are provided with respect to Singleton-like bound and Plotkin-like bound. Moreover, the construction of family of DNA codes is proposed that satisfies reverse and reverse-complement constraints using the Reed-Muller type codes over the rings $\mathcal{R}_θ$.

cs.IT

On Conflict Free DNA Codes

DNA storage has emerged as an important area of research. The reliability of DNA storage system depends on designing the DNA strings (called DNA codes) that are sufficiently dissimilar. In this work, we introduce DNA codes that satisfy a special constraint. Each codeword of the DNA code has a specific property that any two consecutive sub-strings of the DNA codeword will not be the same (a generalization of homo-polymers constraint). This is in addition to the usual constraints such as Hamming, reverse, reverse-complement and $GC$-content. We believe that the new constraint will help further in reducing the errors during reading and writing data into the synthetic DNA strings. We also present a construction (based on a variant of stochastic local search algorithm) to calculate the size of the DNA codes with all the above constraints, which improves the lower bounds from the existing literature, for some specific cases. Moreover, a recursive isometric map between binary vectors and DNA strings is proposed. Using the map and the well known binary codes we obtain few classes of DNA codes with all the constraints including the property that the constructed DNA codewords are free from the hairpin-like secondary structures.

cs.IT

Bounds on Fractional Repetition Codes using Hypergraphs

In the \textit{Distributed Storage Systems} (DSSs), an encoded fraction of information is stored in the distributed fashion on different chunk servers. Recently a new paradigm of \textit{Fractional Repetition} (FR) codes have been introduced, in which, encoded data information is stored on distributed servers, where encoding is done using a \textit{Maximum Distance Separable} (MDS) code and a smart replication of packets. In this work, we have shown that an FR code is equivalent to a hypergraph. Using the correspondence, the properties and the bounds of a hypergraph are directly mapped to the associated FR code. In general, the necessary and sufficient conditions for the existence of an FR code is obtained by using the correspondence. Some of the bounds are new and FR codes meeting these bounds are unknown. It is also shown that any FR code associated with a linear hypergraph is universally good.

cs.IT

On Universally Good Flower Codes

For a Distributed Storage System (DSS), the \textit{Fractional Repetition} (FR) code is a class in which replicas of encoded data packets are stored on distributed chunk servers, where the encoding is done using the Maximum Distance Separable (MDS) code. The FR codes allow for exact uncoded repair with minimum repair bandwidth. In this paper, FR codes are constructed using finite binary sequences. The condition for universally good FR codes is calculated on such sequences. For some sequences, the universally good FR codes are explored.

cs.IT

On Code Rates of Fractional Repetition Codes

In \textit{Distributed Storage Systems} (DSSs), usually, data is stored using replicated packets on different chunk servers. Recently a new paradigm of \textit{Fractional Repetition} (FR) codes have been introduced, in which, data is replicated in a smart way on distributed servers using a \textit{Maximum Distance Separable} (MDS) code. In this work, for a non-uniform FR code, bounds on the FR code rate and DSS code rate are studied. Using matrix representation of an FR code, some universally good FR codes have been obtained.

cs.IT

On Optimal Heterogeneous Regenerating Codes

Heterogeneous Distributed Storage Systems (DSSs) are close to the real world applications for data storage. Each node of the considered DSS, may store different number of packets and each having different repair bandwidth with uniform repair traffic. For such heterogeneous DSS, a failed node can be repaired with the help of some specific nodes. In this work, a family of codes based on graph theory, is constructed which achieves the fundamental bound on file size for the particular heterogeneous DSS.

cs.IT

On Dress Codes with Flowers

Fractional Repetition (FR) codes are well known class of Distributed Replication-based Simple Storage (Dress) codes for the Distributed Storage Systems (DSSs). In such systems, the replicas of data packets encoded by Maximum Distance Separable (MDS) code, are stored on distributed nodes. Most of the available constructions for the FR codes are based on combinatorial designs and Graph theory. In this work, we propose an elegant sequence based approach for the construction of the FR code. In particular, we propose a beautiful class of codes known as Flower codes and study its basic properties.

cs.IT

Tradeoff for Heterogeneous Distributed Storage Systems between Storage and Repair Cost

In this paper, we consider heterogeneous distributed storage systems (DSSs) having flexible reconstruction degree, where each node in the system has dynamic repair bandwidth and dynamic storage capacity. In particular, a data collector can reconstruct the file at time $t$ using some arbitrary nodes in the system and for a node failure the system can be repaired by some set of arbitrary nodes. Using $min$-$cut$ bound, we investigate the fundamental tradeoff between storage and repair cost for our model of heterogeneous DSS. In particular, the problem is formulated as bi-objective optimization linear programing problem. For an arbitrary DSS, it is shown that the calculated $min$-$cut$ bound is tight.

cs.IT

On Heterogeneous Regenerating Codes and Capacity of Distributed Storage Systems

Heterogeneous Distributed Storage Systems (DSS) are close to real world applications for data storage. Internet caching system and peer-to-peer storage clouds are the examples of such DSS. In this work, we calculate the capacity formula for such systems where each node store different number of packets and each having a different repair bandwidth (node can be repaired by contacting a specific set of nodes). The tradeoff curve between storage and repair bandwidth is studied for such heterogeneous DSS. By analyzing the capacity formula new minimum bandwidth regenerating (MBR) and minimum storage regenerating (MBR) points are obtained on the curve. It is shown that in some cases these are better than the homogeneous DSS.

cs.IT

Quorum Sensing for Regenerating Codes in Distributed Storage

Distributed storage systems with replication are well known for storing large amount of data. A large number of replication is done in order to provide reliability. This makes the system expensive. Various methods have been proposed over time to reduce the degree of replication and yet provide same level of reliability. One recently suggested scheme is of Regenerating codes, where a file is divided in to parts which are then processed by a coding mechanism and network coding to provide large number of parts. These are stored at various nodes with more than one part at each node. These codes can generate whole file and can repair a failed node by contacting some out of total existing nodes. This property ensures reliability in case of node failure and uses clever replication. This also optimizes bandwidth usage. In a practical scenario, the original file will be read and updated many times. With every update, we will have to update the data stored at many nodes. Handling multiple requests at the same time will bring a lot of complexity. Reading and writing or multiple writing on the same data at the same time should also be prevented. In this paper, we propose an algorithm that manages and executes all the requests from the users which reduces the update complexity. We also try to keep an adequate amount of availability at the same time. We use a voting based mechanism and form read, write and repair quorums. We have also done probabilistic analysis of regenerating codes.

cs.DC

Reconstruction and Repair Degree of Fractional Repetition Codes

Given a Fractional Repetition code, finding the reconstruction and repair degree in a distributed storage system is an important problem. In this work, we present algorithms for computing the reconstruction and repair degree of fractional repetition codes.

cs.IT