SearcharxivSearch

arXiv subjects

Ryan Gabrys

Publications and source records attributed to Ryan Gabrys.

At least 19 recordsLinked to original sources

Error characterization and error correction approaches in combinatorial DNA-based storage

Data storage in DNA has recently emerged as a promising archival solution, offering space-efficient and long-lasting digital storage. Combinatorial DNA encoding enhances this potential by increasing the logical density through combinations of DNA shortmers, where each sequence position is represented by a set of predefined short DNA fragments, allowing more data to be encoded using fewer synthesis cycles. However, this method introduces unique synthesis and sequencing errors. In this study, we characterize errors in combinatorial DNA-based storage systems. We reveal that asymmetric combinatorial erasure errors, defined as the omission of a single shortmer from the set defining the combinatorial letter, are a prevalent error type, particularly in large-scale systems where read coverage is limited. In two previously published datasets, we observed a high frequency of erasure errors, where missing sequences obstruct the reconstruction of combinatorial letters. We conducted a large-scale experimental proof-of-concept and confirmed that erasure errors become increasingly prominent with reduced sequencing depth: below 50 reads per sequence, their frequency sharply increased. We developed an asymmetric error-correcting code for these errors, utilizing tensor-product codes to integrate standard erasure and substitution-correcting codes (such as Reed-Solomon (RS) codes) with asymmetric Varshamov-Tenengolts codes. We validated its performance in simulations and in a second large-scale experiment directly comparing it with the more straightforward 2D RS scheme. Our method consistently outperformed 2D RS, particularly in erasure-dominated scenarios, and demonstrated superior decoding accuracy under low coverage conditions where 2D RS struggled to decode the data. Our findings demonstrate the importance of error correction schemes tailored to the asymmetric nature of errors in combinatorial DNA.

cs.IT

More on Codes for Combinatorial Composite DNA

In this paper, we focus on constructions of unique-decodable/list-decodable on the recently studied $(t,e)$-composite-asymmetric error-correcting codes ($(t,e)$-CAECCs). Let $X$ be an $m\times n$ binary matrix, in which each row has Hamming weight $w$. When at most $t$ rows of $X$ suffer from errors and in each of these erroneous rows, there are at most $e$ $1 \to 0$ errors, we say that a $(t,e)$-composite-asymmetric-error occurs in $X$. For general $m,n,w,t,e$, we propose new constructions of $(t,e)$-CAECCs with redundancy at most $(t-1)\log(m)+O(1)$, where $O(1)$ is a number independent of the code-length $m$. In particular, this gives a class of $(2,e)$-CAECCs that are optimal in terms of their redundancy. %(in terms of redundancy, regarded as a function of the number of rows $m$) s. When $m$ is a prime power, the redundancy can be further reduced to $(t-1)\log(m)-O(\log(m))$. To further increase the size of these codes, we introduce a combinatorial object called a weak $B_e$-sets. When $e=w$, we show an efficient way to encode/decode our codes. At last, we investigate how much we can gain if we relax the requirement of uniquely decoding to list-decoding. It is shown that when the list size is $t!$ or an exponential function of $t$, there are list-decodable $(t,e)$-CAECCs with constant redundancy. When the list size is two, we show that there are list-decodable $(3,2)$-CAECCs with redundancy $\log(m)+O(1)$.

cs.IT

An Additive Approximation Scheme for Generating Dyadic Codings for the Outputs of an LLM

We study the problem of approximating a discrete probability distribution, such as the next-token distribution of a large language model, by a dyadic distribution induced by a binary tree under encoding rate constraints. The objective is to partition the support of the distribution and assign dyadic probabilities to minimize total variation distance while achieving a prescribed rate. We formulate this task as a tree-based partitioning problem and develop a polynomial-time additive approximation scheme for the rate-constrained setting in the constant-rate regime. Our results provide provable guarantees for near-optimal dyadic approximations and, as an application, yield a principled framework for LLM-based steganography, where the rate maps to bits of hidden information embedded per token and the total variation bound controls statistical detectability.

cs.IT

Efficient Synthesis for Two-Dimensional Strand Arrays with Row Constraints

In large-scale array-based DNA synthesis, optical and chemical coupling between nearby sites can limit simultaneous activations. Motivated by this constraint, we study strands synthesized according to a fixed global synthesis sequence, with at most one strand per row advancing in each cycle. We focus on the fundamental case of two strands in a single row and analyze the expected completion time of row-constrained synthesis. We introduce the laggard-first (LF) policy, a simple rule that always advances the strand with fewer synthesized symbols when a conflict arises, and establish that it is asymptotically optimal among online policies without look-ahead. In the binary case, one-symbol look-ahead strictly improves on the no-look-ahead bound. We further show that even complete advance knowledge does not eliminate the scheduling loss, as even a globally optimal schedule incurs an unavoidable expected overhead that grows linearly with the strand length. Finally, we complement these scheduling results with a dynamic programming algorithm for computing an optimal offline synthesis order and a constant-redundancy binary coding scheme that yields a deterministic worst-case synthesis time guarantee.

cs.IT

Improved Interactive Protocol for Synchronizing From Deletions

Data synchronization is a fundamental problem with applications in diverse fields such as cloud storage, genomics, and distributed systems. This paper addresses the challenge of synchronizing two files, one of which is a subsequence of the other and related through a constant rate of deletions, using an improved communication protocol. Building upon prior work, we integrate advanced multi-deletion correction codes into an existing baseline protocol, which previously relied on single-deletion correction. Our proposed protocol reduces communication cost by leveraging more general partitioning techniques as well as multi-deletion error correction. We derive a generalized upper bound on the expected number of transmitted bits, applicable to a broad class of deletion correction codes. Experimental results demonstrate that our approach outperforms the baseline in communication cost. These findings establish the efficacy of the improved protocol in achieving low-redundancy synchronization in scenarios where deletion errors occur.

cs.IT

Constructing Low-Redundancy Codes via Distributed Graph Coloring

We present a general framework for constructing error-correcting codes using distributed graph coloring under the LOCAL model. Building on the correspondence between independent sets in the confusion graph and valid codes, we show that the color of a single vertex - consistent with a global proper coloring - can be computed in polynomial time using a modified version of Linial's coloring algorithm, leading to efficient encoding and decoding. Our results include: i) uniquely decodable code constructions for a constant number of errors of any type with redundancy twice the Gilbert-Varshamov bound; ii) list-decodable codes via a proposed extension of graph coloring, namely, hypergraph labeling; iii) an incremental synchronization scheme with reduced average-case communication when the edit distance is not precisely known; and iv) the first asymptotically optimal codes (up to a factor of 8) for correcting bursts of unbounded-length edits. Compared to syndrome compression, our approach is more flexible and generalizable, does not rely on a good base code, and achieves improved redundancy across a range of parameters.

cs.IT

Complex DNA Synthesis Sequences

DNA-based storage offers unprecedented density and durability, but its scalability is fundamentally limited by the efficiency of parallel strand synthesis. Existing methods either allow unconstrained nucleotide additions to individual strands, such as enzymatic synthesis, or enforce identical additions across many strands, such as photolithographic synthesis. We introduce and analyze a hybrid synthesis framework that generalizes both approaches: in each cycle, a nucleotide is selected from a restricted subset and incorporated in parallel. This model gives rise to a new notion of a complex synthesis sequence. Building on this framework, we extend the information rate definition of Lenz et al. and analyze an analog of the deletion ball, defined and studied in this setting, deriving tight expressions for the maximal information rate and its asymptotic behavior. These results bridge the theoretical gap between constrained models and the idealized setting in which every nucleotide is always available. For the case of known strands, we design a dynamic programming algorithm that computes an optimal complex synthesis sequence, highlighting structural similarities to the shortest common supersequence problem. We also define a distinct two-dimensional array model with synthesis constraints over the rows, which extends previous synthesis models in the literature and captures new structural limitations in large-scale strand arrays. Additionally, we develop a dynamic programming algorithm for this problem as well. Our results establish a new and comprehensive theoretical framework for constrained DNA, subsuming prior models and setting the stage for future advances in the field.

cs.IT

The Labeled Coupon Collector Problem

We generalize the well-known Coupon Collector Problem (CCP) in combinatorics. Our problem is to find the minimum and expected number of draws, with replacement, required to recover $n$ distinctly labeled coupons, with each draw consisting of a random subset of $k$ different coupons and a random ordering of their associated labels. We specify two variations of the problem, Type-I in which the set of labels is known at the start, and Type-II in which the set of labels is unknown at the start. We show that our problem can be viewed as an extension of the separating system problem introduced by R\'enyi and Katona, provide a full characterization of the minimum, and provide a numerical approach to finding the expectation using a Markov chain model, with special attention given to the case where two coupons are drawn at a time.

cs.DM

Asymptotically Optimal Codes Correcting One Substring Edit

The substring edit error is the operation of replacing a substring $u$ of $x$ with another string $v$, where the lengths of $u$ and $v$ are bounded by a given constant $k$. It encompasses localized insertions, deletions, and substitutions within a window. Codes correcting one substring edit have redundancy at least $\log n+k$. In this paper, we construct codes correcting one substring edit with redundancy $\log n+O(\log \log n)$, which is asymptotically optimal.

cs.IT

Making it to First: The Random Access Problem in DNA Storage

In this paper, we study the Random Access Problem in DNA storage, which addresses the challenge of retrieving a specific information strand from a DNA-based storage system. In this framework, the data is represented by $k$ information strands which represent the data and are encoded into $n$ strands using a linear code. Then, each sequencing read returns one encoded strand which is chosen uniformly at random. The goal under this paradigm is to design codes that minimize the expected number of reads required to recover an arbitrary information strand. We fully solve the case when $k=2$, showing that the best possible code attains a random access expectation of $1+\frac{2}{\sqrt{2}+1}\approx 0.914\cdot 2$ for $q$ large enough. Moreover, we generalize a construction from~\cite{GMZ24}, specifically to $k=3$, for any value of $k$. Our construction uses $B_{k-1}$ sequences over $\mathbb{Z}_{q-1}$, that always exist over large finite fields. We show that for every $k\geq 4$, this generalized construction outperforms all previous constructions in terms of reducing the random access expectation.

cs.IT

HoneyGAN Pots: A Deep Learning Approach for Generating Honeypots

This paper investigates the feasibility and effectiveness of employing Generative Adversarial Networks (GANs) for the generation of decoy configurations in the field of cyber defense. The utilization of honeypots has been extensively studied in the past; however, selecting appropriate decoy configurations for a given cyber scenario (and subsequently retrieving/generating them) remain open challenges. Existing approaches often rely on maintaining lists of configurations or storing collections of pre-configured images, lacking adaptability and efficiency. In this pioneering study, we present a novel approach that leverages GANs' learning capabilities to tackle these challenges. To the best of our knowledge, no prior attempts have been made to utilize GANs specifically for generating decoy configurations. Our research aims to address this gap and provide cyber defenders with a powerful tool to bolster their network defenses.

cs.CR

Robust Gray Codes Approaching the Optimal Rate

Robust Gray codes were introduced by (Lolck and Pagh, SODA 2024). Informally, a robust Gray code is a (binary) Gray code $\mathcal{G}$ so that, given a noisy version of the encoding $\mathcal{G}(j)$ of an integer $j$, one can recover $\hat{j}$ that is close to $j$ (with high probability over the noise). Such codes have found applications in differential privacy. In this work, we present near-optimal constructions of robust Gray codes. In more detail, we construct a Gray code $\mathcal{G}$ of rate $1 - H_2(p) - \varepsilon$ that is efficiently encodable, and that is robust in the following sense. Supposed that $\mathcal{G}(j)$ is passed through the binary symmetric channel $\text{BSC}_p$ with cross-over probability $p$, to obtain $x$. We present an efficient decoding algorithm that, given $x$, returns an estimate $\hat{j}$ so that $|j - \hat{j}|$ is small with high probability.

cs.IT

Quickly-Decodable Group Testing with Fewer Tests: Price-Scarlett and Cheraghchi-Nakos's Nonadaptive Splitting with Explicit Scalars

We modify Cheraghchi-Nakos [CN20] and Price-Scarlett's [PS20] fast binary splitting approach to nonadaptive group testing. We show that, to identify a uniformly random subset of $k$ infected persons among a population of $n$, it takes only $\ln(2 - 4\varepsilon) ^{-2} k \ln n$ tests and decoding complexity $O(\varepsilon^{-2} k \ln n)$, for any small $\varepsilon > 0$, with vanishing error probability. In works prior to ours, only two types of group testing schemes exist. Those that use $\ln(2)^{-2} k \ln n$ or fewer tests require linear-in-$n$ complexity, sometimes even polynomial in $n$; those that enjoy sub-$n$ complexity employ $O(k \ln n)$ tests, where the big-$O$ scalar is implicit, presumably greater than $\ln(2)^{-2}$. We almost achieve the best of both worlds, namely, the almost-$\ln(2)^{-2}$ scalar and the sub-$n$ decoding complexity. How much further one can reduce the scalar $\ln(2)^{-2}$ remains an open problem.

cs.IT

One Code Fits All: Strong stuck-at codes for versatile memory encoding

In this work we consider a generalization of the well-studied problem of coding for ``stuck-at'' errors, which we refer to as ``strong stuck-at'' codes. In the traditional framework of stuck-at codes, the task involves encoding a message into a one-dimensional binary vector. However, a certain number of the bits in this vector are 'frozen', meaning they are fixed at a predetermined value and cannot be altered by the encoder. The decoder, aware of the proportion of frozen bits but not their specific positions, is responsible for deciphering the intended message. We consider a more challenging version of this problem where the decoder does not know also the fraction of frozen bits. We construct explicit and efficient encoding and decoding algorithms that get arbitrarily close to capacity in this scenario. Furthermore, to the best of our knowledge, our construction is the first, fully explicit construction of stuck-at codes that approach capacity.

cs.IT

Tail-Erasure-Correcting Codes

The increasing demand for data storage has prompted the exploration of new techniques, with molecular data storage being a promising alternative. In this work, we develop coding schemes for a new storage paradigm that can be represented as a collection of two-dimensional arrays. Motivated by error patterns observed in recent prototype architectures, our study focuses on correcting erasures in the last few symbols of each row, and also correcting arbitrary deletions across rows. We present code constructions and explicit encoders and decoders that are shown to be nearly optimal in many scenarios. We show that the new coding schemes are capable of effectively mitigating these errors, making these emerging storage platforms potentially promising solutions.

cs.IT

Error-Correcting Codes for Combinatorial Composite DNA

Data storage in DNA is developing as a possible solution for archival digital data. Recently, to further increase the potential capacity of DNA-based data storage systems, the combinatorial composite DNA synthesis method was suggested. This approach extends the DNA alphabet by harnessing short DNA fragment reagents, known as shortmers. The shortmers are building blocks of the alphabet symbols, consisting of a fixed number of shortmers. Thus, when information is read, it is possible that one of the shortmers that forms part of the composition of a symbol is missing and therefore the symbol cannot be determined. In this paper, we model this type of error as a type of asymmetric error and propose code constructions that can correct such errors in this setup. We also provide a lower bound on the redundancy of such error-correcting codes and give an explicit encoder and decoder pair for our construction. Our suggested error model is also supported by an analysis of data from actual experiments that produced DNA according to the combinatorial scheme. Lastly, we also provide a statistical evaluation of the probability of observing such error events, as a function of read depth.

cs.IT

Storage codes and recoverable systems on lines and grids

A storage code is an assignment of symbols to the vertices of a connected graph $G(V,E)$ with the property that the value of each vertex is a function of the values of its neighbors, or more generally, of a certain neighborhood of the vertex in $G$. In this work we introduce a new construction method of storage codes, enabling one to construct new codes from known ones via an interleaving procedure driven by resolvable designs. We also study storage codes on $\mathbb Z$ and ${\mathbb Z}^2$ (lines and grids), finding closed-form expressions for the capacity of several one and two-dimensional systems depending on their recovery set, using connections between storage codes, graphs, anticodes, and difference-avoiding sets.

cs.IT