Searcharxiv⌕ Search

arXiv · 2610.09998

Probabilistic GC-Constraints for Composite DNA

Abstract

This paper addresses the challenge of encoding biochemical GC-content constraints in composite DNA-based data storage. Previous deterministic models impose high rate penalties by avoiding any possibility for a strand to fall outside the allowed range. To account for the stochastic nature of composite DNA, we introduce an $ε$-probabilistic constraint framework, and derive capacity bounds for global constraints using composition types and for local sliding-window constraints via finite-state Markov chains. Furthermore, we propose a capacity-achieving multi-type enumerative encoder.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Humeyra Bodur, Frederik Walter, Antonia Wachter-Zeh. 2026-10-07. Probabilistic GC-Constraints for Composite DNA. https://arxiv.org/abs/2610.09998

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

New Quaternary codes with small Plotkin-defects from two-generator simplicial complexes

A recent characterization of all lengths of Plotkin-optimal quaternary (that is, over the ring $\mathbb{Z}_4$) codes of arbitrary type \cite{tang2025plotkin} also pins down the parameters for which no Plotkin-optimal code exists. In this article, we determine the best achievable parameters in several of these cases, obtaining codes whose minimum Lee distance is one less than the Plotkin bound, namely the codes with Plotkin-defect 1. To the best of our knowledge, this is the first attempt to study quaternary codes with Plotkin-defects. Precisely, we construct infinite families of quaternary $\mathcal{C}_{D}$-codes, where the defining set $D$ is derived utilizing a two-generator simplicial complex, and determine their Lee weight distributions. As a result, we find two quaternary linear code families with Plotkin-defect 1 and report at least 23 new or improved parameters having small (upto 4) Plotkin-defects, including 13 projective and 7 optimal parameters. We additionally report 4 quaternary linear codes with best-known parameters that are also projective. Further, their linear Gray images give two infinite families of distance-optimal, one infinite family of at least almost dimension-optimal binary linear codes and five infinite families of minimal binary linear codes.

cs.IT↗

Distributed Hypothesis Testing Against Dependence

We study distributed hypothesis testing and establish the exact error exponent in single-letter form for new testing problems. In distributed hypothesis testing, a receiver decides between $\mathcal{H}_0:P_{XY}$ and $\mathcal{H}_1:Q_{XY}$ based on $Y^n$ and a rate-limited description of $X^n$. So far, such single-letter forms are known only for testing against independence, studied by Ahlswede and Csiszár, and testing against conditional independence, studied by Rahman and Wagner. In this paper, we study testing against dependence, where $P_{XY}=P_XP_Y$, and show that its error exponent is given by Han's exponent, which is established by single-letterizing a multi-letter version of Han's exponent. We also disprove a previous conjecture of Han claiming that the error exponent is given by the lautum information, as we show that Han's exponent can strictly exceed the latter. We then consider the Cartesian product of testing against dependence and testing against independence, and derive a single-letter characterization of its error exponent. We next study testing against conditional dependence, which is the dependence-testing counterpart of the setting studied by Rahman and Wagner. We derive a single-letter converse bound for the error exponent by introducing and solving a related setting where the side information is also available to the transmitter. We show that our converse bound is tight in some cases by using a conditional-coding version of the quantization scheme. Finally, we extend our converse method for testing against dependence to the general setting, and recover a recently established single-letter converse bound.

cs.IT↗

Language Modeling is Monotone Compression

A long-standing hypothesis in artificial intelligence and neuroscience posits that intelligence is closely related to compression: the ability to compress information efficiently intuitively reflects capacities associated with intelligence and learning. Indeed, recent experimental works verify this intuition by showing connections between the capabilities of large language models (LLMs) and their ability as compressors: for instance, Deletang et al. (ICLR'24) demonstrate that LLMs can be used as powerful compressors, and Huang et al. (COLM'24) show that the compression ability of LLMs is highly correlated with their performance on benchmarks for knowledge and reasoning. In this work, we initiate a theoretical study of this connection. Our main result is that LLMs (formally modeled as next-token predictors) are equivalent to monotone (a.k.a. order-preserving) compression algorithms---namely, compression algorithms where the encoding process preserves the ordering of the inputs---in the sense that the one can be constructed from the other while preserving the same error up to an additive gap of 2. We next show that the monotonicity is required for this equivalence to hold if and only if cryptographic (infinitely-often) one-way functions exist. As a direct corollary, we get a cryptographic result of independent interest: the notion of next-bit pseudoentropy (a computational analogue of entropy) of a distribution is equivalent to monotone incompressibility of the distribution. (Previously, it was only known (Haitner et al., ITCS'23) that incompressibility implies next-bit pseudoentropy.)

cs.IT↗