arXiv · 2602.11951
Learning-Enhanced Composite DNA Data Storage Under Sampling Randomness and IDS Errors
Abstract
DNA data storage offers a high-density, long-term alternative to conventional storage systems, addressing the exponential growth of digital data. Composite DNA extends this paradigm by leveraging mixtures of nucleotides to increase storage capacity beyond the four standard bases. In this work, composite DNA storage is modeled as a multinomial channel, and an analogy to digital modulation is established by representing composite letters on the three-dimensional probability simplex. To mitigate errors caused by sampling randomness, we derive transition probabilities for each constellation point which enables the computation of bit-wise log-likelihood ratios (LLRs) to employ practical channel codes for error correction. The framework is then extended to substitution and insertion-deletion-substitution (IDS) channels by proposing constellation update rules that account for these impairments. For the substitution channel, exact constellation updates enable exact LLR computation. In contrast, for the IDS channel, the large number of possible cases necessitates an approximate update rule, yielding approximate LLRs. To address this limitation, we further propose two learning-enhanced receiver architectures: (1) RhoNet, a learning-assisted model-based method that refines constellation points prior to analytical LLR computation, and (2) LLRNet, which directly estimates LLRs, representing a progressive transition from model-based analytical estimation to fully data-driven processing. Numerical results demonstrate reliable performance with existing low-density parity-check (LDPC) codes.
Explore related subjects
Keep this discovery
Busra Tegin, Tolga M Duman. 2026-02-12. Learning-Enhanced Composite DNA Data Storage Under Sampling Randomness and IDS Errors. https://arxiv.org/abs/2602.11951
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.