Searcharxiv⌕ Search

arXiv · 2610.11031

Language Modeling is Monotone Compression

Abstract

A long-standing hypothesis in artificial intelligence and neuroscience posits that intelligence is closely related to compression: the ability to compress information efficiently intuitively reflects capacities associated with intelligence and learning. Indeed, recent experimental works verify this intuition by showing connections between the capabilities of large language models (LLMs) and their ability as compressors: for instance, Deletang et al. (ICLR'24) demonstrate that LLMs can be used as powerful compressors, and Huang et al. (COLM'24) show that the compression ability of LLMs is highly correlated with their performance on benchmarks for knowledge and reasoning. In this work, we initiate a theoretical study of this connection. Our main result is that LLMs (formally modeled as next-token predictors) are equivalent to monotone (a.k.a. order-preserving) compression algorithms---namely, compression algorithms where the encoding process preserves the ordering of the inputs---in the sense that the one can be constructed from the other while preserving the same error up to an additive gap of 2. We next show that the monotonicity is required for this equivalence to hold if and only if cryptographic (infinitely-often) one-way functions exist. As a direct corollary, we get a cryptographic result of independent interest: the notion of next-bit pseudoentropy (a computational analogue of entropy) of a distribution is equivalent to monotone incompressibility of the distribution. (Previously, it was only known (Haitner et al., ITCS'23) that incompressibility implies next-bit pseudoentropy.)

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Noam Mazor, Andrew Morgan, Rafael Pass. 2026-10-08. Language Modeling is Monotone Compression. https://arxiv.org/abs/2610.11031

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A $q$-Polymatroid Framework for Information Leakage in Secure Linear Network Coding

We study information leakage in secure linear network coding schemes based on nested rank-metric codes. We show that the amount of information leaked to an adversary that observes a subset of network links is characterized by the conditional rank function of a representable $q$-polymatroid associated with the underlying rank-metric code pair. Building on this connection, we introduce the notions of $q$-polymatroid ports and $q$-access structures and describe their structural properties. Moreover, we relate minimal codewords in the rank-metric setting to minimal reconstructing spaces and prove a $q$-analogue of the Brickell--Davenport theorem.

cs.IT↗

New Quaternary codes with small Plotkin-defects from two-generator simplicial complexes

A recent characterization of all lengths of Plotkin-optimal quaternary (that is, over the ring $\mathbb{Z}_4$) codes of arbitrary type \cite{tang2025plotkin} also pins down the parameters for which no Plotkin-optimal code exists. In this article, we determine the best achievable parameters in several of these cases, obtaining codes whose minimum Lee distance is one less than the Plotkin bound, namely the codes with Plotkin-defect 1. To the best of our knowledge, this is the first attempt to study quaternary codes with Plotkin-defects. Precisely, we construct infinite families of quaternary $\mathcal{C}_{D}$-codes, where the defining set $D$ is derived utilizing a two-generator simplicial complex, and determine their Lee weight distributions. As a result, we find two quaternary linear code families with Plotkin-defect 1 and report at least 23 new or improved parameters having small (upto 4) Plotkin-defects, including 13 projective and 7 optimal parameters. We additionally report 4 quaternary linear codes with best-known parameters that are also projective. Further, their linear Gray images give two infinite families of distance-optimal, one infinite family of at least almost dimension-optimal binary linear codes and five infinite families of minimal binary linear codes.

cs.IT↗

Generalized Rank Weight and Extended Generalized Poset Weight Defined For Codes Over Rings: A Galois Connection Approach

In this paper, we study generalized rank weights (GRWs) and extended generalized poset weight (EGPWs) of codes over rings via a Galois connection approach. First, we show that various coding-theoretic properties related to generalized weights, including security drops of a code employed in wire-tap channel of type II, connections between generalized weights of a Gabidulin code and its associated Delsarte code, (generalized) Singleton bound, MDS discrepancy of a code, characterizations of MDS, near MDS, $i$-MDS, MRD, near MRD, $i$-MRD, (dually) quasi-MRD codes as well as evasive property of subspaces, can be reformulated in terms of Galois connections. Next, we study GRWs and rank profiles defined for modules over principal ideal rings, especially those over chain rings. Generalizing GRWs defined for vector spaces over fields, we establish a singleton bound and a Wei-type duality theorem, characterize MRD, near MRD and dually quasi-MRD codes and determine their GRWs; moreover, we characterize $i$-MRD codes and establish a scattered bound for $(h,h)$-evasive codes over chain rings, generalizing counterpart result established for vector space over finite fields. Finally, we propose and study EGPWs and extended poset profiles defined for modules with a composition series, which in fact form a Galois connection. Generalizing EGPWs defined for modules over finite Galois rings, we establish a Wei-type duality theorem for modules over arbitrary quasi-Frobenius rings, which unifies the two Wei-type duality theorems derived in both \cite{32} and \cite{33}.

cs.IT↗