arXiv · 2608.21541
Beyond Sparse Weights: When Is Attention Compressible?
Abstract
KV-cache compression is often justified by attention maps with a few large weights. This is incomplete: large weights may not contain most of the mass, omitted values can cancel, and preserving the attention output may not preserve the task. We separate these questions. Global score gaps -- not threshold counts -- determine how many tokens are needed to retain a target mass. For a realized row, the weighted sum of omitted values is the exact missing statistic. A controlled retrieval--aggregation model explains when truncation helps and when it hurts. These results motivate CertKV, a training-free compressor that reserves one tail-summary slot per head and allocates the rest by value dispersion. Under matched budgets, CertKV is top-two in seven of nine LongBench-v2 settings, remains in the leading compressed tier on 128K RULER, and realizes a ten-fold cache budget in a packed Llama prototype. Compressibility depends on the mass, values, future queries, and task -- not on a sparse-looking map alone.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Chiwun Yang, Xiaoyu Li. 2026-08-21. Beyond Sparse Weights: When Is Attention Compressible?. https://arxiv.org/abs/2608.21541
Cite the original work for its findings. Save a collection to share your selection of sources.