arXiv · 2607.06827
Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs
Abstract
Speech large language models (Speech LLMs) typically encode speech into sequences far longer than text, creating a major efficiency bottleneck during autoregressive decoding. A common remedy is to compress the speech sequence at the adapter level to remove temporal redundancy before it enters the LLM; however, such early downsampling risks discarding fine-grained information that cannot be recovered. We propose SpeechKV, which applies a learned pooling to the KV cache of speech tokens inside the LLM. This design allows the LLM to fuse speech and text internally while directly accelerating decoding. Trained on 71K hours of speech data, SpeechKV compresses the speech to approximately text-level granularity yet maintains performance on par with or even slightly better than the uncompressed baseline, with relative gains of 6.6% on out-of-domain entity recognition and 2.3% on OpenASR, while delivering at least 1.49 times decoding speedup that scales with audio length.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ke-Han Lu, Keqi Deng, Ruchao Fan, Rui Zhao, Jinyu Li. 2026-07-07. Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs. https://arxiv.org/abs/2607.06827
Cite the original work for its findings. Save a collection to share your selection of sources.