arXiv · 2606.10671
FadeMem: Distance-Aware Memory Consolidation for Autoregressive Video Diffusion
Abstract
Autoregressive video generators synthesize long videos by generating successive temporal segments, but their historical KV cache grows with video length. Existing bounded-cache methods reduce this cost with local windows, sink tokens, or compressed memory states, yet they usually assign fixed roles to different parts of the history. We propose FadeMem, a distance-aware KV memory consolidation mechanism that organizes historical KV blocks into a temporal hierarchy under a fixed cache budget. This design is motivated by frequency-dependent temporal decay: fine details decorrelate quickly, while coarse scene structure and identity remain useful over longer horizons. During generation, new history is inserted as fine-grained entries, while older adjacent entries are progressively merged under a power-law temporal allocation schedule, yielding a dense-near, sparse-far memory within one cache. Without architectural changes, FadeMem improves long-range consistency while largely preserving visual quality, and lightweight adaptation further enhances motion dynamics and visual fidelity. Using the same unified schedule under a fixed cache budget, FadeMem also remains effective over multi-minute and hour long video generation and reduces peak memory under a matched KV budget.
Explore related subjects
Keep this discovery
Yu Lu, Junjie Yang, Piotr Koniusz, YuXin Song, Yi Yang. 2026-06-09. FadeMem: Distance-Aware Memory Consolidation for Autoregressive Video Diffusion. https://arxiv.org/abs/2606.10671
Cite the original work for its findings. Save a collection to share your selection of sources.