arXiv · 2609.34814
WaveAlign: Cache-Aware Query-Row Scheduling for Sparse Attention in Long-Video Generation
Abstract
Long-video generation with diffusion transformers (DiTs) produces extremely long token sequences, making attention a dominant inference bottleneck. Dynamic sparse attention reduces computation, but its realized speedup remains limited because irregular query-row execution degrades L2 cache locality and increases HBM traffic. We present WaveAlign, a lightweight, cache-aware query-row reordering framework for dynamic sparse attention. WaveAlign formulates row ordering as an optimization problem and approximates it with two stages. The first stage derives a low-rank SVD representation of sparse-mask rows and groups query rows with similar K/V access patterns, increasing K/V overlap among concurrently scheduled rows. The second stage exploits streaming GPU scheduling by sorting rows within each wave in descending order of their K/V-block counts, so that short rows from the current wave are followed by long rows from the next. This aligns K/V accesses across wave boundaries and enables shared blocks to be reused before eviction. An adaptive skip module avoids unprofitable reordering. By only permuting query and mask rows, WaveAlign preserves sparse-attention semantics and requires no changes to existing methods or backend kernels. Across two GPU architectures, two video DiTs, and four sparse-attention methods, WaveAlign raises the L2 cache hit ratio from 28.48%--36.35% to 79.38%--89.06%, reduces HBM read traffic by up to 92.11%, and achieves up to 1.25x kernel and 1.17x end-to-end generation speedup without quality loss.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zijian Dai, Sen Han, Youhui Bai, Shannon Wang, Kan Wu, Jingkai Huang, Yuhang Wang, Jing Li, Cheng Li. 2026-09-28. WaveAlign: Cache-Aware Query-Row Scheduling for Sparse Attention in Long-Video Generation. https://arxiv.org/abs/2609.34814
Cite the original work for its findings. Save a collection to share your selection of sources.