arXiv · 2601.16032
Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10
Abstract
High-performance attention kernels are essential for Large Language Models. This paper presents analysis of CuTile-based Flash Attention memory behavior and a technique to improve its cache performance. In particular, our analysis on the NVIDIA GB10 (Grace Blackwell) identifies the main cause of L2 cache miss. Leveraging this insight, we introduce a new programming technique called Sawtooth Wavefront Reordering that reduces L2 misses. We validate it in both CUDA and CuTile, observing 50\% or greater reduction in L2 misses and up to 60\% increase in throughput on GB10.
Explore related subjects
Keep this discovery
Yifan Zhu, Yekai Pan, Chen Ding. 2026-01-22. Sawtooth Wavefront Reordering: Enhanced CuTile FlashAttention on NVIDIA GB10. https://arxiv.org/abs/2601.16032
Cite the original work for its findings. Save a collection to share your selection of sources.