arXiv · 2609.37852
Delta-Matching: Closing the Final Gap of Native 8-bit Training for LLMs
Abstract
Reliable FP8 attention remains a barrier to fully native 8-bit large language model training. We derive how forward-backward inconsistencies produce stale delta and empirically show how it distorts training dynamics. Our stale-delta hybrid runs show a modest loss gap at 569M parameters but substantial loss increases and downstream degradation at 1.67B and 5.29B. QK normalization, NoPE (no positional encoding), and lower-learning-rate context extension mitigate or delay degradation without eliminating it. This pattern suggests accumulated optimization error that smaller models and short runs can conceal. We propose Delta-Matching, proving that it restores the softmax gradient's zero-row-sum invariant under the stated numerical assumptions. It enables native block-scaled FP8 in every forward and backward attention-core matmul without architectural changes, smaller global batches, or auxiliary forward outputs. Across tested architectures, scales, and training stages, Delta-Matching matches BF16/FP32 mixed-precision training loss and overall downstream performance. We will release our implementation, trained models, and data recipes.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Haozhan Tang, Hao Kang, Han Cai, Song Han, Chenyan Xiong. 2026-09-29. Delta-Matching: Closing the Final Gap of Native 8-bit Training for LLMs. https://arxiv.org/abs/2609.37852
Cite the original work for its findings. Save a collection to share your selection of sources.