arXiv · 2608.06812
DGEMM with Ozaki Scheme I/II on FP4 Tensor Cores: A Base-13 E2M1 Limb Representation
Abstract
This paper proposes a method and its implementation for emulating FP64 matrix multiplication (DGEMM) by constructing, on FP4 (E2M1; 2 exponent bits and 1 mantissa bit) Tensor Cores, Ozaki schemes I and II, which realize high-precision matrix multiplication on low-precision arithmetic units. Prior implementations were based on INT8 and FP8, and the use of the faster FP4 had not been realized. The key property is that every FP4 value becomes an integer when doubled, and that shifting this integer set by multiples of 13 covers all integers. Converting an arbitrary integer into base-13 FP4 limbs by this property keeps intermediate sums error-free in FP32 accumulators, which makes FP4 Tensor Cores usable for Ozaki schemes I and II. By the same principle, the integer GEMM of INT8 Tensor Cores can also be emulated bit-exactly on FP4 Tensor Cores. When FP4 Tensor Cores have twice the throughput of FP8, Ozaki scheme II on FP4 theoretically achieves slightly higher performance than its FP8 counterpart. This paper further proposes kernel implementation optimizations raising the attained fraction of peak performance, obtaining a measured speedup on top of the theoretical advantage. We verify this on an RTX PRO 6000 Blackwell, achieving performance competitive with that of an existing FP8-based implementation of Ozaki scheme II, and actually exceeding it at a large problem size ($16384^3$).
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Shun-ichiro Hayashi, Daichi Mukunoki, Tetsuya Hoshino, Takahiro Katagiri. 2026-08-07. DGEMM with Ozaki Scheme I/II on FP4 Tensor Cores: A Base-13 E2M1 Limb Representation. https://arxiv.org/abs/2608.06812
Cite the original work for its findings. Save a collection to share your selection of sources.