You Only Convert Twice: Ozaki Scheme II for Chained Tensor Mode Products
This work proposes OzII-RescaleBE, a method that enables persistent execution of chained tensor mode products in Ozaki scheme II, keeping the intermediate tensors in residue form. By performing rescaling and base extension directly on the residue representation, it requires conversion to residues only at the beginning of the chain and reconstruction only at the end. An implementation using CuTe DSL and INT8 Tensor Cores is provided for verification and evaluation. The method is compared with per-mode composition of Ozaki scheme~II (GEMMul8-Composed) and a dense Kronecker-product formulation (GEMMul8-KRON), both implemented using the GEMMul8 library. Across three datasets with different input distributions, OzII-RescaleBE achieves normwise relative errors ranging from $2.5\times10^{-16}$ to $2.6\times10^{-15}$ for chain depths $d=3$ through $7$. The corresponding errors range from $10^{-16}$ to $10^{-15}$ for GEMMul8-KRON and from $1\times10^{-16}$ to $3\times10^{-16}$ for GEMMul8-Composed and the FP64 chain. At $d=8$, the errors of OzII-RescaleBE increase to between $7\times10^{-14}$ and $6\times10^{-13}$. For mode sizes $n=8$ to $256$, OzII-RescaleBE achieves throughputs ranging from 1.47 to 4.42\,GDoF/s. Its throughput is within 12\,\% of GEMMul8-KRON at $n=8$ and $12$ and exceeds it at larger tested sizes. Compared with GEMMul8-Composed, OzII-RescaleBE achieves $2.1$--$60\times$ the throughput for $n\leq128$, but has 6--12\,\% lower throughput at $n=256$. It also uses less device memory than the GEMMul8 baselines in the measured memory comparisons. These results demonstrate that rescaling and base extension enable accurate residue-domain tensor chains while avoiding repeated intermediate conversions to floating point.