Searcharxiv⌕ Search

arXiv subjects

Xuanzhengbo Ren

Publications and source records attributed to Xuanzhengbo Ren.

2 recordsLinked to original sources

You Only Convert Twice: Ozaki Scheme II for Chained Tensor Mode Products

This work proposes OzII-RescaleBE, a method that enables persistent execution of chained tensor mode products in Ozaki scheme II, keeping the intermediate tensors in residue form. By performing rescaling and base extension directly on the residue representation, it requires conversion to residues only at the beginning of the chain and reconstruction only at the end. An implementation using CuTe DSL and INT8 Tensor Cores is provided for verification and evaluation. The method is compared with per-mode composition of Ozaki scheme~II (GEMMul8-Composed) and a dense Kronecker-product formulation (GEMMul8-KRON), both implemented using the GEMMul8 library. Across three datasets with different input distributions, OzII-RescaleBE achieves normwise relative errors ranging from $2.5\times10^{-16}$ to $2.6\times10^{-15}$ for chain depths $d=3$ through $7$. The corresponding errors range from $10^{-16}$ to $10^{-15}$ for GEMMul8-KRON and from $1\times10^{-16}$ to $3\times10^{-16}$ for GEMMul8-Composed and the FP64 chain. At $d=8$, the errors of OzII-RescaleBE increase to between $7\times10^{-14}$ and $6\times10^{-13}$. For mode sizes $n=8$ to $256$, OzII-RescaleBE achieves throughputs ranging from 1.47 to 4.42\,GDoF/s. Its throughput is within 12\,\% of GEMMul8-KRON at $n=8$ and $12$ and exceeds it at larger tested sizes. Compared with GEMMul8-Composed, OzII-RescaleBE achieves $2.1$--$60\times$ the throughput for $n\leq128$, but has 6--12\,\% lower throughput at $n=256$. It also uses less device memory than the GEMMul8 baselines in the measured memory comparisons. These results demonstrate that rescaling and base extension enable accurate residue-domain tensor chains while avoiding repeated intermediate conversions to floating point.

cs.DC↗

Learning-Augmented Performance Model for Tensor Product Factorization in High-Order FEM

Accurate performance prediction is essential for optimizing scientific applications on modern high-performance computing (HPC) architectures. Widely used performance models primarily focus on cache and memory bandwidth, which is suitable for many memory-bound workloads. However, it is unsuitable for highly arithmetic intensive cases such as the sum-factorization with tensor $n$-mode product kernels, which are an optimization technique for high-order finite element methods (FEM). On processors with relatively high single instruction multiple data (SIMD) instruction latency, such as the Fujitsu A64FX, the performance of these kernels is strongly influenced by loop-body splitting strategies. Memory-bandwidth-oriented models are therefore not appropriate for evaluating these splitting configurations, and a model that directly reflects instruction-level efficiency is required. To address this need, we develop a dependency-chain-based analytical formulation that links loop-splitting configurations to instruction dependencies in the tensor $n$-mode product kernel. We further use XGBoost to estimate key parameters in the analytical model that are difficult to model explicitly. Evaluations show that the learning-augmented model outperforms the widely used standard Roofline and Execution-Cache-Memory (ECM) models. On the Fujitsu A64FX processor, the learning-augmented model achieves mean absolute percentage errors (MAPE) between 1% and 24% for polynomial orders ($P$) from 1 to 15. In comparison, the standard Roofline and ECM models yield errors of 42%-256% and 5%-117%, respectively. On the Intel Xeon Gold 6230 processor, the learning-augmented model achieves MAPE values from 1% to 13% for $P$=1 to $P$=14, and 24% at $P$=15. In contrast, the standard Roofline and ECM models produce errors of 1%-73% and 8%-112% for $P$=1 to $P$=15, respectively.

cs.DC↗