arXiv · 2601.09114
A Machine Learning Approach Towards Runtime Optimisation of Matrix Multiplication
Abstract
The GEneral Matrix Multiplication (GEMM) is one of the essential algorithms in scientific computing. Single-thread GEMM implementations are well-optimised with techniques like blocking and autotuning. However, due to the complexity of modern multi-core shared memory systems, it is challenging to determine the number of threads that minimises the multi-thread GEMM runtime. We present a proof-of-concept approach to building an Architecture and Data-Structure Aware Linear Algebra (ADSALA) software library that uses machine learning to optimise the runtime performance of BLAS routines. More specifically, our method uses a machine learning model on-the-fly to automatically select the optimal number of threads for a given GEMM task based on the collected training data. Test results on two different HPC node architectures, one based on a two-socket Intel Cascade Lake and the other on a two-socket AMD Zen 3, revealed a 25 to 40 per cent speedup compared to traditional GEMM implementations in BLAS when using GEMM of memory usage within 100 MB.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Yufan Xia, Marco De La Pierre, Amanda S. Barnard, Giuseppe Maria Junior Barca. 2026-01-14. A Machine Learning Approach Towards Runtime Optimisation of Matrix Multiplication. https://doi.org/10.1109/ipdps54959.2023.00059
Cite the original work for its findings. Save a collection to share your selection of sources.