arXiv · 2610.09766
Performance Portable $\mathrm{SU}(N)$ Lattice Gauge Theory Simulation with Kokkos
Abstract
The increasing diversity of high performance computing systems makes separate, architecture specific implementations of lattice gauge theory algorithms costly to maintain. We present \texttt{kwqft}, a performance portable Kokkos implementation of Wilson pure gauge Monte Carlo simulation for $\mathrm{SU}(N)$ Yang-Mills theory in an arbitrary number of space-time dimensions. The gauge group order $N$ and the dimension $D$ are compile time parameters. A single source targets the Serial, OpenMP, CUDA, HIP, and SYCL execution spaces, with MPI halo exchange overlapped with interior updates. The implementation reproduces the exact two-dimensional plaquette and published three and four dimensional values for gauge groups up to $\mathrm{SU}(17)$. On an NVIDIA A100 the Kokkos CUDA backend is competitive with a native CUDA code, SIMD acceleration improves the OpenMP path on Armv9 processors, and a large scale speedup is demonstrated for an $\mathrm{SU}(4)$ lattice on the LineShine supercomputer, currently ranked first on the TOP500 list.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Wei Sun. 2026-10-07. Performance Portable $\mathrm{SU}(N)$ Lattice Gauge Theory Simulation with Kokkos. https://arxiv.org/abs/2610.09766
Cite the original work for its findings. Save a collection to share your selection of sources.