arXiv · 2609.35015
A performance enhancement of the Payne-Hanek range reduction algorithm
Abstract
Range reduction plays a crucial role in the accuracy and performance of evaluating trigonometric functions, and is often the primary bottleneck for large floating-point inputs. While fast algorithms such as Cody--Waite work efficiently over narrow intervals, the Payne--Hanek algorithm remains the standard technique for accurate reduction across large floating-point inputs. However, existing implementations of Payne--Hanek suffer from high latency due to heavy branching, conversion overheads, and the use of multi-word integer arithmetic, which hinders SIMD vectorization. In this paper, we analyze and present a branch-free variation of the Payne--Hanek algorithm using only floating-point arithmetic. Our method operates directly over large double-precision inputs ($|x| \ge 2^{16}$) and is well suited to hardware with FMA instructions. We formulate the precision constraints in terms of a truncation error budget, construct a compact lookup table indexed by the input exponent, and prove that the reduced argument is accurate to within one ulp for every input. The same routine can serve both as the complete range reduction of a single-stage implementation and as the fast path of a correctly rounded one, and it improves both latency and throughput over existing implementations. The algorithm is currently implemented in the LLVM libc project.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Tue Ly. 2026-09-28. A performance enhancement of the Payne-Hanek range reduction algorithm. https://arxiv.org/abs/2609.35015
Cite the original work for its findings. Save a collection to share your selection of sources.