Searcharxiv⌕ Search

arXiv · 2609.35015

A performance enhancement of the Payne-Hanek range reduction algorithm

Abstract

Range reduction plays a crucial role in the accuracy and performance of evaluating trigonometric functions, and is often the primary bottleneck for large floating-point inputs. While fast algorithms such as Cody--Waite work efficiently over narrow intervals, the Payne--Hanek algorithm remains the standard technique for accurate reduction across large floating-point inputs. However, existing implementations of Payne--Hanek suffer from high latency due to heavy branching, conversion overheads, and the use of multi-word integer arithmetic, which hinders SIMD vectorization. In this paper, we analyze and present a branch-free variation of the Payne--Hanek algorithm using only floating-point arithmetic. Our method operates directly over large double-precision inputs ($|x| \ge 2^{16}$) and is well suited to hardware with FMA instructions. We formulate the precision constraints in terms of a truncation error budget, construct a compact lookup table indexed by the input exponent, and prove that the reduced argument is accurate to within one ulp for every input. The same routine can serve both as the complete range reduction of a single-stage implementation and as the fast path of a correctly rounded one, and it improves both latency and throughput over existing implementations. The algorithm is currently implemented in the LLVM libc project.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Tue Ly. 2026-09-28. A performance enhancement of the Payne-Hanek range reduction algorithm. https://arxiv.org/abs/2609.35015

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

KnottedGraph: Scalable knotted-graph topology for scientific and mathematical discovery

Scientific data span heterogeneous structures, including coordinates, networks, surfaces, volumes and fields, yet their topology can be quantified within a common framework through graph connectivity, cycle structure, genus and spatial embedding. Graph- and homology-based summaries do not determine spatial embedding, while standard knot and link polynomials require extensions to accommodate branching graphs. Here, we introduce KnottedGraph, a computational framework that converts such scientific representations to knotted graphs that retain graph connectivity and spatial embedding together. It constructs projected diagrams and PD codes, enabling various topological analyses, including Yamada-polynomial evaluation for topological classification. For scalable exact evaluation, it combines partial resolutions that leave the same unresolved connections and optimizes their processing order; the resulting algorithm is verified against published topological invariants of knotted graphs with up to 500 crossings. This scalability enables us to introduce an LLM-assisted mathematical-discovery methodology, in which computational topological data generated across knotted-graph families are used to identify candidate closed-form formulas. With this approach, we identify analytical Yamada-polynomials for generic graph motif families exhibiting Abelian and non-Abelian word sequences. Together, these scalable capabilities make knotted-graph topology computationally accessible across scientific domains, enabling large-scale classification and introducing a route from topological data to LLM-assisted AI4Math discovery.

cs.MS↗

DistributedDesignOptimizer: A modular Python framework for Setup, Execution and Processing of Distributed Design Optimization

Multidisciplinary Design Optimization (MDO) enables the application of optimization algorithms to complex engineered systems, ranging from aerospace and automotive to robotics and microelectronics. Such multi-component systems typically consist of coupled subsystems, each characterized by its own design variables, constraints and objectives. Without adequate coordination of these couplings, subsystems may be optimal in isolation while the overall system remains suboptimal or even infeasible. Distributed design optimization coordinates coupled subsystem optimization problems while allowing them to retain control over their local design variables. Although promising coordination methods exist, their application is hindered by the lack of a comprehensive software framework which supports intuitive problem definition, provides suitable distributed coordination algorithms, is modular and extensible, enables interactive (post-)processing, and supports the distributed computation of subsystem executions. This work derives requirements for such software, assesses existing frameworks against them, and introduces DistributedDesignOptimizer, an open-source Python framework to fill this gap.

cs.MS↗

AWE: Adaptive Weight Encoding for Exact Integer Matrix Products with Fewer GEMMs on FP4 Tensor Cores

Emulation of high-accuracy floating-point matrix multiplication, as in the Ozaki scheme, splits the inputs into low-precision components and multiplies them pairwise. These products must be error-free, and each is an integer matrix product times a scale factor. FP4 Tensor Cores are the fastest on the NVIDIA B200 and B300 but cannot hold INT8 operands. The FP4 values scaled by 2 form the set $S = \{0, \pm1, \pm2, \pm3, \pm4, \pm6, \pm8, \pm12\}$, which contains every residue modulo 13, so with carries any integer splits into base-13 digits that FP4 can store. Prior work splits each INT8 operand into 3 such digits (limbs) with weights $(1, 13, 169)$ and multiplies them pairwise, 9 FP4 matrix multiplications (GEMMs) for INT8$\times$INT8. The classical ways to reduce products, such as the Karatsuba and Toom--Cook methods, do not apply as they stand: sums of limbs reach $\pm 24$ and leave $S$. This paper asks how many FP4 GEMMs are needed for one integer matrix product. We propose Adaptive Weight Encoding (AWE): the limbs take freely chosen integer weights, the stored planes are linear combinations of limbs, and the exact product is the sum of the FP4 GEMMs scaled by reconstruction coefficients. For each input range, we searched these choices for encodings with fewer products and found INT8$\times$INT8 in 6 products and INT4$\times$INT8 in 4. The formulation also holds modulo $m$, which covers the residue number systems of Ozaki scheme II: for the FP64 significand, the 75 products of prior work are reduced to 59. The boundary in product count between encodings with and without residues lies near input width 15. We release the encodings found.

cs.MS↗