arXiv · 2609.08424
PENDA: An Efficient Processing Element via Norm-of-Difference for Deep Learning Accelerators
Abstract
Inner product computation dominates the computational cost of deep learning models; thus, accelerating this primitive is key to improving hardware efficiency. However, most existing techniques rely on approximations, which can degrade model accuracy. To preserve exactness while optimizing hardware, this paper presents PENDA (processing element via norm-of-difference architecture), which leverages the law of cosines to recast multiplications as squared-difference operations. Replacing multiply-accumulate units with the proposed norm-of-difference units yields 11~36%, 5~48%, and 11~19% reductions in area, energy, and clock period, respectively, for the PE array of a deep learning accelerator.
Explore related subjects
Keep this discovery
Kai-Chieh Hsu, Tian-Sheuan Chang. 2026-09-08. PENDA: An Efficient Processing Element via Norm-of-Difference for Deep Learning Accelerators. https://arxiv.org/abs/2609.08424
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.