arXiv · 2404.16317
FLAASH: Flexible Accelerator Architecture for Sparse High-Order Tensor Contraction
Abstract
Tensors play a vital role in machine learning (ML) and often exhibit properties best explored while maintaining high-order. Efficiently performing ML computations requires taking advantage of sparsity, but generalized hardware support is challenging. This paper introduces FLAASH, a flexible and modular accelerator design for sparse tensor contraction that achieves over 25x speedup for a deep learning workload. Our architecture performs sparse high-order tensor contraction by distributing sparse dot products, or portions thereof, to numerous Sparse Dot Product Engines (SDPEs). Memory structure and job distribution can be customized, and we demonstrate a simple approach as a proof of concept. We address the challenges associated with control flow to navigate data structures, high-order representation, and high-sparsity handling. The effectiveness of our approach is demonstrated through various evaluations, showcasing significant speedup as sparsity and order increase.
Explore related subjects
Keep this discovery
Gabriel Kulp, Andrew Ensinger, Lizhong Chen. 2024-04-25. FLAASH: Flexible Accelerator Architecture for Sparse High-Order Tensor Contraction. https://arxiv.org/abs/2404.16317
Cite the original work for its findings. Save a collection to share your selection of sources.