arXiv · 2312.09401
Inter-Layer Scheduling Space Exploration for Multi-model Inference on Heterogeneous Chiplets
Abstract
To address increasing compute demand from recent multi-model workloads with heavy models like large language models, we propose to deploy heterogeneous chiplet-based multi-chip module (MCM)-based accelerators. We develop an advanced scheduling framework for heterogeneous MCM accelerators that comprehensively consider complex heterogeneity and inter-chiplet pipelining. Our experiments using our framework on GPT-2 and ResNet-50 models on a 4-chiplet system have shown upto 2.2x and 1.9x increase in throughput and energy efficiency, compared to a monolithic accelerator with an optimized output-stationary dataflow.
Explore related subjects
Keep this discovery
Mohanad Odema, Hyoukjun Kwon, Mohammad Abdullah Al Faruque. 2023-12-14. Inter-Layer Scheduling Space Exploration for Multi-model Inference on Heterogeneous Chiplets. https://arxiv.org/abs/2312.09401
Cite the original work for its findings. Save a collection to share your selection of sources.