arXiv · 2608.13868
Exploring High-Bandwidth Flash for Modern LLM Inference: Opportunities and Challenges
Abstract
This work investigates the potential benefits and technical challenges of using high-bandwidth flash (HBF) for large language model (LLM) inference. HBF has gained increasing attention as a promising solution to mitigate memory-capacity bottlenecks in modern LLM-serving systems, but its benefits and challenges remain largely uninvestigated. To address this gap, we thoroughly analyze HBF-based LLM-serving systems under diverse system configurations and operating scenarios in which HBF serves as a main GPU-memory component to handle both reads and writes. Our analysis shows that, despite its limited write performance, HBF can significantly improve the batch size, throughput, and flexibility of LLM-serving systems while reducing the minimum GPU requirements, but realizing these benefits critically depends on sustaining HBM-comparable read bandwidth and requires significant endurance improvements.
Explore related subjects
Keep this discovery
Dowon Son, Yonggon Park, Hyunuk Cho, Hyungkyu Ham, Onur Mutlu, Sungjin Lee, Gwangsun Kim, Jisung Park. 2026-08-14. Exploring High-Bandwidth Flash for Modern LLM Inference: Opportunities and Challenges. https://doi.org/10.1109/lca.2026.3705817
Cite the original work for its findings. Save a collection to share your selection of sources.