arXiv · 2503.08415
TokenSim: Enabling Hardware and Software Exploration for Large Language Model Inference Systems
Abstract
The increasing demand for large language model (LLM) serving has necessitated significant advancements in the optimization and profiling of LLM inference systems. As these models become integral to a wide range of applications, the need for efficient and scalable serving solutions has grown exponentially. This work introduces TokenSim, a comprehensive hardware and software exploration system designed specifically for LLM inference. TokenSim is characterized by its support for extensible system optimizations including scheduling and memory management. We validate the results with systems running with realworld datasets, achieving an error rate of less than 1%. Furthermore, TokenSim facilitates various insightful explorations into the performance and optimization of LLM serving systems.
Explore related subjects
Keep this discovery
Feiyang Wu, Zhuohang Bian, Guoyang Duan, Tianle Xu, Junchi Wu, Teng Ma, Yongqiang Yao, Ruihao Gong, Youwei Zhuo. 2025-03-11. TokenSim: Enabling Hardware and Software Exploration for Large Language Model Inference Systems. https://arxiv.org/abs/2503.08415
Cite the original work for its findings. Save a collection to share your selection of sources.