arXiv · 2506.04645
Inference economics of language models
Abstract
We develop a theoretical model that addresses the economic trade-off between cost per token versus serial token generation speed when deploying LLMs for inference at scale. Our model takes into account arithmetic, memory bandwidth, network bandwidth and latency constraints; and optimizes over different parallelism setups and batch sizes to find the ones that optimize serial inference speed at a given cost per token. We use the model to compute Pareto frontiers of serial speed versus cost per token for popular language models.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Ege Erdil. 2025-06-05. Inference economics of language models. https://arxiv.org/abs/2506.04645
Cite the original work for its findings. Save a collection to share your selection of sources.