arXiv · 2609.33224
PackServe: SLO-Aware Request Scheduling for Agentic LLM Serving at Scale
Abstract
Request scheduling is a key challenge in large-scale clusters serving agentic large language model (LLM) workloads. An effective scheduler must preserve key-value cache (KVC) reuse across long, shared prefixes, meet token-level latency service-level objectives (SLOs), and minimize GPU resource footprint. Existing schedulers struggle to reconcile these requirements: request consolidation can sacrifice cache locality and increase prefill/decode interference, compromising both SLO attainment and resource efficiency. We present PackServe, a scheduler designed to reduce resource costs while meeting latency SLOs for agentic LLM serving. PackServe uses compact white-box models to predict latency under prefill/decode interference. Guided by these predictions, it packs requests onto fewer serving instances while preserving KVC reuse and SLO constraints, trading available latency headroom for improved per-instance throughput. Evaluation on 64 NVIDIA H20 GPUs shows that PackServe uses up to 16.8% and 24.6% fewer GPU-hours than state-of-the-art schedulers under 30-ms and 50-ms TPOT targets, respectively, while meeting the target TPOT objectives. PackServe has also been deployed in our production cluster comprising over 1000 GPUs, where it reduces the resource footprint by 34.7% compared with the original production scheduler.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zhiyuan Tan, Dejiang Zhu, Jingzhe Jiang, Yihao Zheng, Yang Tian, Tao Wang, Minchen Yu. 2026-09-27. PackServe: SLO-Aware Request Scheduling for Agentic LLM Serving at Scale. https://arxiv.org/abs/2609.33224
Cite the original work for its findings. Save a collection to share your selection of sources.