Searcharxiv⌕ Search

arXiv · 2609.35263

WavePP: High-Throughput Pipeline Parallel LLM Prefill under Prefix Reuse

Abstract

Pipeline parallelism can improve prefill throughput by processing multiple request chunks concurrently across different stages of the model. However, keeping the pipeline fully utilized requires efficient scheduling and request preparation. In systems where stages retain and evict cache state independently, a local cache hit does not guarantee that the same prefix can be reused across the pipeline. Here, coordination overhead can impede request admission cadence and thus reduce overall throughput. In this paper, we present WavePP, a prefill runtime built on top of TensorRT-LLM that addresses these challenges by overlapping request admission with pipeline execution. WavePP asynchronously finds a prefix that can be reused across all stages, protects the cached state, and reserves space for the remaining input while earlier requests continue to execute. It subsequently plans the chunk sizes of each request dynamically to maximize pipeline fill. Each stage then completes the local preparation before executing the request. In the same system and pipeline topology, WavePP improves TensorRT-LLM's prefill throughput in 37 of 40 tested settings on GLM 5.2 and MiniMax M2.7. At concurrency 128 with high cache reuse, these changes increase throughput by factors of 2.91 and 2.02, respectively. Across 28 Kimi K3 settings, WavePP also has the highest measured throughput in all 18 settings at concurrency eight or higher, compared with tensor/expert-parallel and pipeline-parallel baselines from TRT-LLM, SGLang, and vLLM.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Aaryam Sharma. 2026-09-28. WavePP: High-Throughput Pipeline Parallel LLM Prefill under Prefix Reuse. https://arxiv.org/abs/2609.35263

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications

Large Language Models (LLMs) are increasingly deployed in complex multi-agent applications that rely on external function calls. This workload creates severe performance challenges for the KV Cache: spatial contention leads to the eviction of critical agents' caches and temporal underutilization leaves the cache of agents stalled on long-running function calls idling in GPU memory. We present TokenCake, a KV-Cache-centric serving framework that bridges this gap by co-optimizing scheduling and memory management through an agent-aware design. TokenCake's Temporal Scheduler employs an event-driven, opportunistic policy to proactively offload idle KV Caches during function calls and uses predictive uploading to hide data transfer latency. TokenCake's Spatial Scheduler uses dynamic memory partitioning, guided by a hybrid priority metric combining graph structure and runtime state, to reserve GPU memory for critical-path agents. Our evaluation on representative multi-agent benchmarks shows that TokenCake reduces end-to-end latency by up to 47.06% in memory-constrained settings and improves effective GPU utilization by 16.9 percentage points compared to vLLM. TokenCake is publicly available at https://github.com/pkulemonade/TokenCake.

cs.DC↗

SpSYRK: Half the Work in Distributed Sparse Matrix Multiplication

The symmetric rank-$k$ update (SYRK), $C = AA^\top$, computes the dot product of each pair of rows of $A$, producing the Gram matrix $C$. Its sparse variant underpins similarity search in machine learning, graph analytics, and genomics, including Jaccard similarity on datasets too large for a single node. Despite the symmetry in its inputs and outputs, existing distributed sparse matrix multiplication algorithms such as Sparse SUMMA treat sparse SYRK as generic multiplication, computing the full output and materializing the explicit transpose even when the calling application uses only one triangle. Prior distributed $AA^\top$ computations in similarity search and genome assembly inherit this overhead from the underlying SpGEMM. This paper presents SpSYRK and CommSpSYRK, two distributed sparse SYRK algorithms that exploit symmetry. The first, SpSYRK, partitions the off-diagonal blocks of the output between the upper and lower triangular regions of the process grid and computes only the lower-triangular part of each diagonal block, halving per-process computation compared with state-of-the-art distributed SpGEMM. The second, CommSpSYRK, further reorders communication to avoid forming $A^\top$, which reduces per-process communication volume. On 32 nodes of the Perlmutter supercomputer, SpSYRK achieves a 2$\times$ speedup over an optimized Sparse SUMMA on matrices where local multiplication dominates the runtime; the advantage narrows on communication-bound inputs, a dependence that the cost model predicts from the arithmetic intensity. CommSpSYRK fixes this and consistently achieves superior scaling at high process counts. The approach is a drop-in replacement for any application computing $C = AA^\top$ via a distributed SpGEMM routine, and its triangular output can be consumed directly by subsequent operations, reducing both computation and memory footprint.

cs.DC↗

Brain API: An Intent-Aware Control Plane for Policy-Governed Agentic Systems

Contemporary cloud and distributed systems expose control through resource-centric abstractions: services, deployments, network flows, execution graphs. Agentic and tool-augmented systems have meanwhile shifted application logic toward intent-driven, adaptive execution. Existing control planes, workflow engines and service meshes lack abstractions for intent-level decision governance: they cannot represent high-level goals as first-class control objects, cannot enforce policy over the mapping from intent to execution plan, and cannot produce auditable records of why one execution path was chosen over its alternatives. Control logic is therefore embedded in application code, leaving systems brittle, opaque and hard to govern. We propose Brain API, an intent-aware control plane for policy-governed agentic systems. Its central contribution is the decision artifact: a durable, versioned, auditable record of how an intent became an executable plan, capturing which policies applied, which capabilities were evaluated, which alternatives were rejected, and why. A motivating use case is agentic datasets: datasets participating as policy-governed capabilities under residency, compliance and cost constraints. We evaluate a prototype of the decision layer against two external policy corpora we did not author. On the OPA Gatekeeper constraint library it agrees with the library's own published verdicts on 42 of 42 encodable cases, 19 admit and 23 deny. On Cedar example policies, labeled by differential testing against its reference implementation, a deliberately dissimilar domain exposed three defects in our model, including a default-allow assumption that would have inverted every authorization policy. The evaluation covers policy filtering and selection; candidate generation, context signals, ranking and plan synthesis are not measured, nor is decision latency under load.

cs.DC↗