Searcharxiv⌕ Search

arXiv · 2609.37538

AVSG: Accelerated Vectorized Sparse Gather for Efficient KV Cache Offload in Sparse-Attention LLM Serving

Abstract

Dynamic sparse attention reduces long-context attention computation by selecting only a subset of tokens, but still requires access to the full KV cache, leaving serving memory-bound. Offloading the KV cache to host memory reduces device memory pressure but places H2D transfers on the decoding critical path. In DSA, substantial overlap in selected KV entries across decoding steps creates an opportunity for device-resident reuse. Exploiting this reuse efficiently, however, presents three critical challenges: costly matching of selected tokens against entries retained in the HBM buffer, uneven distribution of the remaining H2D transfers across accelerator cores, and retaining frequently accessed KV entries within limited HBM capacity. We present AVSG, an operator that addresses these challenges for DSA while preserving its exact token selections. AVSG uses vectorized hash matching to identify reusable HBM-buffer slots and reserve slots for entries requiring transfer, then evenly partitions these H2D transfers across accelerator cores. Lifetime-based buffer management retains frequently accessed KV entries in the device buffer across decoding steps, and shared slot reservations extend reuse across tokens within a multi-token prediction iteration. On a single NPU, vectorized hash matching is 2.80x faster than scalar dual-pointer matching, miss-only transfer raises effective H2D bandwidth by up to 34.44x over request-level assignment, and an 8K-entry buffer reaches a 94.83% hit rate on real requests. These gains translate into end-to-end improvement: on a serving stack processing a real-world production dataset, AVSG reduces time per output token by 39% and increases output throughput by 1.27x relative to the same KV offload layout without HBM-buffer reuse, demonstrating the benefit of efficient matching, balanced transfer, and effective HBM residency in production-scale serving.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wenwei Kuang, Xiangyu Wang, Chong Wu, Jun Wang, Weijie Zhang, Brian K Chen, Longwen Lan, Ken Zhang. 2026-09-29. AVSG: Accelerated Vectorized Sparse Gather for Efficient KV Cache Offload in Sparse-Attention LLM Serving. https://arxiv.org/abs/2609.37538

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The Symmetric Location Problem: a Song of Efficiency and Robustness

The aim of this Lecture Note is to introduce the Signal Processing (SP) community to a powerful yet still under-utilised tool: the semiparametric statistics. In short, the semiparametric framework allows us to estimate or perform hypothesis testing on a finite-dimensional parameter $\boldsymbolθ \in Θ\subseteq \mathbb{R}^p$ in the presence of an \textit{infinite-dimensional nuisance parameter} $g \in \mathcal{S}$ (i.e. a function), such as the density of the noise. Clearly, this framework is general enough to include almost every SP application. Remarkably, as the title suggests drawing on George R. R. Martin's famous book series, the greatest advantage of semiparametric statistics over parametric and non-parametric ones lies in the fact that it is able to reconcile two seemingly dichotomous concepts: statistical efficiency and distributional robustness. To explain exactly what this means, in this Lecture Note we will focus our attention on the famous and fundamental symmetric location problem.

eess.SP↗

Single-Voxel Wireless NeRF for Spatial Spectrum Prediction

Wireless channel measurements across multiple spatial directions are crucial for AI-driven applications, such as RF digital twins and integrated communication and sensing. However, collecting channel data across large scenes is labor-intensive. Wireless NeRFs address this challenge by learning propagation behavior from sparse measurements and synthesizing channel spatial spectrum magnitude at unseen locations. However, existing wireless NeRFs inherit dense volumetric sampling from vision NeRFs, which requires substantial computation. This paper asks whether such dense sampling is necessary for predicting magnitudes of the wireless spatial spectrum. We empirically show that wireless NeRFs are over-parameterized for this task and introduce SV-INGP, a sparse volumetric sampling variant of Instant Neural Graphics Primitives (INGP). Across real-world and simulated datasets, SV-INGP matches the median Structural Similarity Index Measure (SSIM) of the NeRF2 baseline while reducing training time by 184x. These results generalize across LoS and NLoS scenes, sub-6 and millimeter-wave frequencies, and antenna array tapering configurations, suggesting a simpler and more efficient design path for RF digital twins

eess.SP↗

Best Practices in EEG Analysis: Preprocessing, Modeling, and Machine Learning

Electroencephalography (EEG) analysis requires careful choices in preprocessing, statistical modeling, and machine learning because EEG signals are highly susceptible to artifacts, volume conduction, low signal-to-noise ratio, and substantial inter-subject variability. This chapter provides a practical and methodological guide to modern EEG analysis, spanning EEG preprocessing, artifact removal, filtering, bad-channel detection and interpolation, re-referencing, independent component analysis (ICA), and preprocessing of simultaneous EEG-fMRI recordings. We review major approaches for computational EEG analysis, including event-related potentials (ERPs), time-frequency analysis, functional and effective connectivity, source localization, multivariate decoding, permutation testing, and multiple-comparison correction. We then examine machine-learning methods for EEG, from feature-based classifiers to deep learning and emerging EEG foundation models, with emphasis on cross-subject generalization, limited-data regimes, data leakage, evaluation metrics, and fair benchmarking. Reproducibility is treated as a core requirement throughout, including transparent preprocessing, BIDS-EEG data organization, standardized derivatives, preservation of raw data, and FAIR data practices. The chapter is intended as a practical reference for researchers developing reliable, interpretable, and reproducible EEG analysis and machine-learning pipelines.

eess.SP↗