SOCKET: SOft Collision Kernel EsTimator for Sparse Attention
Exploiting sparsity is key to efficient long-context inference, as attention dominates the cost of autoregressive decoding. Sparse attention reduces this cost by restricting computation to a subset of tokens, but its effectiveness hinges on fast and accurate token scoring and selection at inference time. Data-agnostic approaches offer an attractive way to perform this selection, but often incur substantial memory overhead to maintain high recall. We revisit Locality-Sensitive Hashing (LSH) and introduce SOCKET, a SOft Collision Kernel EsTimator that replaces hard bucket matches with probabilistic, similarity-aware aggregation. Traditional LSH relies on binary collision signals, providing limited information for ranking tokens and necessitating many hash tables for accurate retrieval. In contrast, soft LSH accumulates graded collision evidence across hash tables, closely preserving the true top-$k$ ordering with significantly less memory. This reframes LSH from a candidate-generation mechanism into a principled scoring kernel for sparse attention. Building on this insight, SOCKET enables efficient token selection without ad hoc voting and matches or outperforms existing sparse attention methods across multiple long-context benchmarks and diverse language models. With a custom set of CUDA/Triton kernels for scoring, selection, and attention, SOCKET achieves up to approximately $1.5\times$ higher throughput than FlashAttention. Code is open-sourced at https://github.com/amarka8/SOCKET.