TY - RPRT TI - 98$\times$ Faster LLM Routing Without a Dedicated GPU: Flash Attention, Prompt Compression, and Near-Streaming for the vLLM Semantic Router AU - Xunzhuo Liu AU - Bowei He AU - Xue Liu AU - Andy Luo AU - Haichen Zhang AU - Huamin Chen PY - 2026 UR - https://arxiv.org/abs/2603.12646 ID - 2603.12646 ER -