SearcharxivSearch

arXiv subjects

Sookyung Choi

Publications and source records attributed to Sookyung Choi.

4 recordsLinked to original sources

NELSSA: A GPU-PNM Heterogeneous System for Mixed-Length LLM Serving via Length-based Request Placement

Modern LLMs and their agentic applications are broadening the range of serving workloads, spanning context lengths from a few hundred tokens to hundreds of thousands. As these requests frequently interleave within the same serving window, LLM serving systems must handle highly heterogeneous mixed-length workloads. Such mixed-length workloads expose fundamental inefficiencies in GPU-centric serving architectures, whose throughput depends on large, memory-constrained batches. In this paper, we present NELSSA, an LLM serving system that integrates GPUs with real-world Processing-near-Memory (PNM) accelerator devices to efficiently support mixed-length workloads. NELSSA employs length-based request placement to route short-context requests to GPUs and long-context requests to the PNM tier, incorporating runtime migration to accommodate dynamic context growth without recomputation. We prototype NELSSA as an end-to-end system, implementing device-level sparse attention on PNM, GPU decode kernels, and a host-side runtime that orchestrates scheduling and cross-tier memory movement over a CXL-enabled infrastructure with RPC and RDMA support. Across mixed-length LLM workloads, NELSSA improves decode throughput by up to 5.5x in tokens/sec and reduces P99 latency by up to 15x compared to GPU-only baselines. Our end-to-end prototype and experimental results suggest that integrated GPU-PNM serving, enabled by CXL-based disaggregation, is a promising system paradigm for scalable and flexible LLM infrastructures that support evolving workloads.

cs.AR

MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference

The escalating context length in Large Language Models (LLMs) creates a severe performance bottleneck around the Key-Value (KV) cache, whose memory-bound nature leads to significant GPU under-utilization. This paper introduces Mixture of Shared KV Attention (MoSKA), an architecture that addresses this challenge by exploiting the heterogeneity of context data. It differentiates between per-request unique and massively reused shared sequences. The core of MoSKA is a novel Shared KV Attention mechanism that transforms the attention on shared data from a series of memory-bound GEMV operations into a single, compute-bound GEMM by batching concurrent requests. This is supported by an MoE-inspired sparse attention strategy that prunes the search space and a tailored Disaggregated Infrastructure that specializes hardware for unique and shared data. This comprehensive approach demonstrates a throughput increase of up to 538.7x over baselines in workloads with high context sharing, offering a clear architectural path toward scalable LLM inference.

cs.LG

Hadron and Quarkmonium Exotica

A number of charmonium-(bottmonium-)like states have been observed in $B$-factory experiments. Recently the BESIII experiment has joined this search with a unique data sample collected at the different center of mass energies ranging from 3.9 GeV to 4.42 GeV in which they found new charmonium-like states with non-zero electric charge. We review the status of experimental searchs for quarkomonium-like states and also other types of non-$q\bar{q}$ meson or non-$qqq$ baryons that are predicted by QCD-motivated models. We mainly focus on results from the $B$-factories and BESIII.

hep-ex

Spectroscopy results from Belle

We report recent results on the charmonium and charmoniumlike states based on a large data sample recorded at the $Υ(4S)$ and $Υ(5S)$ resonances with the Belle detector at the KEKB asymmetric-energy $e^+e^-$ collider.

hep-ex