arXiv · 2506.17255
UltraSketchLLM: Sub-1-Bit LLM Compression via Sketch and Hardware-Friendly Operators
Abstract
Large language models (LLMs) require larger GPU memory size these days, necessitating efficient and extreme weight compression methods. Existing compression methods are either theoretically limited by 1 bit per weight or face severe performance degradation and inefficiency. To deploy LLMs in resource-constrained scenarios, we introduce UltraSketchLLM, compressing LLMs with data sketch. It reduces peak GPU memory footprint with a high compression rate down to 0.5 bit per weight. Combined with hardware-friendly implementation, UltraSketchLLM keeps tolerable performance degradation and extremely low latency overhead with 14.9x speedup compared to naive sketch solution.
Explore related subjects
Keep this discovery
Sunan Zou, Xueting Sun, Ziyun Zhang, Guojie Luo. 2025-06-08. UltraSketchLLM: Sub-1-Bit LLM Compression via Sketch and Hardware-Friendly Operators. https://doi.org/10.1145/3770743.3804314
Cite the original work for its findings. Save a collection to share your selection of sources.