arXiv · 2603.03383
Accelerating OpenPangu Inference on NPU via Speculative Decoding
Abstract
To mitigate the Memory Wall bottleneck encountered by Large Language Models (LLMs) during inference on \textbf{NPU} hardware, and addressing the scarcity of native support for mainstream speculative decoding algorithms on domestic infrastructure, this study presents an end-to-end speculative inference acceleration scheme for OpenPangu-7B.
Explore related subjects
Keep this discovery
Yuntao Dai, Jing Wu, Hang Gu, Teng Wang. 2026-03-03. Accelerating OpenPangu Inference on NPU via Speculative Decoding. https://arxiv.org/abs/2603.03383
Cite the original work for its findings. Save a collection to share your selection of sources.