arXiv · 2606.09905
Toward a Small ML Runtime Stack for Raspberry Pi 5 QPUs
Abstract
We present a QPU-first ML runtime stack for Raspberry Pi 5's VideoCore VII QPU, built on top of the py-videocore7 assembly library. The system comprises reusable tiled matrix-multiplication substrate, GEMM-backed convolution, a single-head attention-style core, persistent executors, and integer execution based on smul24 instructions. For dense integer kernels, packed INT16-input with INT32 accumulation achieves nearly two orders of magnitude higher throughput over NumPy. Across operations (min/max, pooling, convolution, attention), we report improved performance over both PyTorch and NumPy. Our preliminary results indicate that Raspberry QPUs can serve as a practical execution substrate towards accelerating AI model execution at the edge.
Explore related subjects
Keep this discovery
Yiannis Hadjiyianni, Panagiotis Michelakis, Dimitrios Stamoulis. 2026-06-06. Toward a Small ML Runtime Stack for Raspberry Pi 5 QPUs. https://arxiv.org/abs/2606.09905
Cite the original work for its findings. Save a collection to share your selection of sources.