arXiv · 2603.19340
Benchmarking NIST-Standardised ML-KEM and ML-DSA on ARM Cortex-M0+: Latency, Rejection-Sampling Variance, and Memory on the RP2040
Abstract
Internet of Things devices with 10 to 20 year lifespans need post-quantum migration, yet the finalised NIST standards remain sparsely benchmarked on the most constrained 32-bit ARM class. We present, to our knowledge, the first isolated algorithm-level benchmarks of ML-KEM (FIPS 203) and ML-DSA (FIPS 204) on ARM Cortex-M0+: all six parameter sets on the RP2040 at 125 MHz, using PQClean reference C. A full ML-KEM-512 key exchange completes in 35.7 ms, 17x faster than an mbedTLS ECDH P-256 baseline on the same hardware. ML-DSA signing varies widely under rejection sampling (coefficient of variation 66.0 to 73.5%), and we recover the rejection-sampling iteration count of individual signatures from wall-clock time alone, a timing side channel that reveals attempt counts but not their content. The Cortex-M0+ incurs only a roughly 2.0x slowdown relative to published Cortex-M4 reference C despite lacking a 64-bit multiplier, DSP, and SIMD instructions. A second physical RP2040 agrees to within 0.075%, showing that build configuration, not silicon, dominates measurement variance. All code and data are released as an open-source benchmark suite.
Explore related subjects
Keep this discovery
Rojin Chhetri, Sijan Dhakal, Asmita Gautam. 2026-03-19. Benchmarking NIST-Standardised ML-KEM and ML-DSA on ARM Cortex-M0+: Latency, Rejection-Sampling Variance, and Memory on the RP2040. https://doi.org/10.5281/zenodo.19393777
Cite the original work for its findings. Save a collection to share your selection of sources.