arXiv · 2504.19651
Fooling the Decoder: An Adversarial Attack on Quantum Error Correction
Abstract
Neural network decoders are essential for fault-tolerant quantum computation, yet their internal mechanisms remain poorly understood. Leading machine learning decoders use recurrent and transformer models (e.g., AlphaQubit), and reinforcement learning (RL) is increasingly used to train such models. In this work we ask whether an RL surface code decoder, DeepQ, can be made to fail systematically. Our white-box search repeatedly draws candidate syndrome volumes from the unmodified noise model and retains the one the decoder's learned $Q$-function is most pessimistic about. At code distance 5 under phenomenological depolarising noise, it cuts the average logical qubit lifetime from roughly $10^5$ cycles to 60. We quantify what this costs: the retained error histories are individually plausible, but the selected trajectory is roughly a $10^{-18}$ tail event under undisturbed noise, so the effect is bought entirely by the ability to select among physically realisable error configurations. Controls confirm this is not an artefact of a weak target: the decoder resists noise fluctuations, tolerates an MWPM referee, and is fault tolerant at checkable distances. The result is a robustness upper bound for one decoder under one noise model, but to our knowledge the first adversarial attack of a machine learning QEC decoder.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Jerome Lenssen, Alexandru Paler. 2025-04-28. Fooling the Decoder: An Adversarial Attack on Quantum Error Correction. https://arxiv.org/abs/2504.19651
Cite the original work for its findings. Save a collection to share your selection of sources.