Conflict-Free Color-Clustered Sequential Belief-Propagation Decoding of Quantum LDPC Codes via Reinforcement Learning
Belief-propagation (BP) decoding for quantum low-density parity-check (QLDPC) codes is attractive due to its low complexity and low latency, but it is often limited by short cycles, degeneracy, and convergence failures. Reinforcement-learning-based sequential BP decoding (RL-S) improves BP by learning a syndrome-dependent variable-node (VN) update order, but its VN-by-VN schedule has limited within-iteration parallelism. In this paper, we propose a conflict-free color-clustered extension of RL-S. We construct a VN conflict graph in which two VNs are adjacent if they share an X-type or Z-type check, and color this graph so that same-color VNs have disjoint check neighborhoods. This also prevents VNs from the same Tanner 4- or 6-cycle from being updated simultaneously. During decoding, the trained VN-level Q-table selects a seed VN, and all remaining VNs with the same color are updated in parallel using the same pre-batch messages. For the [[288,12,18]] bivariate-bicycle code over the depolarizing channel, our proposed decoder achieves block-error-rate performance close to VN-level RL-S while reducing the scheduling decisions from 288 VNs to 11 color classes per BP iteration.