Searcharxiv⌕ Search

arXiv subjects

Jeonggeun Seo

Publications and source records attributed to Jeonggeun Seo.

3 recordsLinked to original sources

High-Throughput Normalized Min-Sum Belief Propagation Decoding for Quantum LDPC Codes with Near-Memory Processing

Real-time quantum error correction requires classical decoders to process growing syndrome workloads with low and predictable latency. For quantum low-density parity-check (qLDPC) codes, iterative belief propagation (BP) repeatedly updates messages over sparse Tanner graphs, creating substantial memory-access and data-movement demands. We map normalized Min-Sum BP decoding of the [[144,12,12]] Bivariate Bicycle qLDPC code onto a DPU-based Processing-in-Memory (PIM) architecture. Within each DPU, 11 tasklets cooperatively decode one syndrome, while multiple DPUs process independent syndrome instances in parallel. Using uPIMulator and a data-qubit Pauli error model with ideal syndrome measurements, we compare throughput, per-syndrome processing time, logical error rate (LER), and single-syndrome tail latency against a 16-logical-CPU baseline. At a component-wise physical error probability of p=0.001 and one BP iteration, the projected aggregate kernel throughput of 2,560 DPUs reaches 1.071 x 10^7 decodes/s, compared with 1.22 x 10^6 decodes/s for the CPU, an 8.8x improvement. From two iterations onward, the measured LER remains below the physical error probability for every evaluated value of p. For one to five iterations, the maximum sampled serialized X+Z DPU compute latency remains below the 1 ms decoder-side reference for trapped-ion QEC, reaching approximately 0.873 ms at five iterations. These results show that near-memory processing can provide high aggregate throughput and sub-millisecond compute latency for qLDPC BP decoding under the evaluated conditions.

quant-ph↗

Reducing Postselection Overhead in Magic-State Cultivation by In-Patch Multiplexing

Fault-tolerant quantum computing requires high-fidelity logical magic states for implementing non-Clifford operations. Magic-state cultivation provides a lower-overhead route to logical magic-state preparation, but its efficiency is limited by postselection loss during the early injection-and-cultivation stages. In this work, we propose an in-patch multiplexing scheme that uses early-stage idle resources within a single logical patch to create multiple local cultivation opportunities. A candidate that passes the early stages is forwarded to the standard escape pathway, while the escape stage and the decoder-based acceptance procedure are kept identical to those of the single-site baseline. Under a uniform depolarizing noise model with idle noise, the proposed protocol substantially reduces the injection-and-cultivation discard rate and the expected number of attempts required to obtain an accepted early-stage candidate. At a physical error rate of \(p=2\times10^{-3}\), the injection-and-cultivation expected attempts are reduced by \(45.46\%\) for \(d_1=3\) and by \(72.91\%\) for \(d_1=5\), relative to the single-site MSC baseline. In the direct full-cycle evaluation including escape, the expected attempts per kept logical output are further reduced by \(49.04\%\) for \(d_1=3\) and by \(78.69\%\) for \(d_1=5\) at the same physical error rate. The full-cycle cost curves are shifted toward smaller expected attempts, while the final logical-error behavior remains governed by the escape-stage gap threshold. These results show that in-patch multiplexing can reduce postselection overhead while preserving the standard magic-state cultivation framework.

quant-ph↗

Constraint-Optimal Driven Allocation for Scalable QEC Decoder Scheduling

Fault-tolerant quantum computing (FTQC) requires fast and accurate decoding of Quantum Error Correction (QEC) syndromes. However, in large-scale systems, the number of available decoders is much smaller than the number of logical qubits, leading to a fundamental resource shortage. To address this limitation, Virtualized Quantum Decoder (VQD) architectures have been proposed to share a limited pool of decoders across multiple qubits. While the Minimize Longest Undecoded Sequence (MLS) heuristic has been introduced as an effective scheduling policy within the VQD framework, its locally greedy decision-making structure limits its ability to consider global circuit structure, causing inefficiencies in resource balancing and limited scalability. In this work, we propose Constraint-Optimal Driven Allocation (CODA), an optimization-based scheduling algorithm that leverages global circuit structure to minimize the longest undecoded sequence length. Across 19 benchmark circuits, CODA achieves an average 74\% reduction in the longest undecoded sequence length. Crucially, while the theoretical search space scales exponentially with circuit size, CODA effectively bypasses this combinatorial explosion. Our evaluation confirms that the scheduling time scales linearly with the number of qubits, determined by physical resource constraints rather than the combinatorial search space, ensuring robust scalability for large-scale FTQC systems. These results demonstrate that CODA provides a global optimization-based, scalable scheduling solution that enables efficient decoder virtualization in large-scale FTQC systems.

quant-ph↗