Searcharxiv⌕ Search

arXiv subjects

Yasunari Suzuki

Publications and source records attributed to Yasunari Suzuki.

At least 19 recordsLinked to original sources

NAQsim: Full-Stack Architecture Simulation Framework for Fast and Space-Efficient Neutral Atom Quantum Computing

Technological advances in neutral-atom platforms have opened a new path toward designing efficient protocols for fault-tolerant quantum computing (FTQC). However, each physical operation is still orders of magnitude slower than on other platforms, such as superconducting qubits. Therefore, architects must identify fast and efficient FTQC architectures, which require reliable modeling tools to explore various design choices across the software, classical hardware, and quantum device stacks of neutral-atom platforms. In this paper, we propose NAQsim, an open-source simulation framework for neutral-atom FTQC architectures based on transversal gates. As its key feature, NAQsim enables detailed full-stack architecture evaluation that opens opportunities to explore previously overlooked performance bottlenecks. As a first use case for NAQsim, we identify one such bottleneck, patch rotations, perform full-stack co-optimization, and finally derive a near-rotation-free architecture (D3-ROT). For practical FTQC benchmark workloads, D3-ROT achieves a 2.27x speedup from this single bottleneck alone, with minimal footprint overhead. These results point to a much broader space of full-stack optimizations that NAQsim makes accessible for transversal surface-code architectures and beyond.

quant-ph↗

Do Not Let CNOTs Overwhelm the Decoder: Scheduling Transversal Gates for Fast FTQC

Transversal CNOT (TCNOT) gates can accelerate fault-tolerant quantum computation (FTQC) in the surface code by reducing the number of syndrome extraction rounds required between logical operations from $O(d)$ to $O(1)$. This is particularly attractive for quantum platforms with long-range connectivity, such as neutral atoms. However, dense TCNOT schedules substantially increase the classical decoding workload. TCNOTs propagate errors across multiple surface-code patches, enlarging the spatiotemporal region that must be decoded jointly. Consequently, denser TCNOT schedules increase decoding latency and memory requirements and potentially exceed available decoder capacity. Moreover, because the detector error model (DEM) of each decoding window depends on the TCNOT schedule, exhaustively precomputing all possible window-level DEMs is infeasible, requiring just-in-time (JIT) DEM compilation. Thus, the practical benefit of TCNOT gates is limited not only by quantum hardware performance but also by classical decoding and DEM-compilation capacity. We introduce PACE, a decoder-aware scheduling framework for TCNOT-based FTQC. PACE first mitigates the decoder-side costs of aggressive TCNOT scheduling through three complementary techniques. Hybrid Window Decoding assigns different decoders for each decoding window according to its DEM structure. DEM Stitch generates schedule-specific window-level DEMs just in time by assembling reusable precompiled fragments. Sub-window Parallel Decoding decomposes large windows into smaller sub-windows with graph-coloring formulation. Building on these techniques, PACE then performs decoder-aware scheduling to maximize TCNOT concurrency within the available decoder resources. Our evaluation shows the trade-off between quantum acceleration and classical decoding cost, revealing the limitations of current decoding systems for TCNOT-based FTQC.

quant-ph↗

$N$-Party Hadamard Test for Distributed Quantum Computation

Quantum computers promise computational advantages over classical computers, but hardware-imposed limitations remain a major obstacle. The Hadamard test mitigates these limitations by estimating expectation values associated with resource-intensive quantum operations using simple quantum circuits at the cost of additional classical sampling, and therefore underlies many quantum algorithms. However, in distributed quantum computing (DQC), which offers a promising route to scalability, its use is hindered by the need for nonlocal controlled operations. Here we introduce an $N$-party Hadamard test for DQC that estimates the same expectation values as the standard Hadamard test without implementing nonlocal controlled operations. The protocol instead uses pre-shared entanglement together with local operations and classical communication, which are standard resources in DQC settings. To demonstrate its utility, we apply it to unitary operations for clustered Hamiltonian simulation and to projectors for stabilizer-state preparation, showing lower sampling overheads than previous approaches by exploiting pre-shared entangled ancilla states. Moreover, we numerically demonstrate Bell-state preparation from Werner states to show favorable sampling efficiency and noise robustness relative to conventional purification, circuit knitting/cutting, and probabilistic error cancellation. Our work provides a general strategy for bringing Hadamard-test-based algorithms to DQC, facilitating practical and flexible quantum computation.

quant-ph↗

No More Hooks in the Surface Code: Distance-Preserving Syndrome Extraction for Arbitrary Layouts at Minimum Depth

Hook errors are a major challenge in implementing logical operations with the surface code, because they can reduce the fault distance below the code distance. This motivates syndrome-extraction circuits that suppress hook-error effects for the stabilizer layouts that appear during logical operations. However, the existing methods either increase circuit depth or require simultaneous execution of measurements and CNOT gates, both of which introduce additional overheads and degrade the threshold. We propose the ZX interleaving syndrome extraction, which preserves the full fault distance $d$ for any surface-code layout with regular stabilizer tiles at minimum depth, i.e., four layers of CNOT gates, without requiring additional circuit depth or simultaneous execution of measurements and CNOT gates. The key idea is to interleave the Z and X stabilizer tiles so that hook-error edges in the decoding graph are shortened and effectively eliminated. Numerical simulations under uniform depolarizing noise for memory and lattice-surgery experiments confirm that the proposed method achieves a full fault distance of $d$, whereas the best existing minimum-depth approach achieves $d-1$. Since the full fault distance is achievable for any regular tiling layout of the surface code, the proposed method may serve as an indispensable technique for practical fault-tolerant quantum computation.

quant-ph↗

Bounded-depth spacetime lattice surgery for resource-efficient fault-tolerant quantum computation

Fault-tolerant quantum computing based on lattice surgery requires place-and-route compilation with low spacetime overhead. Routing, in particular, faces a basic tension between suppressing path conflicts through greater spatial allocation and exploiting the time direction to realize ancilla-efficient spacetime routing. Existing approaches do not fully resolve this trade-off while retaining compatibility with inner factory layouts and termination guarantees. Here we introduce double-slice routing, a constant-depth spacetime-routing method that uses two consecutive time slices with a guarantee that its kink-parity correction terminates under both planar and stacked architectures. We numerically benchmark the resulting compiler on Hamiltonian-simulation workloads to show that double-slice routing reduces compilation cost by up to a factor of 2.4 over a single-slice baseline. Compared to projective routing, an existing method that allows an unbounded number of time slices per path, double-slice routing achieves smaller circuit volume with only a marginal execution-time penalty. Combined with a cultivation-compatible mapping optimization, the overall improvement reaches up to 7.5-fold over a naive single-slice compilation baseline. These results identify double-slice routing as a practically useful operating point in lattice-surgery compilation and show the substantial benefit in joint optimization of mapping and routing.

quant-ph↗

QuBE/Qubex: an integrated hardware-software system for superconducting qubit experiments with broadband control

Achieving high-fidelity operation in large-scale superconducting qubit systems requires not only control hardware with broad frequency coverage, low crosstalk, and tight synchronization but also software that coordinates system configuration, experiment execution, and data analysis. Here we present an integrated qubit-control system that combines broadband microwave hardware with a pulse-level software stack for scalable superconducting qubit experiments. The hardware provides broadband microwave coverage, including an instantaneous span of up to 1.6 GHz from a control output, while the software reduces setup and calibration overhead through automated configuration and built-in experiment workflows. We validate the system on a 64-qubit fixed-frequency transmon chip through full-chip frequency identification and representative demonstrations, including multi-unit far-detuned cross-resonance calibration and benchmarking that yields a measured two-qubit gate fidelity of 98.34%, and multilevel readout beyond the computational subspace. By disclosing the hardware architecture and releasing the software stack as open source, this work provides an inspectable hardware-software foundation for scalable superconducting qubit control experiments.

quant-ph↗

Quantum Magic in early FTQC: From Diagonal Clifford Hierarchy No-Go Theorems to Architecture Design Blueprints

We address the circuit-design problem of maximizing quantum magic in early fault-tolerant quantum computing (early FTQC), where logical dynamics natively take the form of alternating Clifford layers and diagonal non-Clifford layers. To render this optimization analytically tractable, we first prove a uniqueness theorem: for operational magic functionals built from Pauli expectation values, the axioms of faithfulness and tensor-product additivity force a Rényi-type dependence on the Pauli-spectrum. Leveraging the closed phase-polynomial description of the diagonal Clifford hierarchy, we derive exact Pauli-spectrum expressions and tight bounds for a shallow-layer model. These bounds expose a zero-magic mechanism and prove that maximal magic strictly requires graph-state preconditioning. Consequently, we establish our first no-go theorem: hierarchy level alone cannot universally order operational magic. Extending our framework to the $N$-layer model motivated by the Space-Time Efficient Analog Rotation (STAR) architecture, we obtain an exact iterative update rule for the Pauli spectrum. This yields a second no-go theorem: no state-independent sequence of operations can guarantee monotonic magic improvement. Together, these theorems demonstrate that algebraic gate structures are fundamentally insufficient to dictate resource generation. To overcome this, we reframe early FTQC gate selection as a state-aware, differentiable optimization over continuous analog parameters. Finally, we identify a severe kinematic expressibility bottleneck in architectures restricted to single-qubit $Z$-rotations and show that introducing nonlinear diagonal phases, such as multi-qubit $Z$-rotation, shatters this bottleneck. This provides a fundamental principle for demonstrating early FTQC, establishing scalable magic generation as a foundational benchmark for evaluating early FTQC architectures.

quant-ph↗

Accelerating BP-based decoders for QLDPC Codes with Local Syndrome-Based Preprocessing

Due to the high error rate of qubits, detecting and correcting errors is essential for achieving fault-tolerant quantum computing (FTQC). Quantum low-density parity-check (QLDPC) codes are one of the most promising quantum error correction (QEC) methods due to their high encoding rates. BP (Belief Propagation)-based decoders are widely used and highly competitive for QLDPC codes because BP offers inherent parallelism and strong scalability. However, BP-based decoders still suffer from high decoding latency, a large portion of which is spent in the iterative BP stage. In this paper, we propose a lightweight preprocessing step that utilizes local patterns in the syndrome to detect likely trivial error events and provide them as hints to BP-based decoders. These hints accelerate BP convergence and thereby reduce the overall decoding time. The proposed preprocessing step offers a broadly compatible approach to reducing the latency of BP-based QLDPC decodes. On the bivariate bicycle code $[[144,12,12]]$ at low physical error rates, our method achieves a $10\times$ speedup in decoding time for BP-OSD, and more than $2\times$ speedup for both BP-LSD and Relay-BP. Our method maintains the logical error rate when combined with BP-OSD and Relay-BP, while further achieving a significant reduction in logical error rate when combined with BP-LSD.

quant-ph↗

A $\boldsymbol{2d \times d \times d}$ Spacetime Volume Implementation of a Logical S Gate in the Surface Code

The logical S gate implemented via twist defect braiding in the surface code is one of the major sources of overhead in fault-tolerant quantum computing, since an S-gate correction is required in every logical T-gate teleportation. Existing logical S-gate implementations require spacetime volumes of \(2d \times 2d \times d\) or \(2d \times 1.5d \times d\), where $d$ is the code distance of the surface code. To the best of our knowledge, their circuit-level implementations have not yet been shown, hindering quantitative comparisons of fault distances and logical error rates. In this work, we provide these missing circuit-level implementations. Additionally, we propose a novel twist defect braiding protocol that reduces the spacetime volume to \(2d \times d \times d\). First, we construct an implementation of the proposed method using constant-length non-local gates, and then refine it to utilize only nearest-neighbor two-qubit gates on a square grid, without requiring additional two-qubit gate depth beyond that of standard syndrome extraction circuits. Through numerical simulations, we evaluate the fault distances and logical error rates for both existing and proposed methods. Our results show that, although the proposed method reduces the fault distance by one or three, its logical error rates remain comparable to those of existing methods at large code distances (\(d \ge 5\)) and at physical error rates near \(p = 10^{-3}\). This demonstrates that the proposed method is promising for near-term fault-tolerant quantum computing.

quant-ph↗

Design automation and space-time reduction for surface-code logical operations using a SAT-based EDA kernel compatible with general encodings

Fault-tolerant quantum computers (FTQCs) based on surface codes and lattice surgery have been widely studied, and there is strong demand for a framework that can identify logical operations with low space-time cost, verify their functionality and fault tolerance, and demonstrate their optimality within a given search space, much like electronic design automation (EDA) in classical circuit design. In this paper, we propose KOVAL-Q, an EDA kernel that verifies and optimizes surface-code logical operations by formulating them as a satisfiability (SAT) problem. Compared with existing SAT-based frameworks such as LaSsynth, our method can handle logical qubits with more flexible surface-code encodings, both as target configurations and as intermediate states. This extension enables the optimization of advanced layouts, such as fast blocks, and broadens the search space for logical operations. We demonstrate that KOVAL-Q can determine the minimum execution time of fundamental logical operations in given spatial layouts, such as $d$-cycle logical CNOTs and $2d$-cycle patch rotations. Their use reduces the execution time of widely studied FTQC applications by about 10% under a simplified scheduling model. KOVAL-Q consists of three subkernels corresponding to different types of constraints, which facilitates its integration as a submodule into scalable heuristic frameworks. Thus, our proposal provides an essential framework for optimizing and validating core FTQC subroutines.

quant-ph↗

Efficient and high-performance routing of lattice-surgery paths on three-dimensional lattice

Encoding logical qubits with surface codes and performing multi-qubit logical operations with lattice surgery is one of the most promising approaches to demonstrate fault-tolerant quantum computing. Thus, a method to efficiently schedule a sequence of lattice-surgery operations is vital for high-performance fault-tolerant quantum computing. A possible strategy to improve the throughput of lattice-surgery operations is splitting a large instruction into several small instructions, such as Bell state preparation and measurements, and executing a part of them in advance. However, scheduling methods to fully utilize this idea have yet to be explored. In this paper, we propose a fast and high-performance scheduling algorithm for lattice-surgery instructions leveraging this strategy. We achieved this by converting the scheduling problem of lattice-surgery instructions to a graph problem of embedding 3D paths into a 3D lattice, which enables us to explore efficient scheduling by solving path search problems in the 3D lattice. Based on this reduction, we propose a method to solve the path-finding problems, the look-ahead Dijkstra projection. We numerically show that this method reduced the execution time of benchmark programs generated from quantum phase estimation algorithms by 3.8 times compared with a naive method based on greedy algorithms. Our study establishes the relation between the lattice-surgery scheduling and graph search problems, which leads to further theoretical analysis on compiler optimization of fault-tolerant quantum computing.

quant-ph↗

Logical entanglement distribution between distant 2D array qubits

Sharing logical entangled pairs between distant quantum nodes is a key process to achieve fault tolerant quantum computation and communication. However, there is a gap between current experimental specifications and theoretical requirements for sharing logical entangled states while improving experimental techniques. Here, we propose an efficient logical entanglement distribution protocol based on surface codes for two distant 2D qubit array with nearest-neighbor interaction. A notable feature of our protocol is that it allows post-selection according to error estimations, which provides the tunability between the infidelity of logical entanglements and the success probability of the protocol. With this feature, the fidelity of encoded logical entangled states can be improved by sacrificing success rates. We numerically evaluated the performance of our protocol and the trade-off relationship, and found that our protocol enables us to prepare logical entangled states while improving fidelity in feasible experimental parameters. We also discuss a possible physical implementation using neutral atom arrays to show the feasibility of our protocol.

quant-ph↗

Efficient tomography of microwave photonic cluster states

Entanglement among a large number of qubits is a crucial resource for many quantum algorithms. Such many-body states have been efficiently generated by entangling a chain of itinerant photonic qubits in the optical or microwave domain. However, it has remained challenging to fully characterize the generated many-body states by experimentally reconstructing their exponentially large density matrices. Here, we develop an efficient tomography method based on the matrix-product-operator formalism and demonstrate it on a cluster state of up to 35 microwave photonic qubits by reconstructing its $2^{35} \times 2^{35}$ density matrix. The full characterization enables us to detect the performance degradation of our photon source which occurs only when generating a large cluster state. This tomography method is generally applicable to various physical realizations of entangled qubits and provides an efficient benchmarking method for guiding the development of high-fidelity sources of entangled photons.

quant-ph↗

Trade-offs between Quantum and Classical Resources in the Linear Combination of Unitaries

The randomized linear combination of unitaries (LCU) method with many applications to early fault-tolerant quantum computing algorithms has been proposed. This quantum algorithm computes the same expectation values as the original, fully coherent LCU algorithm using a shallower quantum circuit with a single ancilla qubit, at the cost of a quadratically larger sampling overhead. In this work, we propose a quantum algorithm intermediate between the original and randomized LCU that manages the trade-off between the sampling overhead and circuit complexity. Our algorithm divides the set of unitary operators into several groups and then randomly samples LCU circuits from these groups to evaluate the target expectation value. Notably, we reveal that across all grouping strategies, the mechanism of the sampling overhead reduction can be solely characterized by a metric we call the reduction factor. Moreover, we analytically prove an underlying monotonicity of the reduction factor in the group size: larger group sizes entail smaller sampling overhead. Finally, our framework enables a more flexible algorithmic design by systematically yielding intermediate implementations of LCU-based algorithms; we provide intermediate implementations of non-Hermitian dynamics simulation, ground-state property estimation, and quantum error detection. Besides, we demonstrate this principle by deriving intermediate trade-off scaling in sample complexity and ancillary space for quantum linear system solver.

quant-ph↗

Addressing requirements for crosstalk-free quantum-gate operation in many-body nanofiber cavity QED systems

A distributed network architecture in which flying photons connect individual modules containing stationary atomic qubits is a promising approach for scaling up neutral-atom based quantum-computing platforms. We consider an all-fiber based platform consisting of nanofiber cavity QED systems interconnected via conventional optical fibers. Each nanofiber cavity is strongly coupled to multiple atoms through its evanescent field, and atom pairs within one cavity (local) or two distant cavities (remote) are addressed for performing photon-mediated quantum logic gates on them by controlling the effective light-matter coupling via local AC Stark shifts and atom-fiber distance. We numerically evaluate the required parameters for achieving nearly crosstalk-free gate operation using these targeting methods by calculating average gate fidelities, success probabilities, and Pauli error rates for both local and remote controlled-Z gates. For the case of perfect addressing, we also analytically determine the theoretical optimum gate performance as limited by cavity reflectivity, cooperativity, and qubit level-splitting.

quant-ph↗

Network-Based Quantum Computing: an efficient design framework for many-small-node distributed fault-tolerant quantum computing

In fault-tolerant quantum computing, a large number of physical qubits are required to construct a single logical qubit, and a single quantum node may be able to hold only a small number of logical qubits. In such a case, the idea of distributed fault-tolerant quantum computing (DFTQC) is important to demonstrate large-scale quantum computation using small-scale nodes. However, the design of distributed systems on small-scale nodes, where each node can store only one or a few logical qubits for computation, has not been explored well yet. In this paper, we propose network-based quantum computation (NBQC) to efficiently realize distributed fault-tolerant quantum computation using many small-scale nodes. A key idea of NBQC is to let computational data continuously move throughout the network while maintaining the connectivity to other nodes. We numerically show that, for practical benchmark tasks, our method achieves shorter execution times than circuit-based strategies and more node-efficient constructions than measurement-based quantum computing. Also, if we are allowed to specialize the network to the structure of quantum programs, such as peak access frequencies, the number of nodes can be significantly reduced. Thus, our methods provide a foundation in designing DFTQC architecture exploiting the redundancy of many small fault-tolerant nodes.

quant-ph↗

Unitary-transformed projective squeezing: applications for circuit-knitting and state-preparation of non-Gaussian states

Continuous-variable (CV) quantum computing is a promising candidate for quantum computation because it can, even with one mode, utilize infinite-dimensional Hilbert spaces and can efficiently handle continuous values. Although photonic platforms have been considered as a leading platform for CV computation, hybrid systems that use both qubits and bosonic modes, e.g., superconducting hardware, have shown significant advances because they can prepare non-Gaussian states by utilizing the nonlinear interaction between the qubits and the bosonic modes. However, the size of hybrid hardware is currently restricted. Moreover, the fidelity of the non-Gaussian state is also restricted. This work extends the projective squeezing method to establish a formalism for projecting quantum states onto the states that are unitary-transformed from the squeezed vacuum at the expense of the sampling cost. Based on this formalism, we propose methods for simulating larger quantum devices and projecting states onto the cubic phase state, a typical non-Gaussian state, with a higher squeezing level and higher nonlinearity. To make implementation practical, we can, by leveraging the interactions in hybrid systems of qubits and bosonic modes, apply the smeared projector by using either the linear-combination-of-unitaries or virtual quantum error detection algorithms. We numerically verify the performance of our methods and show that projection can suppress the effect of photon-loss errors.

quant-ph↗

Selective Excitation of Superconducting Qubits with a Shared Control Line through Pulse Shaping

In conventional architectures of superconducting quantum computers, each qubit is connected to its own control line, leading to a commensurate increase in the number of microwave lines as the system scales. Frequency-multiplexed qubit control addresses this problem by enabling multiple qubits to share a single microwave line. However, it can cause unwanted excitation of non-target qubits, especially when the detuning between qubits is smaller than the pulse bandwidth. Here, we propose a selective-excitation-pulse (SEP) technique that suppresses unwanted excitations by shaping a drive pulse to create null points at non-target qubit frequencies. In a proof-of-concept experiment with three fixed-frequency transmon qubits, we demonstrate that the SEP technique achieves single-qubit gate fidelities comparable to those obtained with conventional Gaussian pulses while effectively suppressing unwanted excitations in non-target qubits. These results highlight the SEP technique as a promising tool for enhancing frequency-multiplexed qubit control.

quant-ph↗