SearcharxivSearch

arXiv subjects

Linghang Kong

Publications and source records attributed to Linghang Kong.

At least 19 recordsLinked to original sources

Bunny Codes: Broadening Superconducting Quantum Error Correction Capability through Advanced Control Engineering

Drawing on advances in superconducting qubit control schemes that unlock enriched native gate sets at the hardware level, we systematically examine how harnessing this enlarged physical two-qubit gate pool---specifically CNOT and CXSWAP---streamlines syndrome extraction for certain qLDPC codes with nonlocal stabilizers. Through an exhaustive search, we discover a set of qLDPC codes with various stabilizer weights and distances that can be implemented on the two-dimensional nearest-neighbor qubit connectivity native to superconducting hardware while achieving performance equivalent to that of the direct CNOT implementation requiring long-range interactions. We refer to those codes as Bunny codes. Across all code distances we examine, the best Bunny codes with weight-6 stabilizers in periodic boundary conditions have a code rate approximately $3\times$ that of the toric code; when converted to open boundary conditions, they retain an approximately $2\times$ code rate advantage over the rotated surface code. In circuit-level simulation, we find that some Bunny codes exhibit logical error rates an order of magnitude lower than toric codes with comparable code rates. Our results demonstrate that high-performance quantum error correction can be achieved using an expanded gate set rather than long-range couplers, thereby significantly reducing hardware complexity.

quant-ph

Learning to Decode in Parallel: Self-Coordinating Neural Network for Real-Time Quantum Error Correction

Fast, reliable decoders are pivotal components for enabling fault-tolerant quantum computation (FTQC). Neural network decoders like AlphaQubit have demonstrated potential, achieving higher accuracy than traditional human-designed decoding algorithms. However, existing implementations of neural network decoders lack the parallelism required to decode the syndrome stream generated by a superconducting logical qubit in real time. Moreover, integrating AlphaQubit with sliding window-based parallel decoding schemes presents non-trivial challenges: AlphaQubit is trained solely to output a single bit corresponding to the global logical correction for an entire memory experiment, rather than local physical corrections that can be easily integrated. We address this issue by training a recurrent, transformer-based neural network specifically tailored for parallel window decoding. While it still outputs a single bit, we derive training labels from a consistent set of local corrections and train on various types of decoding windows simultaneously. This approach enables the network to self-coordinate across neighboring windows, facilitating high-accuracy parallel decoding of arbitrarily long memory experiments. As a result, we overcome the throughput bottleneck that previously precluded the use of AlphaQubit-type decoders in FTQC. Our work presents the first scalable, neural-network-based parallel decoding framework that simultaneously achieves SOTA accuracy and the stringent throughput required for real-time quantum error correction. Using an end-to-end experimental workflow, we benchmark our decoder on the Zuchongzhi 3.2 superconducting quantum processor on surface codes with distances up to 7, demonstrating its superior accuracy. Moreover, we demonstrate that, using our approach, a single TPU v6e is capable of decoding surface codes with distances up to 25 within 1us per decoding round.

quant-ph

Quantum Design Automation: Foundations, Challenges, and the Road Ahead

Quantum computing is transitioning from laboratory research to industrial deployment, yet significant challenges persist: system scalability and performance, fabrication yields, and the advancement of algorithms and applications. We emphasize that in building quantum computers -- spanning quantum chips, system integration, instruction sets, algorithms, and middleware such as quantum error correction schemes -- design is everywhere. In this paper, we advocate for a holistic design perspective in quantum computing, a perspective we argue is pivotal to unlocking innovative co-design opportunities and addressing the aforementioned key challenges. To equip readers with sufficient background for exploring co-optimization opportunities, we detail how interconnected computational methods and tools collaborate to enable end-to-end quantum computer design. This coverage encompasses critical stages -- such as chip layout design automation, high-fidelity system-level simulation, Hamiltonian derivation for quantum system modeling, control pulse simulation, decoherence analysis, and physical verification and testing -- followed by quantum instruction set design. We then proceed to quantum system and software development, including quantum circuit synthesis, quantum error correction and fault tolerance, and logic verification and testing. Through these discussions, we illustrate with concrete examples -- including co-optimizing quantum instruction sets with algorithmic considerations, customizing error correction circuits to hardware-specific constraints, and streamlining quantum chip design through tailored code design, among others. We hope that the detailed end-to-end design workflow as well as these examples will foster dialogue between the hardware and software communities, ultimately facilitating the translation of meaningful research findings into future quantum hardware implementations.

quant-ph

Learning Neural Decoding with Parallelism and Self-Coordination for Quantum Error Correction

Fast, reliable decoders are pivotal components for enabling fault-tolerant quantum computation. Neural network decoders like AlphaQubit have demonstrated significant potential, achieving higher accuracy than traditional human-designed decoding algorithms. However, existing implementations of neural network decoders lack the parallelism required to decode the syndrome stream generated by a superconducting logical qubit in real time. Moreover, integrating AlphaQubit with sliding window-based parallel decoding schemes presents non-trivial challenges: AlphaQubit is trained solely to output a single bit corresponding to the global logical correction for an entire memory experiment, rather than local physical corrections that can be easily integrated. We address this issue by training a recurrent, transformer-based neural network specifically tailored for sliding-window decoding. While our network still outputs a single bit per window, we derive training labels from a consistent set of local corrections and train on various types of decoding windows simultaneously. This approach enables the network to self-coordinate across neighboring windows, facilitating high-accuracy parallel decoding of arbitrarily long memory experiments. As a result, we resolve the throughput limitation that previously prohibited the application of AlphaQubit-type decoders in fault-tolerant quantum computation.

quant-ph

LATTE: A Decoding Architecture for Quantum Computing with Temporal and Spatial Scalability

Quantum error correction allows inherently noisy quantum devices to emulate an ideal quantum computer with reasonable resource overhead. As a crucial component, decoding architectures have received significant attention recently. In this paper, we introduce LATTE, a FPGA-CPU hybrid decoding architecture aiming to address the key requirements of scaling up in lattice surgery quantum computation -- Latency, Accuracy, Throughput and Transmission Bandwidth, in an Eclectic manner. LATTE follows a hierarchical design: (1) A fully streaming and asynchronous block decoding system on CPU to enable parallelization both temporally and spatially. (2) A super-light yet accurate neural local decoding unit integrated with quantum control hardware on FPGA, which remains \emph{transparent} to the block decoding system, effectively reducing transmission bandwidth and accelerating the decoding process. LATTE delivers accuracy on par with the base decoder while achieving real-time decoding throughput and significantly reducing both bandwidth requirements and computational resources, enabling a level of scalability far beyond previous approaches. Under circuit-level noise $p=0.001$, LATTE achieves over $\mathbf{90\%}$ reduction in transmission bandwidth and a $\mathbf{6.4\times}$ speedup on average in single-block decoding. In the \emph{streaming decoding} scenario: (1) LATTE achieves constant and low latency ($\mathbf{16\times}$-$\mathbf{20\times}$ speedup over existing streaming decoding implementations) in arbitrarily long quantum memory experiments, with near-optimal resources -- merely $\mathbf{2}$ threads are sufficient for decoding the surface code with distance up to $17$. (2) LATTE minimizes latency in multi-patch measurement experiments through highly parallelized decoding operations. These combined efforts ensure sufficient scalability for large-scale fault-tolerant quantum computing.

quant-ph

Louvre: Relaxing Hardware Requirements of Quantum LDPC Codes by Routing with Expanded Quantum Instruction Set

Generalized bicycle codes (GB codes) represent a promising family of quantum low-density parity-check codes, characterized by high code rates and relatively local qubit connectivity. A subclass of the GB code called bivariate bicycle codes (BB codes) has garnered significant interest due to their compatibility with two-layer connectivity architectures on superconducting quantum processors. However, one key limitation of BB codes is their high qubit connectivity degree requirements (degree 6), which exacerbates the noise susceptibility of the system. Building on the recent progress in implementing multiple two-qubit gates on a single chip, this work introduces Louvre -- a routing-based framework designed to reduce qubit connectivity requirements in GB codes. Specifically, Louvre-7 achieves degree reduction while preserving the depth of the syndrome extraction circuit, whereas Louvre-8 further minimizes the connectivity by slightly increasing the circuit depth. When applied to BB codes, these two schemes could reduce the average degree to 4.5 and 4, respectively. Crucially, Louvre eliminates some of the long-range, error-prone connections, which is a distinct advantage over prior approaches. Numerical simulations demonstrate that Louvre-7 has an indistinguishable logical error rate as the standard syndrome extraction circuits of GB codes, while Louvre-8 only incurs a slight error rate penalty. Furthermore, by reordering some of the gates in the circuit, we can reduce the coupler length without degrading the performance. Though most of our analysis focuses on GB codes defined on periodic boundary conditions, we further discuss the adaptability of Louvre to open-boundary lattices and defect-containing grids, underscoring its broader applicability in practical quantum error correction architectures.

quant-ph

Benchmarking fault-tolerant quantum computing hardware via QLOPS

It is widely recognized that quantum computing has profound impacts on multiple fields, including but not limited to cryptography, machine learning, materials science, etc. To run quantum algorithms, it is essential to develop scalable quantum hardware with low noise levels and to design efficient fault-tolerant quantum computing (FTQC) schemes. Currently, various FTQC schemes have been developed for different hardware platforms. However, a comprehensive framework for the analysis and evaluation of these schemes is still lacking. In this work, we propose Quantum Logical Operations Per Second (QLOPS) as a metric for assessing the performance of FTQC schemes on quantum hardware platforms. This benchmarking framework will integrate essential relevant factors, e.g., the code rates of quantum error-correcting codes, the accuracy, throughput, and latency of the decoder. Through a resource analysis of factoring RSA-2048, we demonstrate that QLOPS reflects the practical requirements of quantum algorithm execution. This framework will enable the identification of bottlenecks in quantum hardware, providing potential directions for their development. Moreover, our results will help establish a comparative framework for evaluating FTQC designs. As this benchmarking approach considers practical applications, it may assist in estimating the hardware resources needed to implement quantum algorithms and offers preliminary insights into potential timelines.

quant-ph

Routing-based technique for defect mitigation in quantum error correction

As quantum chips scale up for large-scale computation, hardware defects become inevitable and must be carefully addressed. In this work, we introduce Halma, a defect mitigation technique empowered by an expanded native gate set that incorporates the iSWAP gate alongside the conventional CNOT gate. Halma emerges as a supplementary technique within the defect mitigation toolbox, offering effective mitigation of ancilla qubit defects encountered during surface code stabilizer measurements while maintaining compatibility with existing superstabilizer-based methodologies. Halma introduces zero reduction in the spacelike distance of the code without further sacrifice to the timelike distance. Numerical simulation suggests that in comparison to previous methods, Halma could provide an order of magnitude improvement in the average logical error rate under realistic experimental settings, leading to a $\sim3\times$ reduction in the footprint of a teraquop. These results clearly demonstrate the capability of Halma in easing the near-term realization of fault-tolerant quantum computing on hardware with fabrication defects, and exemplifies how leveraging intrinsic hardware capabilities can enhance quantum hardware performance.

quant-ph

Convergence efficiency of quantum gates and circuits

We consider quantum circuit models where the gates are drawn from arbitrary gate ensembles given by probabilistic distributions over certain gate sets and circuit architectures, which we call stochastic quantum circuits. Of main interest in this work is the speed of convergence of stochastic circuits with different gate ensembles and circuit architectures to unitary t-designs. A key motivation for this theory is the varying preference for different gates and circuit architectures in different practical scenarios. In particular, it provides a versatile framework for devising efficient circuits for implementing $t$-designs and relevant applications including random circuit and scrambling experiments, as well as benchmarking the performance of gates and circuit architectures. We examine various important settings in depth. A key aspect of our study is an "ironed gadget" model, which allows us to systematically evaluate and compare the convergence efficiency of entangling gates and circuit architectures. Particularly notable results include i) gadgets of two-qubit gates with KAK coefficients $\left(\fracπ{4}-\frac{1}{8}\arccos(\frac{1}{5}),\fracπ{8},\frac{1}{8}\arccos(\frac{1}{5})\right)$ (which we call $χ$ gates) directly form exact 2- and 3-designs; ii) the iSWAP gate family achieves the best efficiency for convergence to 2-designs under mild conjectures with numerical evidence, even outperforming the Haar-random gate, for generic many-body circuits; iii) iSWAP + complete graph achieve the best efficiency for convergence to 2-designs among all graph circuits. A variety of numerical results are provided to complement our analysis. We also derive robustness guarantees for our analysis against gate perturbations. Additionally, we provide cursory analysis on gates with higher locality and found that the Margolus gate outperforms various other well-known gates.

quant-ph

TEAFormers: TEnsor-Augmented Transformers for Multi-Dimensional Time Series Forecasting

Multi-dimensional time series data, such as matrix and tensor-variate time series, are increasingly prevalent in fields such as economics, finance, and climate science. Traditional Transformer models, though adept with sequential data, do not effectively preserve these multi-dimensional structures, as their internal operations in effect flatten multi-dimensional observations into vectors, thereby losing critical multi-dimensional relationships and patterns. To address this, we introduce the Tensor-Augmented Transformer (TEAFormer), a novel method that incorporates tensor expansion and compression within the Transformer framework to maintain and leverage the inherent multi-dimensional structures, thus reducing computational costs and improving prediction accuracy. The core feature of the TEAFormer, the Tensor-Augmentation (TEA) module, utilizes tensor expansion to enhance multi-view feature learning and tensor compression for efficient information aggregation and reduced computational load. The TEA module is not just a specific model architecture but a versatile component that is highly compatible with the attention mechanism and the encoder-decoder structure of Transformers, making it adaptable to existing Transformer architectures. Our comprehensive experiments, which integrate the TEA module into three popular time series Transformer models across three real-world benchmarks, show significant performance enhancements, highlighting the potential of TEAFormers for cutting-edge time series forecasting.

cs.LG

A Classical Architecture For Digital Quantum Computers

Scaling bottlenecks the making of digital quantum computers, posing challenges from both the quantum and the classical components. We present a classical architecture to cope with a comprehensive list of the latter challenges {\em all at once}, and implement it fully in an end-to-end system by integrating a multi-core RISC-V CPU with our in-house control electronics. Our architecture enables scalable, high-precision control of large quantum processors and accommodates evolving requirements of quantum hardware. A central feature is a microarchitecture executing quantum operations in parallel on arbitrary predefined qubit groups. Another key feature is a reconfigurable quantum instruction set that supports easy qubit re-grouping and instructions extensions. As a demonstration, we implement the widely-studied surface code quantum computing workflow, which is instructive for being demanding on both the controllers and the integrated classical computation. Our design, for the first time, reduces instruction issuing and transmission costs to constants, which do not scale with the number of qubits, without adding any overheads in decoding or dispatching. Rather than relying on specialized hardware for syndrome decoding, our system uses a dedicated multi-core CPU for both qubit control and classical computation, including syndrome decoding. This simplifies the system design and facilitates load-balancing between the quantum and classical components. We implement recent proposals as decoding firmware on a RISC-V system-on-chip (SoC) that parallelizes general inner decoders. By using our in-house Union-Find and PyMatching 2 implementations, we can achieve unprecedented decoding capabilities of up to distances 47 and 67 with the currently available SoCs, under realistic and optimistic assumptions of physical error rate $p=0.001 and p=0.0001, respectively, all in just 1 \textmu s.

quant-ph

Quantum Instruction Set Design for Performance

A quantum instruction set is where quantum hardware and software meet. We develop new characterization and compilation techniques for non-Clifford gates to accurately evaluate different quantum instruction set designs. We specifically apply them to our fluxonium processor that supports mainstream instruction $\mathrm{iSWAP}$ by calibrating and characterizing its square root $\mathrm{SQiSW}$. We measure a gate fidelity of up to $99.72\%$ with an average of $99.31\%$ and realize Haar random two-qubit gates using $\mathrm{SQiSW}$ with an average fidelity of $96.38\%$. This is an average error reduction of $41\%$ for the former and a $50\%$ reduction for the latter compared to using $\mathrm{iSWAP}$ on the same processor. This shows designing the quantum instruction set consisting of $\mathrm{SQiSW}$ and single-qubit gates on such platforms leads to a performance boost at almost no cost.

quant-ph

Linear Cross Entropy Benchmarking with Clifford Circuits

With the advent of quantum processors exceeding $100$ qubits and the high engineering complexities involved, there is a need for holistically benchmarking the processor to have quality assurance. Linear cross-entropy benchmarking (XEB) has been used extensively for systems with $50$ or more qubits but is fundamentally limited in scale due to the exponentially large computational resources required for classical simulation. In this work we propose conducting linear XEB with Clifford circuits, a scheme we call Clifford XEB. Since Clifford circuits can be simulated in polynomial time, Clifford XEB can be scaled to much larger systems. To validate this claim, we run numerical simulations for particular classes of Clifford circuits with noise and observe exponential decays. When noise levels are low, the decay rates are well-correlated with the noise of each cycle assuming a digital error model. We perform simulations of systems up to 1,225 qubits, where the classical processing task can be easily dealt with by a workstation. Furthermore, using the theoretical guarantees in Chen et al. (arXiv:2203.12703), we prove that Clifford XEB with our proposed Clifford circuits must yield exponential decays under a general error model for sufficiently low errors. Our theoretical results explain some of the phenomena observed in the simulations and shed light on the behavior of general linear XEB experiments.

quant-ph

Near-optimal covariant quantum error-correcting codes from random unitaries with symmetries

Quantum error correction and symmetries play central roles in quantum information science and physics. It is known that quantum error-correcting codes that obey (are covariant with respect to) continuous symmetries in a certain sense cannot correct erasure errors perfectly (a well-known result in this regard being the Eastin-Knill theorem in the context of fault-tolerant quantum computing), in contrast to the case without symmetry constraints. Furthermore, several quantitative fundamental limits on the accuracy of such covariant codes for approximate quantum error correction are known. Here, we consider the quantum error correction capability of uniformly random covariant codes. In particular, we analytically study the most essential cases of $U(1)$ and $SU(d)$ symmetries, and show that for both symmetry groups the error of the covariant codes generated by Haar-random symmetric unitaries, i.e., unitaries that commute with the group actions, typically scale as $O(n^{-1})$ in terms of both the average- and worst-case purified distances against erasure noise, saturating the fundamental limits to leading order. We note that the results hold for symmetric variants of unitary 2-designs, and comment on the convergence problem of symmetric random circuits. Our results not only indicate (potentially efficient) randomized constructions of optimal $U(1)$- and $SU(d)$-covariant codes, but also reveal fundamental properties of random symmetric unitaries, which yield important solvable models of complex quantum systems (including black holes and many-body spin systems) that have attracted great recent interest in quantum gravity and condensed matter physics. We expect our construction and analysis to find broad relevance in both physics and quantum computing.

quant-ph

A framework for randomized benchmarking over compact groups

Characterization of experimental systems is an essential step in developing and improving quantum hardware. A collection of protocols known as Randomized Benchmarking (RB) was developed in the past decade, which provides an efficient way to measure error rates in quantum systems. In a recent paper (arxiv:2010.07974), a general framework for RB was proposed, which encompassed most of the known RB protocols and overcame the limitation on error models in previous works. However, even this general framework has a restriction: it can only be applied to a finite group of gates. This does not meet the need posed by experiments, in particular the demand for benchmarking non-Clifford gates and continuous gate sets on quantum devices. In this work we generalize the RB framework to continuous groups of gates and show that as long as the noise level is reasonably small, the output can be approximated as a linear combination of matrix exponential decays. As an application, we numerically study the fully randomized benchmarking protocol (i.e. RB with the entire unitary group as the gate set) enabled by our proof. This provides a unified way to estimate the gate fidelity for any quantum gate in an experiment.

quant-ph

Charge-conserving unitaries typically generate optimal covariant quantum error-correcting codes

Quantum error correction and symmetries play central roles in quantum information science and physics. It is known that quantum error-correcting codes covariant with respect to continuous symmetries cannot correct erasure errors perfectly (an important case being the Eastin-Knill theorem), in contrast to the case without symmetry constraints. Furthermore, there are fundamental limits on the accuracy of such covariant codes for approximate quantum error correction. Here, we consider the quantum error correction capability of random covariant codes. In particular, we show that $U(1)$-covariant codes generated by Haar random $U(1)$-symmetric unitaries, i.e. unitaries that commute with the charge operator (or conserve the charge), typically saturate the fundamental limits to leading order in terms of both the average- and worst-case purified distances against erasure noise. We note that the results hold for symmetric variants of unitary 2-designs, and comment on the convergence problem of charge-conserving random circuits. Our results not only indicate (potentially efficient) randomized constructions of optimal $U(1)$-covariant codes, but also reveal fundamental properties of random charge-conserving unitaries, which may underlie important models of complex quantum systems in wide-ranging physical scenarios where conservation laws are present, such as black holes and many-body spin systems.

quant-ph

A Separation of Out-of-time-ordered Correlation and Entanglement

The out-of-time-ordered correlation (OTOC) and entanglement are two physically motivated and widely used probes of the "scrambling" of quantum information, a phenomenon that has drawn great interest recently in quantum gravity and many-body physics. We argue that the corresponding notions of scrambling can be fundamentally different, by proving an asymptotic separation between the time scales of the saturation of OTOC and that of entanglement entropy in a random quantum circuit model defined on graphs with a tight bottleneck, such as tree graphs. Our result counters the intuition that a random quantum circuit mixes in time proportional to the diameter of the underlying graph of interactions. It also provides a more rigorous justification for an argument in our previous work arXiv:1807.04363, that black holes may be slow information scramblers, which in turn relates to the black hole information problem. The bounds we obtained for OTOC are interesting in their own right in that they generalize previous studies of OTOC on lattices to the geometries on graphs in a rigorous and general fashion.

quant-ph

The performance of the quantum adiabatic algorithm on spike Hamiltonians

Perturbed Hamming weight problems serve as examples of optimization instances for which the adiabatic algorithm provably out performs classical simulated annealing. In this work we study the efficiency of the adiabatic algorithm for solving the "the Hamming weight with a spike" problem by using several methods to compute the scaling of the spectral gap at the critical point, which apply for various ranges of the height and width of the barrier. Our main result is a rigorous polynomial lower bound on the minimum spectral gap for the adiabatic evolution when the bit-symmetric cost function has a thin but polynomially high barrier. This is accomplished by the use of a variational argument with an improved ansatz for the ground state, along with a comparison to the spectrum of the system when no spike term is present. We also give a more detailed treatment of the spin coherent path-integral instanton method which was used by Farhi, Goldstone, and Gutmann in arXiv:quant-ph/0201031, and consider its applicability for estimating the gap for different scalings of barrier height and width. We adapt the discrete WKB method for an abruptly changing potential, and apply it to the construction of approximate wave functions which can be used to estimate the gap. Finally, the improved ansatz for the ground state leads to a method for predicting the location of avoided crossings in the excited states of the energy spectrum of the thin spike Hamiltonian, and we use a recursion relation to determine the ordering of some of these avoided crossings, which may be a useful step towards understanding the diabatic cascade phenomenon which occurs in spike Hamiltonians.

quant-ph