SearcharxivSearch

arXiv subjects

Siyuan Niu

Publications and source records attributed to Siyuan Niu.

At least 19 recordsLinked to original sources

AC/DC: Automated Compilation for Dynamic Circuits

Dynamic quantum circuits incorporate mid-circuit measurements (MCMs) and feed-forward operations are crucial for manipulating quantum information. They have been broadly used in quantum error correction and quantum teleportation. Recently, they are utilized to prepare certain states and long-range entangling gates as well as reduce resource overhead in quantum algorithms. In this paper, we present AC/DC, a novel Automated Compilation framework for generating Dynamic quantum Circuits that prepare any unitary operators or states, leveraging numerical optimization-based circuit synthesis methods. The first contribution is introducing optimization objective functions incorporating MCMs and feed-forward operations. The second contribution is embedding these into a popular open-source quantum circuit synthesis framework. We demonstrate generating dynamic circuits for long range entangling gates, circuit optimization, lattice simulations, and state preparation, with validation through simulation and quantum hardware. Furthermore, we perform a noise analysis to assess the impact of MCM and gate errors, identifying scenarios where dynamic circuits provide significant benefits. The dynamic circuits generated by our framework show substantial improvements in reducing circuit depth and, in some cases, the number of gates. To our knowledge, this is the first practical procedure to generate dynamic quantum circuits, paving the way for enhanced circuit generation and optimization methods for near-term quantum computers.

quant-ph

Scalable Benchmarking Framework for Dynamic Quantum Circuits

Dynamic quantum circuits with mid-circuit measurements (MCMs) and feed-forward operations play a crucial role in various applications, such as quantum error correction and quantum algorithms. With advancements in quantum hardware enabling the implementation of MCM and feed-forward loops, the use of dynamic circuits has become increasingly prevalent. There is a significant need for a benchmarking framework specially designed for dynamic circuits to capture their unique properties, as current benchmarking tools are designed primarily for unitary circuits and cannot be trivially extended to dynamic circuits. We propose dynamarq, a scalable and hardware-agnostic benchmarking framework for dynamic circuits. We collect a set of dynamic circuit benchmarks spanning various applications and propose a broad set of circuit features to characterize the structure of these dynamic circuits. We run them on two IBM quantum processors and the Quantinuum Helios-1E emulator, and propose scalable, application-dependent fidelity scores for each benchmark based on hardware execution results. We perform statistical modeling to identify correlations between circuit features and fidelity scores, and demonstrate highly accurate fidelity prediction using our model. Our model parameters are also transferable across hardware backends and calibration cycles. Our framework facilitates the understanding of dynamic circuit structures and provides insights for designing and optimizing dynamic circuits to achieve high execution fidelity on quantum hardware.

quant-ph

Parallel Circuit Execution for Scalable Quantum Computation

Today's quantum processors have tens to hundreds of physical qubits, but reliable execution of arbitrary circuits remains limited to fewer than 30 entangled qubits across hardware modalities. Building upon prior work, we introduce an error- and topology-aware method for mapping multiple independent circuits onto disjoint regions of a single large-scale QPU for parallel circuit execution. For applications with many similarly sized circuits, such as observable estimation for Hamiltonian simulation, this approach can reduce billed QPU execution time, with ideal speedup proportional to the number of usable partitions. We demonstrate the approach on IBM's 156-qubit ibm_boston processor using standard QED-C benchmark and Hamiltonian-based observable-estimation workloads. Compared with standard sequential execution, parallel execution reduces billed execution time by 3.5-5.5x while retaining 83-92% of the sequential fidelity. We further evaluate the scaling of parallel circuit execution using GPU-accelerated classical simulation, distributing measurement circuits across GPUs via MPI. Using CUDA-Q on the NERSC Perlmutter system, we achieve up to 13.8x speedup on 16 GPUs (86% parallel efficiency) for an H2 electronic-structure simulation, with scaling evaluated across multiple Hamiltonians and circuit counts. These results provide an indication of the performance ceiling that parallel execution on future quantum hardware may eventually approach. Both execution modes are implemented as a runtime option within the QED-C Application-Oriented Benchmark suite. Together, the results show that circuit-level parallelism can reduce execution cost on current quantum hardware and simulation time on GPU clusters, with the potential for greater benefits as device quality and qubit counts increase.

quant-ph

Challenges in Barren Plateau Mitigation with Dynamic Parameterized Quantum Circuits

Variational quantum algorithms (VQAs) are a promising paradigm for quantum advantage, yet their trainability is severely hampered by barren plateaus (BPs). Several recent works have proposed dynamic parameterized quantum circuits (DPQCs), which interleave unitary layers with parameterized CPTP maps, such as engineered dissipation, feedforward gadgets, and periodic resets, as a possible strategy for mitigating BPs. We unify this class of circuits into a formalization for DPQCs.We identify constraints on the nature and the structure of DPQCs if they are to prevent a significant number of parameters from becoming untrainable. Using purification and Pauli-path analysis, we further identify a mechanism by which the cost function can remain anti-concentrated even when many parameters remain untrainable. Our analysis reveals ways to design DPQCs that do not have an exponentially concentrated cost function, and our results suggest that BP mitigation via DPQCs is at least as hard as designing BP-free unitaries.

quant-ph

Performance Analysis of QAOA Across Distributed Quantum Network Topologies Using SwitchQNet

Quantum data-center (QDC) architectures aim to scale distributed quantum computing (DQC) by interconnecting multiple quantum processing units (QPUs), but their performance depends strongly on how algorithmic communication patterns interact with entanglement generation, switch reconfiguration, and network topology. This paper studies the Quantum Approximate Optimization Algorithm (QAOA) as a graph-structured optimization workload for QDC-based distributed quantum computing. We adapt QAOA to SwitchQNet, a distributed quantum compiler framework that schedules communication and entanglement generation over switch-based QDC networks, by adding a routing generator that converts graph-dependent two-qubit cost interactions into remote-CX communication requests across QPUs. Using this extension, we evaluate QAOA instances across Clos, fat-tree, and spine-leaf topologies, measuring communication latency, EPR-pair overhead, EPR wait time, retry overhead, and sensitivity to buffer size, look-ahead depth, communication-qubit count, EPR latency, and EPR fidelity assumptions. The results show that QAOA obtains modest but consistent latency reductions, highlighting its value as a diagnostic benchmark for studying the interaction between algorithm structure, entanglement management, and quantum-network architecture.

quant-ph

Towards Scalable Quaternary Message-Passing Decoding for Quantum Error Correction

The scalability and interpretability of message-passing (MP) decoding, such as (quaternary) Belief Propagation, remain open challenges in quantum error correction. Even for surface codes, arguably the first testbed for decoding methods, studies of improved MP decoders have mostly been restricted to small distances ($d \lesssim 19$). Moreover, the mismatch with established message-passing theory limits the decoder's interpretability, making it unclear whether MP decoding can sustain its effectiveness at large system sizes. This work takes a step toward a more principled and interpretable MP decoding framework, with the goal of making MP-based decoding more reliable and bridging theory and practice. We introduce a dilution method, which allows a quaternary Min-Sum (MS) decoder to exhibit an apparent depolarizing threshold of $16\%$ up to distance $20$, outperforming Minimum-Weight Perfect Matching in finite-length regimes. Notably, for $X$-noise, the standard MS decoder under dilution has worst-case complexity $O(N \log^2 d)$ and outperforms BP-OSD at $d=65$. The observed $\sim 9\%$ threshold may correspond to a true asymptotic threshold. Finally, we give a graph-dilution argument that interprets the success of the dilution method and offers insight into when MP algorithms can genuinely scale. Taken together, these results provide encouraging progress toward scalable and interpretable MP decoding in quantum error correction.

quant-ph

Error Mitigation in Dynamic Circuits for Hamiltonian Simulation

Dynamic quantum circuits integrate mid-circuit measurements and feed-forward operations to enable real-time classical processing and conditional quantum logic. These capabilities are central to key quantum protocols such as quantum error correction, and have recently demonstrated significant potential for reducing quantum resources, including circuit depth and gate count, across a range of applications. However, executing dynamic circuits on real quantum hardware introduces a critical trade-off: while resource requirements decrease, circuit fidelity degrades due to high error rates of mid-circuit measurements, as well as the decoherence errors accumulated during the extended idle periods introduced by both mid-circuit measurements and feed-forward operations. In this paper, we systematically investigate the impact of standard error mitigation techniques on dynamic circuit applications pertaining to Hamiltonian simulation and ground state estimation of physically relevant systems like the Heisenberg model. We explore dynamical decoupling (DD) as a strategy to suppress decoherence and crosstalk errors during idle windows introduced by mid-circuit measurements and feed-forward delays, and also examine error mitigation via zero-noise extrapolation (ZNE). Through experiments conducted on IBM quantum hardware, we benchmark effective combinations of these strategies that maximize the practical benefits of dynamic quantum circuits in these applications. We demonstrate that a combination of DD and ZNE is effective in mitigating the errors introduced during mid-circuit measurements and feed-forward operations, as well as the errors arising from faulty measurements. This approach yields a energy gap improvement of at least 60% in ground state estimation and reduces observed error of time-evolved states by up to 99% for the Ising model and up to 20% for the Heisenberg model.

quant-ph

Estimating The Energy Consumption of Quantum Computing from A Full System Aspect

Quantum computing promises disruptive capabilities, yet its energy footprint has received far less attention than its asymptotic speedups. We present a first-order, full-system energy model for quantum computing in an high performance computing (HPC) context. The model separates costs common to NISQ and FTQC, such as system maintenance and classical processing, from regime-specific ones such as error mitigation for NISQ and error correction for FTQC. We instantiate the model on 96- and 100-qubit Heisenberg time-evolution simulations on IBM Eagle r3 and a representative VQE workload, and sketch the FTQC energy pipeline. We find that NISQ energy is dominated by the QEM sampling multiplier, while FTQC cost shifts to physical-qubit overhead set by the code distance and magic states. Our model provides actionable insights into the energy consumption of both NISQ and FTQC workloads, and paves the way toward energy-efficient quantum advantage.

quant-ph

Metriq: A Collaborative Platform for Benchmarking Quantum Computers

The fragmented landscape of quantum computer benchmarks, characterized by system-specific tools and inconsistent evaluation methodologies, hinders reliable cross-platform performance assessment. We introduce Metriq, an open-source collaborative platform for reproducible cross-platform quantum benchmarking that integrates benchmark definition and execution, data collection, and public presentation into a unified workflow. The Metriq benchmark suite spans both system-level metrics that characterize fundamental device properties such as entanglement quality, gate performance, and circuit speed, as well as application-inspired protocols that assess performance on quantum machine learning, optimization, and quantum simulation tasks. Benchmarks are chosen to scale with processor size, and the framework incorporates cost and resource estimation to support practical evaluation. Using Metriq, we collect and publicly release results from more than ten quantum computers across multiple hardware vendors, enabling systematic cross-platform comparison. The resulting curated dataset also reveals the practical strengths and limitations of individual benchmarks, creating a feedback loop that informs the ongoing refinement of the suite. To summarize performance across the benchmark suite, we introduce the Metriq Score, a composite index aggregating benchmark outcomes. We further present cross-benchmark analyses enabled by the shared dataset and their correlations with hardware calibration metrics. Through open development and data sharing, Metriq provides a practical foundation for reproducible benchmarking of quantum computers as hardware and benchmarking methods continue to evolve.

quant-ph

Architectural Foundations for Checkpointing and Restoration in Quantum HPC Systems

In this work, we explore the design of the checkpointing and restoration for quantum HPC that leverages dynamic circuit technology to enable restartable and resilient quantum execution. Rather than attempting to checkpoint quantum states, our approach redefines checkpointing as a control flow and algorithmic state problem. By exploiting mid-circuit measurements, classical feed forward, and conditional execution supported by dynamic circuits, we capture sufficient program state to allow correct restoration of quantum workflows after interruption or failure. This design aligns naturally with iterative and staged quantum algorithms such as variational eigensolvers, quantum approximate optimization, and time-stepping methods commonly used in quantum simulation and scientific computing.

quant-ph

A Compilation Framework for Quantum Circuits with Mid-Circuit Measurement Error Awareness

Mid-circuit measurement (MCM) provides the capability for qubit reuse and dynamic control in quantum processors, enabling more resource-efficient algorithms and supporting error-correction procedures. However, MCM introduces several sources of error, including measurement-induced crosstalk, idling-qubit decoherence, and reset infidelity, and these errors exhibit pronounced qubit-dependent variability within a single device. Since existing compilers such as the Qiskit-compiler and QR-Map (the state-of-art qubit reuse compiler) do not account for this variability, circuits with frequent MCM operations often experience substantial fidelity loss. In thie paper, we propose MERA, a compilation framework that performs MCM-error-aware layout, routing, and scheduling. MERA leverages lightweight profiling to obtain a stable per-qubit MCM error distribution, which it uses to guide error-aware qubit mapping and SWAP insertions. To further mitigate MCM-related decoherence and crosstalk, MERA augments as-late-as-possible scheduling with context-aware dynamic decoupling. Evaluated on 27 benchmark circuits, MERA achieves 24.94% -- 52.00% fidelity improvement over the Qiskit compiler (optimization level 3) without introducing additional overhead. On QR-Map-generated circuits, it improves fidelity by 29.26% on average and up to 122.58% in the best case, demonstrating its effectiveness for dynamic circuits dominated by MCM operations.

quant-ph

Software for Creating Scalable Benchmarks from Quantum Algorithms

Creating scalable, reliable, and well-motivated benchmarks for quantum computers is challenging: straightforward approaches to benchmarking suffer from exponential scaling, are insensitive to important errors, or use poorly-motivated performance metrics. Furthermore, curated benchmarking suites cannot include every interesting quantum circuit or algorithm, which necessitates a tool that enables the easy creation of new benchmarks. In this work, we introduce a software tool for creating scalable and reliable benchmarks that measure a well-motivated performance metric (process fidelity) from user-chosen quantum circuits and algorithms. Our software, called $\texttt{scarab}$, enables the creation of efficient and robust benchmarks even from circuits containing thousands or millions of qubits, by employing efficient fidelity estimation techniques, including mirror circuit fidelity estimation and subcircuit volumetric benchmarking. $\texttt{scarab}$ provides a simple interface that enables the creation of reliable benchmarks by users who are not experts in the theory of quantum computer benchmarking or noise. We demonstrate the flexibility and power of $\texttt{scarab}$ by using it to turn existing inefficient benchmarks into efficient benchmarks, to create benchmarks that interrogate hardware and algorithmic trade-offs in Hamiltonian simulation, to quantify the in-situ efficacy of approximate circuit compilation, and to create benchmarks that use subcircuits to measure progress towards executing a circuit of interest.

quant-ph

Platform-Agnostic Modular Architecture for Quantum Benchmarking

We present a platform-agnostic modular architecture that addresses the increasingly fragmented landscape of quantum computing benchmarking by decoupling problem generation, circuit execution, and results analysis into independent, interoperable components. Supporting over 20 benchmark variants ranging from simple algorithmic tests like Bernstein-Vazirani to complex Hamiltonian simulation with observable calculations, the system integrates with multiple circuit generation APIs (Qiskit, CUDA-Q, Cirq) and enables diverse workflows. We validate the architecture through successful integration with Sandia's $\textit{pyGSTi}$ for advanced circuit analysis and CUDA-Q for multi-GPU HPC simulations. Extensibility of the system is demonstrated by implementing dynamic circuit variants of existing benchmarks and a new quantum reinforcement learning benchmark, which become readily available across multiple execution and analysis modes. Our primary contribution is identifying and formalizing modular interfaces that enable interoperability between incompatible benchmarking frameworks, demonstrating that standardized interfaces reduce ecosystem fragmentation while preserving optimization flexibility. This architecture has been developed as a key enhancement to the continually evolving QED-C Application-Oriented Performance Benchmarks for Quantum Computing suite.

quant-ph

A Practical Framework for Assessing the Performance of Observable Estimation in Quantum Simulation

Simulating dynamics of physical systems is a key application of quantum computing, with potential impact in fields such as condensed matter physics and quantum chemistry. However, current quantum algorithms for Hamiltonian simulation yield results that are inadequate for real use cases and suffer from lengthy execution times when implemented on near-term quantum hardware. In this work, we introduce a framework for evaluating the performance of quantum simulation algorithms, focusing on the computation of observables, such as energy expectation values. Our framework provides end-to-end demonstrations of algorithmic optimizations that utilize Pauli term groups based on k-commutativity, generate customized Clifford measurement circuits, and implement weighted shot distribution strategies across these groups. These demonstrations span multiple quantum execution environments, allowing us to identify critical factors influencing runtime and solution accuracy. We integrate enhancements into the QED-C Application-Oriented Benchmark suite, utilizing problem instances from the open-source HamLib collection. Our results demonstrate a 27.1% error reduction through Pauli grouping methods, with an additional 37.6% improvement from the optimized shot distribution strategy. Our framework provides an essential tool for advancing quantum simulation performance using algorithmic optimization techniques, enabling systematic evaluation of improvements that could maximize near-term quantum computers' capabilities and advance practical quantum utility as hardware evolves.

quant-ph

Multi-qubit Dynamical Decoupling for Enhanced Crosstalk Suppression

Dynamical decoupling (DD) is one of the simplest error suppression methods, aiming to enhance the coherence of qubits in open quantum systems. Moreover, DD has demonstrated effectiveness in reducing coherent crosstalk, one major error source in near-term quantum hardware, which manifests from two types of interactions. Static crosstalk exists in various hardware platforms, including superconductor and semiconductor qubits, by virtue of always-on qubit-qubit coupling. Additionally, driven crosstalk may occur as an unwanted drive term due to leakage from driven gates on other qubits. Here we explore a novel staggered DD protocol tailored for multi-qubit systems that suppresses the decoherence error and both types of coherent crosstalk. We develop two experimental setups -- an "idle-idle" experiment in which two pairs of qubits undergo free evolution simultaneously and a "driven-idle" experiment in which one pair is continuously driven during the free evolution of the other pair. These experiments are performed on an IBM Quantum superconducting processor and demonstrate the significant impact of the staggered DD protocol in suppressing both types of coherent crosstalk. When compared to the standard DD sequences from state-of-the-art methodologies with the application of X2 sequences, our staggered DD protocol enhances circuit fidelity by 19.7% and 8.5%, respectively, in addressing these two crosstalk types.

quant-ph

Coqa: Blazing Fast Compiler Optimizations for QAOA

The Quantum Approximate Optimization Algorithm (QAOA) is one of the most promising candidates for achieving quantum advantage over classical computers. However, existing compilers lack specialized methods for optimizing QAOA circuits. There are circuit patterns inside the QAOA circuits, and current quantum hardware has specific qubit connectivity topologies. Therefore, we propose Coqa to optimize QAOA circuit compilation tailored to different types of quantum hardware. Our method integrates a linear nearest-neighbor (LNN) topology and efficiently map the patterns of QAOA circuits to the LNN topology by heuristically checking the interaction based on the weight of problem Hamiltonian. This approach allows us to reduce the number of SWAP gates during compilation, which directly impacts the circuit depth and overall fidelity of the quantum computation. By leveraging the inherent patterns in QAOA circuits, our approach achieves more efficient compilation compared to general-purpose compilers. With our proposed method, we are able to achieve an average of 30% reduction in gate count and a 39x acceleration in compilation time across our benchmarks.

quant-ph

Adaptive variational simulation for open quantum systems

Emerging quantum hardware provides new possibilities for quantum simulation. While much of the research has focused on simulating closed quantum systems, the real-world quantum systems are mostly open. Therefore, it is essential to develop quantum algorithms that can effectively simulate open quantum systems. Here we present an adaptive variational quantum algorithm for simulating open quantum system dynamics described by the Lindblad equation. The algorithm is designed to build resource-efficient ansatze through the dynamical addition of operators by maintaining the simulation accuracy. We validate the effectiveness of our algorithm on both noiseless simulators and IBM quantum processors and observe good quantitative and qualitative agreement with the exact solution. We also investigate the scaling of the required resources with system size and accuracy and find polynomial behavior. Our results demonstrate that near-future quantum processors are capable of simulating open quantum systems.

quant-ph

Powerful Quantum Circuit Resizing with Resource Efficient Synthesis

In the noisy intermediate-scale quantum era, mid-circuit measurement and reset operations facilitate novel circuit optimization strategies by reducing a circuit's qubit count in a method called resizing. This paper introduces two such algorithms. The first one leverages gate-dependency rules to reduce qubit count by 61.6% or 45.3% when optimizing depth as well. Based on numerical instantiation and synthesis, the second algorithm finds resizing opportunities in previously unresizable circuits via dependency rules and other state-of-the-art tools. This resizing algorithm reduces qubit count by 20.7% on average for these previously impossible-to-resize circuits.

quant-ph