Searcharxiv⌕ Search

arXiv subjects

Gokul Subramanian Ravi

Publications and source records attributed to Gokul Subramanian Ravi.

At least 19 recordsLinked to original sources

C-Phase-Aware Compilation for Efficient Fault-Tolerant Quantum Execution

Achieving practical quantum advantage on fault-tolerant quantum computers (FTQC) is fundamentally constrained by the substantial spatial and temporal overheads required to map logical operations onto physical hardware. Existing compilation approaches typically adopt coarse-grained, slice-based abstractions that overlook fine-grained microarchitectural effects, such as routing contention, leading to inefficient resource utilization and limited alignment between algorithm structure and hardware capabilities. We introduce Qomet, a microarchitecture-aware compiler that tightly couples algorithmic properties with lattice surgery (LS) execution. By exploiting C-Phase gate commutativity, Qomet translates sequential operations into simultaneous multi-target interactions, natively leveraging LS to eliminate false dependencies and expose instruction-level parallelism. To support this, Qomet employs an adaptive, event-driven scheduler that captures precise spatial and routing constraints to overlap instructions temporally. By minimizing grid idling and routing contention, Qomet achieves a geometric-mean execution speedup of 4.29$\times$ and a maximum speedup of 59.7$\times$ across realistic workloads.

quant-ph↗

Improving Join Order Optimization on Gate-Based Quantum Computers via Structured Parameter Initialization

Join Order Optimization (JOO) is one of the most computationally expensive tasks in relational query optimization due to the exponential growth of possible join plans with increasing query size. Recent work has explored quantum and quantum-inspired approaches for solving JOO by reformulating the problem as a Quadratic Unconstrained Binary Optimization (QUBO) problem suitable for optimization using quantum hardware. However, many existing approaches have limited scalability on current gate-based quantum devices. In addition, little work has investigated the role of initialization strategies in improving the performance of gate-based quantum optimization for database workloads. In this work, we investigate gate-based quantum join order optimization using the Quantum Approximate Optimization Algorithm (QAOA) initialized with Scalable Parameter Initialization for QAOA (SPIQ). SPIQ is used to efficiently identify high-quality initial points in the quantum solution landscape for QAOA executed on a gate-based quantum computer. We evaluate the interaction between QUBO encoding, SPIQ initialization, and gate-based optimization on small-scale join ordering problems involving 3 and 4 relations. Our results show that structured initialization improves optimization stability and increases convergence toward high-quality join plans compared to uninformed initialization approaches. Across these small-scale, simulation-based instances, SPIQ increases the sampling frequency of the optimal join order by up to approximately 5$\times$ and yields final-state energies significantly lower than a randomly initialized QAOA. Overall, this work enhances existing gate-based quantum optimization while providing an initial proof of concept for applying SPIQ initialization to database query optimization workloads.

cs.DB↗

Scalable Clifford-Based Classical Initialization for the Quantum Approximate Optimization Algorithm

Variational Quantum Algorithms (VQAs), such as the Quantum Approximate Optimization Algorithm (QAOA), offer a promising route to tackling combinatorial optimization problems on near and intermediate-term quantum devices. However, their performance critically depends on the choice of initial parameters, and the limited expressiveness of the QAOA ansatz makes identifying effective initializations both difficult and unscalable. To address this, we propose a framework, Scalable Parameter Initialization for QAOA (SPIQ), that employs a relaxed QAOA ansatz to enable classical search over a set of Clifford-preparable quantum states that yield high-quality solutions. These states serve as superior QAOA initializations, driving rapid convergence while significantly reducing the quantum circuit evaluations needed to reach high-quality solutions and consequently lowering quantum-device cost. We present a scalable, application-agnostic initialization framework that achieves an absolute accuracy improvement of up to 80% over state-of-the-art initialization and reduces initial-state diversity by up to 10,000x across QUBO, PUBO, and PCBO problems spanning tens to hundreds of qubits. We further benchmark its performance on a wide range of problem formulations and instances derived from real-world datasets, demonstrating consistent and scalable improvements. Furthermore, we introduce two complementary strategies for selecting high-quality Clifford points identified by our search procedure and using them to seed multi-start optimization, thereby enhancing exploration and improving solution quality.

quant-ph↗

Mitigating Classical Resource Costs in Quantum Error Correction via Generalized qLDPC Predecoding

Large-scale fault-tolerant quantum computing (FTQC) will require quantum-classical interfaces (QCIs) that orchestrate real-time decoding over thousands to millions of logical qubits simultaneously. To scale FTQC systems, complex decoding resources must be shared between logical qubits, creating resource contention bottlenecks in the QCI. Mitigating this contention via optimal resource allocation remains an open problem. Lightweight predecoding techniques can reduce decoder utilization and average latency, both of which ease contention for shared decoding resources. To date, both decoder allocation and predecoding work is limited to the surface code. As focus shifts towards general qLDPC codes, slower decoding exacerbates resource contention, while code complexity precludes manual predecoder design. To address this gap, we introduce an automated framework designed to generate predecoders for arbitrary qLDPC codes. By independently handling up to 99.98% of the decoding workload, these predecoders reduce decoder utilization up to 4,090$\times$, including up to 81.19% decrease in expensive OSD post-processing and 59.96% decrease in extra RelayBP legs. An efficient, pipelined hardware architecture enables simultaneous decoding of ~1,800 BB code logical qubits on a single FPGA, while cryogenic ASIC implementation supports ~50,000-500,000 BB code logical qubits within a 1.5 W power budget at 4 K.

quant-ph↗

ExtraFerm: An Extended Matchgate Simulator

We present and open source Extraferm, a quantum circuit simulator tailored to chemistry applications. More specifically, our simulator can compute the Born-rule probabilities of samples obtained from circuits containing particle number-conserving matchgates and controlled-phase gates. We support both approximate and exact calculation of probabilities, and for approximate probability calculation, our simulator's runtime is exponential only in the magnitudes of the circuit's controlled-phase gate angles. This makes our simulator useful for simulating certain systems that are beyond the reach of conventional state vector methods. We demonstrate our simulator's utility by simulating the local cluster unitary Jastrow (LUCJ) ansatz and integrating it with sample-based quantum diagonalization (SQD) to improve the accuracy of molecular ground-state energy estimates with negligible computational overhead. More generally, we highlight a regime in which our simulator achieves substantially superior latency scaling and exponentially superior memory scaling over a tensor network simulator and a state vector simulator. As an efficient and flexible tool for simulating quantum chemistry circuits, our simulator enables new opportunities for enhancing near-term quantum algorithms in chemistry and related domains.

quant-ph↗

CryoZip: An Efficient Cryogenic Compressor for Quantum Error Correction Syndromes

Scaling fault tolerant quantum computing is increasingly constrained by the limited bandwidth and power budget across the 4 K to room temperature (RT) interface. We present CryoZip, a cross stack cryogenic compression framework that cooperates with a lightweight cryogenic quantum error correction (QEC) predecoder to reduce 4 K to RT syndrome transmission under realistic, circuit level noise. CryoZip targets sparse syndrome vectors with a sliding window compression architecture sized under strict decoding latency constraints to maximize energy efficiency. We implement and evaluate the design in 22 nm FDSOI characterized at 4 K, using vector based power, performance, and area analysis to obtain realistic hardware data. CryoZip achieves up to 48x compression, 1.8x higher than state of the art compressors, across various QEC codes while delivering 4 to 26x energy savings. When paired with a QEC predecoder, it yields over 14,238x bandwidth reduction, while energy savings rise to 42x when accounting for realistic QEC interface overheads.

quant-ph↗

Stalls and Spequlation: Pipelined Execution for Fault Tolerant Quantum Computation

Fault-tolerant quantum computation requires the coordinated action of three distinct systems: classical control logic, quantum hardware, and classical error decoders. Current scheduling models treat logical operations as atomic, hiding the fact that these subsystems operate sequentially and spend significant time idle. We present a pipelined execution framework that decomposes each logical operation into its component stages i.e. Control, Execute, and Decode. Building on this, we discuss some speculation strategies that allow successor operations to begin processing before their predecessors have completed decoding. We evaluate our framework on several common benchmarks and show that pipelining with speculation reduces total pipeline steps by 20-40% compared to a no-speculation baseline. The most aggressive strategy consistently outperforms conservative alternatives, even though partial rollback is needed at times, because the per-rollback penalty is small relative to the parallelism gained. We further show that speculation facilitates load balancing by distributing work more evenly across the heterogeneous subsystems of a fault-tolerant quantum computer, converting idle time into useful computation while also saving on execution time.

quant-ph↗

Classical State Preparation for Variational Quantum Algorithms via Reinforcement Learning

Variational Quantum Algorithms (VQAs) potentially offer a pathway to practical quantum advantage, but their optimization is heavily hindered by barren plateaus and numerous local minima. While classically simulable Clifford circuits can warm-start VQAs to accelerate convergence, existing heuristic-based initialization methods struggle to scale within vast combinatorial search spaces. To overcome this bottleneck, we propose CRiSP (a Clifford Reinforcement Learning agent for State Preparation), a framework that formulates discrete prefix selection as a sequential decision-making problem. CRiSP utilizes Neural-Guided Monte Carlo Tree Search, driven by a Transformer-based policy trained via self-play, to insert learned Clifford gates before fixed parameterized rotations. This enables the construction of high-quality initial states entirely through polynomial-time classical stabilizer simulation without altering the underlying circuit architecture. By integrating a curriculum learning strategy that progressively expands the search horizon, the agent efficiently scales to deep circuits. Evaluated on QAOA benchmarks of up to $22$ qubits and $1{,}370$ parameters, CRiSP outperforms state-of-the-art Clifford initialization methods by a mean of $3.17\times$ (max $45.02\times$) in average energy accuracy and $2.44\times$ (max $16.01\times$) in best-achieved energy accuracy. Assessments on VQE tasks further demonstrate the framework's robustness and generalizability.

quant-ph↗

Price and Payoff: Non-Determinism in Fault Tolerant Quantum Computation

A promising approach to achieving scalable fault-tolerant quantum computation is the use of quantum error correction (QEC) codes augmented with magic states i.e. resource states produced via distillation, cultivation, or $R_z$ synthesis and teleported into the circuit as needed. Because magic-state production dominates the space-time volume of fault-tolerant programs, system architects must decide how many production units to allocate. Current approaches rely on deterministic analysis that either provisions for worst-case peak demand (wasting valuable qubit resources on factories that are never simultaneously utilized) or assumes average demand, which increases execution time. In this work, we build a simulation framework that couples circuit scheduling with different stochastic magic state production models, and use it to quantify the impact of non-determinism on circuit execution. We show that non-determinism has a dual effect that deterministic models cannot capture: it inflates total execution time (the price), while deflating peak per-cycle resource demand (the payoff). For distillation-based architectures, this demand smoothing shifts the space-time-optimal provisioning point: fewer factories are needed to minimize space-time volume than deterministic analysis predicts. Across benchmarks, stochastic-aware provisioning reduces space-time volume by up to 27% compared to the deterministic optimum for distillation, while requiring up to 30% fewer factories. We characterize these effects across each preparation mechanism, map the resulting design-space tradeoffs, and demonstrate that static resource estimation systematically mis-characterizes the cost of fault-tolerant execution. Our results establish that stochastic-aware analysis is necessary for right-sizing the factory allocations and should replace deterministic heuristics as the standard methodology for FTQC resource planning.

quant-ph↗

Accelerating BP-based decoders for QLDPC Codes with Local Syndrome-Based Preprocessing

Due to the high error rate of qubits, detecting and correcting errors is essential for achieving fault-tolerant quantum computing (FTQC). Quantum low-density parity-check (QLDPC) codes are one of the most promising quantum error correction (QEC) methods due to their high encoding rates. BP (Belief Propagation)-based decoders are widely used and highly competitive for QLDPC codes because BP offers inherent parallelism and strong scalability. However, BP-based decoders still suffer from high decoding latency, a large portion of which is spent in the iterative BP stage. In this paper, we propose a lightweight preprocessing step that utilizes local patterns in the syndrome to detect likely trivial error events and provide them as hints to BP-based decoders. These hints accelerate BP convergence and thereby reduce the overall decoding time. The proposed preprocessing step offers a broadly compatible approach to reducing the latency of BP-based QLDPC decodes. On the bivariate bicycle code $[[144,12,12]]$ at low physical error rates, our method achieves a $10\times$ speedup in decoding time for BP-OSD, and more than $2\times$ speedup for both BP-LSD and Relay-BP. Our method maintains the logical error rate when combined with BP-OSD and Relay-BP, while further achieving a significant reduction in logical error rate when combined with BP-LSD.

quant-ph↗

Computer Science Challenges in Quantum Computing: Early Fault-Tolerance and Beyond

Quantum computing is entering a period in which progress will be shaped as much by advances in computer science as by improvements in hardware. The central thesis of this report is that early fault-tolerant quantum computing shifts many of the primary bottlenecks from device physics alone to computer-science-driven system design, integration, and evaluation. While large-scale, fully fault-tolerant quantum computers remain a long-term objective, near- and medium-term systems will support early fault-tolerant computation with small numbers of logical qubits and tight constraints on error rates, connectivity, latency, and classical control. How effectively such systems can be used will depend on advances across algorithms, error correction, software, and architecture. This report identifies key research challenges for computer scientists and organizes them around these four areas, each centered on a fundamental question.

quant-ph↗

TreeVQA: A Tree-Structured Execution Framework for Shot Reduction in Variational Quantum Algorithms

Variational Quantum Algorithms (VQAs) are promising for near- and intermediate-term quantum computing, but their execution cost is substantial. Each task requires many iterations and numerous circuits per iteration, and real-world applications often involve multiple tasks, scaling with the precision needed to explore the application's energy landscape. This demands an enormous number of execution shots, making practical use prohibitively expensive. We observe that VQA costs can be significantly reduced by exploiting execution similarities across an application's tasks. Based on this insight, we propose TreeVQA, a tree-based execution framework that begins by executing tasks jointly and progressively branches only as their quantum executions diverge. Implemented as a VQA wrapper, TreeVQA integrates with typical VQA applications. Evaluations on scientific and combinatorial benchmarks show shot count reductions of $25.9\times$ on average and over $100\times$ for large-scale problems at the same target accuracy. The benefits grow further with increasing problem size and precision requirements.

quant-ph↗

Pinball: A Cryogenic Predecoder for Surface Code Decoding Under Circuit-Level Noise

Scaling fault tolerant quantum computers, especially cryogenic systems based on the surface code, to millions of qubits is challenging due to poorly-scaling data processing and power consumption overheads. One key hurdle is the design of real-time quantum error correction (QEC) decoders, which demands high data rates for error processing; this is particularly apparent in systems with cryogenic qubits and room temperature (RT) decoders. In response, cryogenic predecoding using lightweight logic has been proposed to handle sparse errors in the cryogenic domain. However, prior work only accounts for a subset of error sources in real-world quantum systems with limited accuracy, often degrading performance below useful levels in practical scenarios. Moreover, prior reliance on SFQ logic precludes detailed architecture-technology co-optimization. To address these limitations, this paper introduces Pinball, a comprehensive design in cryogenic CMOS of a QEC predecoder for the surface code tailored to realistic, circuit-level noise. By accounting for error generation and propagation through QEC circuits, our design achieves higher predecoding accuracy, outperforming logical error rates (LER) of the current state-of-the-art (SOTA) cryogenic predecoder by nearly six orders of magnitude. Remarkably, despite operating under much stricter power and area constraints, Pinball also reduces LER by 32.58x and 5x, respectively, compared to SOTA RT predecoder and RT ensemble configurations. By increasing cryogenic coverage, we also reduce syndrome bandwidth up to 3780.72x. Through co-design with 4 K-characterized 22nm FDSOI technology, we achieve peak power consumption under 0.56 mW. Voltage/frequency scaling and body biasing enable 22.2x lower typical power consumption, yielding up to 67.4x total energy savings. Assuming a 1.5 W 4 K power budget, our predecoder supports up to 2,668 logical qubits at d=21.

quant-ph↗

Enhancing the Clique Local Decoder to Correct Length-2 Space Errors in the Surface Code

The growing demand for fault-tolerant quantum computing drives the need for efficient, scalable Quantum Error Correction (QEC) strategies. Conventional decoders designed for worst-case error scenarios incur significant overhead, prompting the development of local decoders, that leverage the sparse and often trivial nature of many quantum errors, to support the conventional decoders. The previously proposed Clique decoder addresses this by handling isolated, length-1 space and time errors within the cryogenic environment with minimal hardware costs, thereby mitigating I/O bandwidth constraints between cryogenic quantum systems and room-temperature processors. Building on this foundation, we propose Clique_L2 that extends the Clique-based approach by relaxing some original constraints and incorporating additional low-cost logic to also correct length-2 error chains in space, which become non-trivial occurrences at higher physical error rates and code distances. This enhanced capability not only further reduces out-of-the-fridge data transmission but also adapts more effectively to clustered errors observed under a variety of noise models. Specifically, under data-qubit-only errors and uniformly random noise, Clique_L2 achieves up to 8.95x decoding bandwidth reduction over the original Clique (or Clique_L1) decoder, especially beneficial at higher code distances. When clustered errors and longer error chains are more likely to occur, Clique_L2 achieves up to 18.3x decoding bandwidth reduction over Clique_L1, achieving substantial benefits across a wide range of physical qubit error rates.

quant-ph↗

Variational Quantum Algorithms in the era of Early Fault Tolerance

Quantum computing roadmaps predict the availability of 10,000 qubit devices within the next 3-5 years. With projected two-qubit error rates of 0.1%, these systems will enable certain operations under quantum error correction (QEC) using lightweight codes, offering significantly improved fidelities compared to the NISQ era. However, the high qubit cost of QEC codes like the surface code (especially at near-threshold physical error rates) limits the error correction capabilities of these devices. In this emerging era of Early Fault Tolerance (EFT), it will be essential to use QEC resources efficiently and focus on applications that derive the greatest benefit. In this work, we investigate the implementation of Variational Quantum Algorithms in the EFT regime (EFT-VQA). We introduce partial error correction (pQEC), a strategy that error-corrects Clifford operations while performing Rz rotations via magic state injection instead of the more expensive T-state distillation. Our results show that pQEC can improve VQA fidelities by 9.27x over standard approaches. Furthermore, we propose architectural optimizations that reduce circuit latency by ~2x, and achieve qubit packing efficiency of 66% in the EFT regime.

quant-ph↗

Clifford Assisted Optimal Pass Selection for Quantum Transpilation

The fidelity of quantum programs in the NISQ era is limited by high levels of device noise. To increase the fidelity of quantum programs running on NISQ devices, a variety of optimizations have been proposed. These include mapping passes, routing passes, scheduling methods and standalone optimisations which are usually incorporated into a transpiler as passes. Popular transpilers such as those proposed by Qiskit, Cirq and Cambridge Quantum Computing make use of these extensively. However, choosing the right set of transpiler passes and the right configuration for each pass is a challenging problem. Transpilers often make critical decisions using heuristics since the ideal choices are impossible to identify without knowing the target application outcome. Further, the transpiler also makes simplifying assumptions about device noise that often do not hold in the real world. As a result, we often see effects where the fidelity of a target application decreases despite using state-of-the-art optimisations. To overcome this challenge, we propose OPTRAN, a framework for Choosing an Optimal Pass Set for Quantum Transpilation. OPTRAN uses classically simulable quantum circuits composed entirely of Clifford gates, that resemble the target application, to estimate how different passes interact with each other in the context of the target application. OPTRAN then uses this information to choose the optimal combination of passes that maximizes the target application's fidelity when run on the actual device. Our experiments on IBM machines show that OPTRAN improves fidelity by 87.66% of the maximum possible limit over the baseline used by IBM Qiskit. We also propose low-cost variants of OPTRAN, called OPTRAN-E-3 and OPTRAN-E-1 that improve fidelity by 78.33% and 76.66% of the maximum permissible limit over the baseline at a 58.33% and 69.44% reduction in cost compared to OPTRAN respectively.

quant-ph↗

Predictive Window Decoding for Fault-Tolerant Quantum Programs

Real-time decoding is a key ingredient in future fault-tolerant quantum systems, yet many decoders are too slow to run in real time. Prior work has shown that parallel window decoding schemes can scalably meet throughput requirements in the presence of increasing decoding times, given enough classical resources. However, windowed decoding schemes require that some decoding tasks be delayed until others have completed, which can be problematic during time-sensitive operations such as T gate teleportation, leading to suboptimal program runtimes. To alleviate this, we introduce a speculative window decoding scheme. Taking inspiration from branch prediction in classical computer architecture our decoder utilizes a light-weight speculation step to predict data dependencies between adjacent decoding windows, allowing multiple layers of decoding tasks to be resolved simultaneously. Through a state-of-the-art compilation pipeline and a detailed simulator, we find that speculation reduces application runtimes by 40% on average compared to prior parallel window decoders.

quant-ph↗

Quantum-centric Supercomputing for Materials Science: A Perspective on Challenges and Future Directions

Computational models are an essential tool for the design, characterization, and discovery of novel materials. Hard computational tasks in materials science stretch the limits of existing high-performance supercomputing centers, consuming much of their simulation, analysis, and data resources. Quantum computing, on the other hand, is an emerging technology with the potential to accelerate many of the computational tasks needed for materials science. In order to do that, the quantum technology must interact with conventional high-performance computing in several ways: approximate results validation, identification of hard problems, and synergies in quantum-centric supercomputing. In this paper, we provide a perspective on how quantum-centric supercomputing can help address critical computational problems in materials science, the challenges to face in order to solve representative use cases, and new suggested directions.

quant-ph↗