SearcharxivSearch

arXiv subjects

Costin Iancu

Publications and source records attributed to Costin Iancu.

At least 19 recordsLinked to original sources

AC/DC: Automated Compilation for Dynamic Circuits

Dynamic quantum circuits incorporate mid-circuit measurements (MCMs) and feed-forward operations are crucial for manipulating quantum information. They have been broadly used in quantum error correction and quantum teleportation. Recently, they are utilized to prepare certain states and long-range entangling gates as well as reduce resource overhead in quantum algorithms. In this paper, we present AC/DC, a novel Automated Compilation framework for generating Dynamic quantum Circuits that prepare any unitary operators or states, leveraging numerical optimization-based circuit synthesis methods. The first contribution is introducing optimization objective functions incorporating MCMs and feed-forward operations. The second contribution is embedding these into a popular open-source quantum circuit synthesis framework. We demonstrate generating dynamic circuits for long range entangling gates, circuit optimization, lattice simulations, and state preparation, with validation through simulation and quantum hardware. Furthermore, we perform a noise analysis to assess the impact of MCM and gate errors, identifying scenarios where dynamic circuits provide significant benefits. The dynamic circuits generated by our framework show substantial improvements in reducing circuit depth and, in some cases, the number of gates. To our knowledge, this is the first practical procedure to generate dynamic quantum circuits, paving the way for enhanced circuit generation and optimization methods for near-term quantum computers.

quant-ph

Scalable Benchmarking Framework for Dynamic Quantum Circuits

Dynamic quantum circuits with mid-circuit measurements (MCMs) and feed-forward operations play a crucial role in various applications, such as quantum error correction and quantum algorithms. With advancements in quantum hardware enabling the implementation of MCM and feed-forward loops, the use of dynamic circuits has become increasingly prevalent. There is a significant need for a benchmarking framework specially designed for dynamic circuits to capture their unique properties, as current benchmarking tools are designed primarily for unitary circuits and cannot be trivially extended to dynamic circuits. We propose dynamarq, a scalable and hardware-agnostic benchmarking framework for dynamic circuits. We collect a set of dynamic circuit benchmarks spanning various applications and propose a broad set of circuit features to characterize the structure of these dynamic circuits. We run them on two IBM quantum processors and the Quantinuum Helios-1E emulator, and propose scalable, application-dependent fidelity scores for each benchmark based on hardware execution results. We perform statistical modeling to identify correlations between circuit features and fidelity scores, and demonstrate highly accurate fidelity prediction using our model. Our model parameters are also transferable across hardware backends and calibration cycles. Our framework facilitates the understanding of dynamic circuit structures and provides insights for designing and optimizing dynamic circuits to achieve high execution fidelity on quantum hardware.

quant-ph

Approximate synthesis of general single-qubit unitaries over the Clifford+$\sqrt{T}$ gate set

For the standard Clifford+$T$ gate set, deterministic, ancilla-free synthesis now attains the minimal $T$-count for general single-qubit unitaries (Morisaki et al., arXiv:2510.05816). The $\sqrt{T}$ gate rotates by half the angle of $T$, generating a finer lattice of implementable operations. It was assumed that access to this magic state lowers the cost of deterministic and ancilla-free synthesis of general single-qubit unitaries, but no direct Clifford+$\sqrt{T}$ algorithm existed for this case. We provide one by extending the integer lattice-point enumeration method of Morisaki et al. We adopt a resource state cost model based on the magic-state catalysis approach of Gidney and Fowler (arXiv:1812.01238). On Haar-random targets synthesized to precisions ranging from $\varepsilon=10^{-3}$ to $10^{-8}$, the cost of Clifford+$\sqrt{T}$ circuits scales as $2.4\log_2(1/\varepsilon)$ compared to $3.0\log_2(1/\varepsilon)$ for the provably $T$-count-optimal Clifford+$T$ circuits. Once a one-time catalyst state is amortized, the Clifford+$\sqrt{T}$ circuits are never costlier than their Clifford+$T$ counterparts.

quant-ph

PureMagic: A Dynamic Scheduler for Lattice Surgery

Fault-tolerant quantum computation on surface codes requires magic states for universal computation. Traditional distillation factories deliver magic states deterministically but consume large areas of logical qubits, forcing static, peripheral placement. Magic state cultivation reduces magic state preparation to a single logical qubit, but is inherently stochastic, making static scheduling infeasible. We introduce PureMagic, a dynamic scheduler that eliminates dedicated bus patches by repurposing all ancilla patches for both routing and cultivation. When a patch is needed for routing, cultivation is interrupted and restarted afterward, naturally cutting off the long tail of cultivation times and ensuring no ancilla is ever idle. We also introduce a weight limit on Tableau transpilation that trades gate count for parallelism, which PureMagic is particularly well-suited to exploit. Across 29 benchmark circuits, PureMagic achieves 43% to 152% efficiency improvement over bus routing, uses 19% to 80% fewer logical qubits, and reduces average magic state preparation time by 4.5x. Compared to DASCOT, a state-of-the-art static scheduler, PureMagic is up to 21x more efficient when magic state preparation costs are included. PureMagic's scheduled volumes fall between the conservative and optimistic FLASQ theoretical lower bounds, demonstrating near-optimal use of ancilla resources.

quant-ph

Collective neutrino oscillations: Many-body non-forward effects and non-classicality

Neutrino evolution in dense astrophysical environments is typically described either within a quantum kinetic framework, which neglects the build-up of multi-body correlations, or through simplified many-body calculations that allow significant entanglement to develop. In this work, we compare these two approaches in a simple neutrino-gas configuration, with particular emphasis on the role of non-forward scattering processes. These effects are incorporated either through a collision term in the kinetic description, or by considering the full neutrino-neutrino many-body Hamiltonian. We highlight differences between the two descriptions in both their characteristic timescales and asymptotic behavior. Motivated by the natural suitability of quantum computing for many-body calculations, we further investigate the non-classicality of neutrino evolution, discussing Trotter error scaling, along with the associated costs of constructing quantum circuits in terms of entangling gates and non-Clifford gates. We find that the resources needed for neutrino many-body evolution are on the low end of typical high-energy physics problems and on the mid to high end with respect to quantum chemistry problems. For the full Hamiltonian, resource requirements increase relative to the truncated version. We emphasize the importance of efficient fermion-to-qubit encodings, which are essential for reducing the substantial computational resources required for such simulations.

hep-ph

Multi-Qubit Dyadic Phase Fixing for Fault-Tolerant Quantum Compilation

Fault-tolerant quantum computing requires translating application-level quantum circuits into the Clifford+$T$ gate set, where the $T$ gate is the dominant resource cost. Phase kickback is an ancilla-based technique that can dramatically reduce $T$-count for rotations with dyadic angles, but has previously been limited to highly structured circuit families. We present Dyadic Phase Fixing (DPF), a general multi-qubit synthesis tool that extends phase kickback to general quantum circuits. DPF uses numerical unitary synthesis to greedily extract dyadic angle rotations from any input circuit. Combined with a decision matrix to automatically size the final phase gradient register, our end-to-end workflow achieves up to 70% reduction in $T$-count compared to \texttt{gridsynth} and up to 60% compared to Repeat-Until-Success synthesis on a diverse set of benchmarks. We map these compiled circuits to a surface-code architecture to evaluate space-time volume, demonstrating up to a 60\% reduction in this metric as well. However, for some circuits and mapping strategies the two metrics diverge significantly, demonstrating that $T$-count alone is a useful but incomplete proxy for fault-tolerant program costs.

quant-ph

Linear Complexity Fermionic Simulation on Quantum Devices with Hardware Connectivity Constraints

Simulating fermionic systems on quantum hardware requires compiling fermionic Hamiltonians into executable quantum circuits. Existing approaches treat each compilation stage independently, applying heuristics with localized objectives that produce circuits with superquartic gate count and depth scaling and compilation times reaching several hours for large instances. We present Accordion, an end-to-end framework that co-designs the fermion-to-qubit mapping with circuit synthesis and hardware routing. Accordion fixes the Jordan Wigner mapping, which despite its higher Pauli weight produces Pauli operators with structural regularity that enables provably efficient circuit generation. For full-rank all-to-all electronic structure Hamiltonians, we prove O(N^4) gate count and circuit depth, matching the information-theoretic lower bound imposed by the Theta(N^4) second excitation terms. On linear, IBM heavy-hex, and square-grid architectures, Accordion reduces gate count by up to 79% and circuit depth by up to 77% relative to the best baseline.

cs.AR

Logical Resource Estimation for Quantum State Preparation with Compilation

Quantum state preparation is a fundamental primitive in quantum algorithms for encoding classical data into quantum amplitudes. We compare the cost of preparing general $n$-qubit states with real amplitudes using two common paradigms: rotation-based methods, based on controlled rotations, and sampling-based methods, based on a structured representation of the target state. Although these approaches are often theoretically compared using CNOT count and $T$-count, their relative performance in total gate count remains less well understood practically. We compare representative rotation-based and sampling-based methods using $T$-count and total gate count, and analyze how compilation overhead affects their relative performance. We also develop a software package for compiling state preparation circuits, designed as a practical subroutine for more general quantum computations. Numerical experiments on resource states and quantum states related to quantum chemistry, condensed matter physics, and simulation via Magnus expansion over a range of target accuracies $ε$ support the analysis. Our results show that sampling-based methods achieve asymptotically lower $T$-count and retain an overall advantage after accounting for total gate count and compilation overhead.

quant-ph

Optimizing PyTorch Inference with LLM-Based Multi-Agent Systems

Maximizing performance on available GPU hardware is an ongoing challenge for modern AI inference systems. Traditional approaches include writing custom GPU kernels and using specialized model compilers to tune high-level code for specific GPU targets. Recent work shows that LLM-based multi-agent systems can effectively perform such tuning, often outperforming existing compilers and eliminating the need for manual kernel development. However, the dynamics of multi-agent systems for this task remain unexplored. In this work, we present a logical framework for comparing multi-agent PyTorch optimization systems. Our evaluation shows that exploit-heavy strategies perform best when paired with error-fixing agents, and that performance correlates with the granularity of optimization steps. The best implementation achieves an average 2.88x speedup over PyTorch Eager (1.85x over torch.compile) on an H100 GPU across diverse tasks in KernelBench, a benchmark suite covering a range of machine learning architectures in PyTorch. Code is publicly available at: https://github.com/pike-project/pike

cs.MA

TopoLS: Lattice Surgery Compilation via Topological Program Transformations

Lattice surgery is a leading approach for implementing fault-tolerant logical operations in surface code quantum computing, but compiling efficient lattice surgery layouts remains challenging. Existing compilers are largely circuit-centric and operate directly on gate sequences, limiting their ability to exploit the topological flexibility of merge-split operations and minimize space--time volume. We present TopoLS, a topology-centric compiler that uses ZX diagrams as an intermediate representation for lattice surgery compilation. TopoLS combines semantic-preserving ZX-level program transformations, including spider fusion and topology-aware slicing, with a Monte Carlo Tree Search (MCTS)-based synthesis procedure that constructs pipe-diagram embeddings by jointly optimizing placement and routing in 3D space--time. To scale to large circuits, TopoLS further introduces topology-aware partitioning that decomposes the compilation task into bounded subproblems and limits the routing frontier during embedding. Across evaluated benchmarks, TopoLS achieves an average $46\%$ reduction in space--time volume over prior circuit-centric compilers, with improvements ranging from $25\%$ to $90\%$, and exhibits strong empirical scalability on large benchmark families. Compared with SAT-based formulations that become intractable on larger instances, TopoLS offers a practical end-to-end solution for optimized lattice surgery compilation. TopoLS has been integrated into the TQEC ecosystem, enabling downstream circuit-level simulation and resource estimation workflows.

quant-ph

AlphaSyndrome: Tackling the Syndrome Measurement Circuit Scheduling Problem for QEC Codes

Quantum error correction (QEC) is essential for scalable quantum computing, yet repeated syndrome-measurement cycles dominate its spacetime and hardware cost. Although stabilizers commute and admit many valid execution orders, different schedules induce distinct error-propagation paths under realistic noise, leading to large variations in logical error rate. Outside of surface codes, effective syndrome-measurement scheduling remains largely unexplored. We present AlphaSyndrome, an automated synthesis framework for scheduling syndrome-measurement circuits in general commuting-stabilizer codes under minimal assumptions: mutually commuting stabilizers and a heuristic decoder. AlphaSyndrome formulates scheduling as an optimization problem that shapes error propagation to (i) avoid patterns close to logical operators and (ii) remain within the decoder's correctable region. The framework uses Monte Carlo Tree Search (MCTS) to explore ordering and parallelism, guided by code structure and decoder feedback. Across diverse code families, sizes, and decoders, AlphaSyndrome reduces logical error rates by 80.6% on average (up to 96.2%) relative to depth-optimal baselines, matches Google's hand-crafted surface-code schedules, and outperforms IBM's schedule for the Bivariate Bicycle code.

cs.ET

Application Scale Quantum Circuit Compilation with Controlled Error

Compilation and optimization of quantum circuits are critical components in the execution of algorithms on quantum computers. These components must successfully balance two competing priorities: minimizing the number of expensive resources, such as two-qubit gates or arbitrary angle single-qubit rotations, and minimizing the approximation error of the compiled circuit to the ideal target unitary describing the quantum algorithm. We develop a practical workflow for managing and optimizing this tradeoff, which enables quantum circuit compilation and optimization at scales of hundreds of qubits. Our workflow is able to tackle circuits at such large scales while providing rigorous guarantees on circuit output error by leveraging circuit partitioning and the notion of averaging over circuit ensembles. We demonstrate our workflow on several benchmark algorithmic circuits acting on up to 380 qubits, and show that it can simultaneously achieve substantial reductions in resource-intensive gates and control output errors, offering a practical and scalable strategy for both near-term and fault-tolerant quantum computing.

quant-ph

High-Precision Multi-Qubit Clifford+T Synthesis by Unitary Diagonalization

Resource-efficient and high-precision approximate synthesis of quantum circuits expressed in the Clifford+T gate set is vital for Fault-Tolerant quantum computing. Efficient optimal methods are known for single-qubit RZ unitaries, otherwise the problem is generally intractable. Search-based methods, like simulated annealing, empirically generate low resource cost approximate implementations of general multi-qubit unitaries so long as low precision (Hilbert-Schmidt distances of e>10^-2) can be tolerated. These algorithms build up circuits that directly invert target unitaries. We instead leverage search-based methods to first approximately diagonalize a unitary, then perform the inversion analytically. This lets difficult continuous rotations be bypassed and handled in a post-processing step. Our approach improves both the implementation precision and run time of synthesis algorithms by orders of magnitude when evaluated on unitaries from real quantum algorithms. On benchmarks previously synthesizable only with analytical techniques like the Quantum Shannon Decomposition, diagonalization uses an average of 95% fewer non-Clifford gates.

quant-ph

Intermediate-temperature topological Uhlmann phase on IBM quantum computers

A spin-1 system can exhibit an intermediate-temperature topological regime with a quantized Uhlmann phase sandwiched by topologically trivial low- and high-temperature regimes. We present a quantum circuit consisting of system and ancilla qubits plus a probe qubit which prepares an initial state corresponding to the purified state of a spin-1 system at finite temperature, evolves the system according to the Uhlmann process, and measures the Uhlmann phase via expectation values of the probe qubit. Although classical simulations suggest the quantized Uhlmann phase is observable on IBM's noisy intermediate-scale quantum (NISQ) computers, an implementation of the circuit without any optimization exceeds the gate count for the error budget and results in unresolved signals. Through a series of optimization with Qiskit and BQSQit, the gate count can be substantially reduced, making the jumps of the Uhlmann phase more visible. A recent hardware upgrade of IBM quantum computers further improves the signals and leads to a clearer demonstration of interesting finite-temperature topological phenomena on NISQ hardware.

quant-ph

QTurbo: A Robust and Efficient Compiler for Analog Quantum Simulation

Analog quantum simulation leverages native hardware dynamics to emulate complex quantum systems with great efficiency by bypassing the quantum circuit abstraction. However, conventional compilation methods for analog simulators are typically labor-intensive, prone to errors, and computationally demanding. This paper introduces QTurbo, a powerful analog quantum simulation compiler designed to significantly enhance compilation efficiency and optimize hardware execution time. By generating precise and noiseresilient pulse schedules, our approach ensures greater accuracy and reliability, outperforming the existing state-of-theart approach.

quant-ph

Leveraging AI for Productive and Trustworthy HPC Software: Challenges and Research Directions

We discuss the challenges and propose research directions for using AI to revolutionize the development of high-performance computing (HPC) software. AI technologies, in particular large language models, have transformed every aspect of software development. For its part, HPC software is recognized as a highly specialized scientific field of its own. We discuss the challenges associated with leveraging state-of-the-art AI technologies to develop such a unique and niche class of software and outline our research directions in the two US Department of Energy--funded projects for advancing HPC Software via AI: Ellora and Durban.

cs.SE

Making Neural Networks More Suitable for Approximate Clifford+T Circuit Synthesis

Machine Learning with deep neural networks has transformed computational approaches to scientific and engineering problems. Central to many of these advancements are precisely tuned neural architectures that are tailored to the domains in which they are used. In this work, we develop deep learning techniques and architectural modifications that improve performance on reinforcement learning guided quantum circuit synthesis-the task of constructing a circuit that implements a given unitary matrix. First, we propose a global phase invariance operation which makes our architecture resilient to complex global phase shifts. Second, we demonstrate how augmenting data with small random unitary perturbations during training enables more robust learning. Finally, we show how encoding numerical data with techniques from image processing allow networks to better detect small but significant changes in data. Our work enables deep learning approaches to better synthesize quantum circuits that implement unitary matrices.

quant-ph

HATT: Hamiltonian Adaptive Ternary Tree for Optimizing Fermion-to-Qubit Mapping

This paper introduces the Hamiltonian-Adaptive Ternary Tree (HATT) framework to compile optimized Fermion-to-qubit mapping for specific Fermionic Hamiltonians. In the simulation of Fermionic quantum systems, efficient Fermion-to-qubit mapping plays a critical role in transforming the Fermionic system into a qubit system. HATT utilizes ternary tree mapping and a bottom-up construction procedure to generate Hamiltonian aware Fermion-to-qubit mapping to reduce the Pauli weight of the qubit Hamiltonian, resulting in lower quantum simulation circuit overhead. Additionally, our optimizations retain the important vacuum state preservation property in our Fermion-to-qubit mapping and reduce the complexity of our algorithm from $O(N^4)$ to $O(N^3)$. Evaluations and simulations of various Fermionic systems demonstrate $5\sim20\%$ reduction in Pauli weight, gate count, and circuit depth, alongside excellent scalability to larger systems. Experiments on the Ionq quantum computer also show the advantages of our approach in noise resistance in quantum simulations.

quant-ph