SearcharxivSearch

arXiv subjects

Dawei Ding

Publications and source records attributed to Dawei Ding.

At least 19 recordsLinked to original sources

Efficient Compilation for Hamiltonian Simulation via Global Binary Symplectic Form Simplification

Hamiltonian simulation is a core quantum workload, underpinning variational quantum algorithms and Trotterized time evolution. Such programs are expressed as Pauli exponential sequences, exhibiting structural patterns that are highly amenable to high-level synthesis and optimization. Existing compilers, however, fail to fully unlock the optimization potential of their global algebraic structure, even when employing advanced graph- or tableau-based methods. We present Symphony, a holistic compilation approach built on the binary symplectic form (BSF) representation of Pauli strings. Unlike prior group-wise BSF simplification and path-based Pauli network synthesis, Symphony applies generalized controlled-Pauli Clifford transformations directly to a global BSF tableau, adaptively reducing active Pauli rows and emitting eligible two-qubit blocks other than single-qubit rotations in a forward Clifford frame. Following algebraic simplification, Symphony performs a causality-preserving block rescheduling heuristic that respects frame-induced dependencies while exposing extensive two-qubit block parallelism opportunities. This streamlined compilation style comprehensively exploits simultaneous simplification and commutativity opportunities, achieving efficient global optimization without relying on computationally expensive heuristics or long-horizon searches. Across the generic Hamiltonian simulation benchmarks in HamLib, Symphony achieves average reductions of 59% in two-qubit gate count and 91% in circuit depth. It strictly Pareto-dominates prior state-of-the-art compilers, requiring 1.14--1.58$\times$ fewer two-qubit gates and especially shrinking two-qubit circuit depth by a substantial factor of 1.87--5.67$\times$ on average.

quant-ph

On the KAK Decomposition and Equivalence Classes

The KAK decomposition is a fundamental tool in Lie theory and quantum computing. Despite its widespread use, the mathematical foundations remain incomplete, particularly regarding the precise conditions for the decomposition and the characterization of equivalence classes under multiplication by elements of $K$. Here, we present a mathematical theory of the KAK decomposition for connected compact semisimple Lie groups and derive the decomposition for $\mathrm{SU}(4)$. In particular, we clarify the relationship between various definitions of a Cartan decomposition in the literature and give a complete proof of a general KAK decomposition theorem. We then distinguish two distinct notions of KAK equivalence classes, double coset equivalence and projective equivalence, thereby addressing mathematical inconsistencies regarding KAK classification in the literature. Specifically, for $\mathrm{SU}(4)$, we show that local equivalence classes under multiplication by $\mathrm{SU}(2)\otimes \mathrm{SU}(2)$ are geometrically represented not by the usual "Weyl chamber" as claimed in the existing literature. Instead, the "Weyl chamber" is only recovered by the projective-local equivalence which disregards global phases. We develop a systematic theory for determining equivalence and uniqueness for both notions of equivalence. Our work establishes a rigorous Lie-theoretic foundation for the theory of quantum gates and circuits.

quant-ph

Quantum Telepathy: A Quantum Technology with Near-Term Applications

Quantum telepathy is the concept of using quantum entanglement to solve real-world problems involving decision coordination between parties with restricted communication. One possible reason for this restriction is a latency constraint: some pairs of parties do not have enough time to communicate with each other before they have to produce their outputs. Example scenarios include high frequency trading and distributed systems. Another reason is physical or operational isolation: for some pairs of parties, there is an obstacle to communication. Example scenarios include locating a stray traveler by a rescue team and coordination within a network where nodes are owned by competing firms. In this paper we give a concise overview of the different application areas of quantum telepathy. We find that these real-world problems can be modeled as a nonlocal game or its generalizations. We also discuss possible physical implementations. Quantum telepathy guarantees a quantum advantage via Bell's theorem and can directly solve real-world problems, such as reducing risk in high frequency trading or balancing data loads efficiently in ad hoc networks. Moreover, this quantum advantage can be physically realized with existing or near-term quantum hardware.

quant-ph

TakeAD: Preference-based Post-optimization for End-to-end Autonomous Driving with Expert Takeover Data

Existing end-to-end autonomous driving methods typically rely on imitation learning (IL) but face a key challenge: the misalignment between open-loop training and closed-loop deployment. This misalignment often triggers driver-initiated takeovers and system disengagements during closed-loop execution. How to leverage those expert takeover data from disengagement scenarios and effectively expand the IL policy's capability presents a valuable yet unexplored challenge. In this paper, we propose TakeAD, a novel preference-based post-optimization framework that fine-tunes the pre-trained IL policy with this disengagement data to enhance the closed-loop driving performance. First, we design an efficient expert takeover data collection pipeline inspired by human takeover mechanisms in real-world autonomous driving systems. Then, this post optimization framework integrates iterative Dataset Aggregation (DAgger) for imitation learning with Direct Preference Optimization (DPO) for preference alignment. The DAgger stage equips the policy with fundamental capabilities to handle disengagement states through direct imitation of expert interventions. Subsequently, the DPO stage refines the policy's behavior to better align with expert preferences in disengagement scenarios. Through multiple iterations, the policy progressively learns recovery strategies for disengagement states, thereby mitigating the open-loop gap. Experiments on the closed-loop Bench2Drive benchmark demonstrate our method's effectiveness compared with pure IL methods, with comprehensive ablations confirming the contribution of each component.

cs.RO

TARDis: Time Attenuated Representation Disentanglement for Incomplete Multi-Modal Tumor Segmentation and Classification

The accurate diagnosis and segmentation of tumors in contrast-enhanced Computed Tomography (CT) are fundamentally driven by the distinctive hemodynamic profiles of contrast agents over time. However, in real-world clinical practice, complete temporal dynamics are often hard to capture by strict radiation dose limits and inconsistent acquisition protocols across institutions, leading to a prevalent missing modality problem. Existing deep learning approaches typically treat missing phases as absent independent channels, ignoring the inherent temporal continuity of hemodynamics. In this work, we propose Time Attenuated Representation Disentanglement (TARDis), a novel physics-aware framework that redefines missing modalities as missing sample points on a continuous Time-Attenuation Curve. We first hypothesize that the latent feature can be disentangled into a time-invariant static component (anatomy) and a time-dependent dynamic component (perfusion). We achieve this via a dual-path architecture: a quantization-based path using a learnable embedding dictionary to extract consistent anatomical structures, and a probabilistic path using a Hemodynamic Conditional Variational Autoencoder to model dynamic enhancement conditioned on the estimated scan time. This design allows the network to infer missing hemodynamic features by sampling from the learned latent distribution. Extensive experiments on a large-scale multi-modal private abdominal CT dataset (2,282 patients) and two public datasets demonstrate that TARDis significantly outperforms state-of-the-art incomplete modality frameworks. Notably, our method maintains robust diagnostic performance even in extreme data-sparsity scenarios, highlighting its potential for reducing radiation exposure while maintaining diagnostic precision.

cs.CV

Reconfigurable Quantum Instruction Set Computers for High Performance Attainable on Hardware

The performance of current quantum hardware is severely limited. While expanding the quantum ISA with high-fidelity, expressive basis gates is a key path forward, it imposes significant gate calibration overhead and complicates compiler optimization. As a result, even though more powerful ISAs have been designed, their use remains largely conceptual rather than practical. To move beyond these hurdles, we introduce the concept of "reconfigurable quantum instruction set computers" (ReQISC), which incorporates: (1) a unified microarchitecture capable of directly implementing arbitrary 2Q gates equivalently, i.e., SU(4) modulo 1Q rotations, with theoretically optimal gate durations given any 2Q coupling Hamiltonians; (2) a compilation framework tailored to ReQISC primitives for end-to-end synthesis and optimization, comprising a program-aware pass that refines high-level representations, a program-agnostic pass for aggressive circuit-level optimization, and an SU(4)-aware routing pass that minimizes hardware mapping overhead. We detail the hardware implementation to demonstrate the feasibility, in terms of both pulse control and calibration of this superior gate scheme on realistic hardware. By leveraging the expressivity of SU(4) and the time minimality realized by the underlying microarchitecture, the SU(4)-based ISA achieves remarkable performance, with a 4.97-fold reduction in average pulse duration to implement arbitrary 2Q gates, compared to the usual CNOT/CZ scheme on mainstream flux-tunable transmons. Supported by the end-to-end compiler, ReQISC outperforms the conventional CNOT-ISA, SOTA compiler, and pulse implementation counterparts, in significantly reducing 2Q gate counts, circuit depth, pulse duration, qubit mapping overhead, and program fidelity losses. For the first time, ReQISC makes the theoretical benefits of continuous ISAs practically feasible.

quant-ph

Unifying Qubit Routing Across Diverse Quantum ISAs via Canonical Representation

Qubit mapping/routing is a critical stage in compilation for both near-term and fault-tolerant quantum computers, yet existing scalable methods typically impose several times the routing overhead in terms of circuit depth or duration. This inefficiency stems from a fundamental disconnect: compilers rely on an abstract routing model (e.g., three-CX-unrolled SWAP insertion) that completely ignores the idiosyncrasies of native gates supported by physical devices. Recent hardware breakthroughs have enabled high-precision implementations of diverse instruction set architectures (ISAs) beyond standard CX-based gates. Advanced ISAs involving gates such as $\mathrm{\sqrt{iSWAP}}$ and $\mathrm{ZZ}(\theta)$ gates offer superior circuit synthesis capabilities and can be realized with higher fidelities. However, systematic compiler optimization strategies tailored to these advanced ISAs are lacking. To address this, we propose Canopus, a unified qubit mapping/routing framework applicable to diverse quantum ISAs. Built upon the canonical representation of two-qubit gates, Canopus centers on qubit routing to perform deep co-optimization in an ISA-aware approach. Canopus leverages the two-qubit canonical representation and the monodromy polytope theory to model the synthesis cost for more intelligent SWAP insertion during qubit routing. We also formalize the commutation relations between two-qubit gates through the canonical form, providing a generalized approach to commutativity-based optimization. Experiments show that Canopus consistently reduces routing overhead by 15%-35% compared to state-of-the-art methods across various backend ISAs and device topologies. More broadly, this work establishes a coherent method for co-exploration of program patterns, quantum ISAs, and hardware topologies, yielding concrete guidelines for hardware-software co-design.

quant-ph

Quantum Nonlocality under Latency Constraints

Bell inequalities are bounds on the correlations between different parties obeying a local hidden variable theory. Here, "local" refers to spacetime locality: the parties cannot communicate their inputs because they must produce their outputs faster than the speed-of-light delay between them. In other words, the parties must satisfy a certain latency constraint. In this work, we explicitly incorporate spacetime locality into the formulation of Bell inequalities by imposing such a latency constraint. When the latency constraint is sufficiently tight such that no parties can communicate, this becomes a standard Bell scenario. When the latency constraint is relaxed such that a subset of the parties can communicate, we no longer have a Bell scenario, but we can again find a divide between classical and quantum behaviors. Hence, we observe that the classical-quantum gap should actually be a function of time. To study these more general scenarios, we introduce the mathematical framework of latency-constrained games, which models time-evolving input and output processes for spatially separated parties subject to finite communication speeds. This framework allows us to systematically study the weirdness of quantum mechanics in the "low-latency regime" where the speed-of-light delay is non-negligible. Latency-constrained games can describe real-time decision-making in real-world settings that are latency-sensitive, such as high-frequency trading and distributed systems, and can reveal the utility of quantum correlations in these settings.

quant-ph

PHOENIX: Pauli-Based High-Level Optimization Engine for Instruction Execution on NISQ Devices

Variational quantum algorithms (VQA) based on Hamiltonian simulation represent a specialized class of quantum programs well-suited for near-term quantum computing applications due to its modest resource requirements in terms of qubits and circuit depth. Unlike the conventional single-qubit (1Q) and two-qubit (2Q) gate sequence representation, Hamiltonian simulation programs are essentially composed of disciplined subroutines known as Pauli exponentiations (Pauli strings with coefficients) that are variably arranged. To capitalize on these distinct program features, this study introduces PHOENIX, a highly effective compilation framework that primarily operates at the high-level Pauli-based intermediate representation (IR) for generic Hamiltonian simulation programs. PHOENIX exploits global program optimization opportunities to the greatest extent, compared to existing SOTA methods despite some of them also utilizing similar IRs. Experimental results demonstrate that PHOENIX outperforms SOTA VQA compilers across diverse program categories, backend ISAs, and hardware topologies.

quant-ph

VOILA: Complexity-Aware Universal Segmentation of CT images by Voxel Interacting with Language

Satisfactory progress has been achieved recently in universal segmentation of CT images. Following the success of vision-language methods, there is a growing trend towards utilizing text prompts and contrastive learning to develop universal segmentation models. However, there exists a significant imbalance in information density between 3D images and text prompts. Moreover, the standard fully connected layer segmentation approach faces significant challenges in handling multiple classes and exhibits poor generalizability. To address these challenges, we propose the VOxel Interacting with LAnguage method (VOILA) for universal CT image segmentation. Initially, we align voxels and language into a shared representation space and classify voxels on the basis of cosine similarity. Subsequently, we develop the Voxel-Language Interaction framework to mitigate the impact of class imbalance caused by foreground-background discrepancies and variations in target volumes. Furthermore, a Complexity-Aware Sampling method is proposed to focus on region hard to segment, achieved by generating pseudo-heatmaps from a trainable Gaussian mixture distribution. Our results indicate the proposed VOILA is capable to achieve improved performance with reduced parameters and computational cost during training. Furthermore, it demonstrates significant generalizability across diverse datasets without additional fine-tuning.

cs.CV

Coordinating Decisions via Quantum Telepathy

Quantum telepathy is the phenomenon where two non-communicating parties can exhibit correlated behaviors that are impossible to achieve using classical resources. This is also known as Bell inequality violation and is made possible by quantum entanglement. In this work, we present a conceptual framework for applying quantum telepathy to real-world problems. In general, the problems involve multiple parties making local observations that need to coordinate their decisions but are unable to communicate. We argue this inability is actually quite prevalent in the modern era where the decision-making timescales of computer processors are so short that the speed of light delay is appreciable in comparison. We highlight the example of high-frequency trading (HFT), where trades are made at microsecond timescales, but the speed of light delay between different stock exchanges are on the order of 100 microseconds to 10 milliseconds. Due to the maturity of Bell inequality violation experiments, experimental realization of quantum telepathy schemes that can attain a quantum advantage for real-world problems $\textit{is already almost immediately possible}$. We demonstrate this by conducting a case study for a concrete HFT scenario that gives rise to a generalization of the CHSH game and evaluate different possible physical implementations for achieving a quantum advantage. It is well known that Bell inequality violation is a rigorous mathematical proof of a quantum advantage over any classical strategy and does not need any complexity-theoretic assumptions such as $\text{BQP}\neq\text{BPP}$. Moreover, fault tolerance is not necessary to realize a quantum advantage: for example, violating the CHSH inequality only requires single-qubit operations on two entangled physical qubits.

quant-ph

One Gate Scheme to Rule Them All: Introducing a Complex Yet Reduced Instruction Set for Quantum Computing

The design and architecture of a quantum instruction set are paramount to the performance of a quantum computer. This work introduces a gate scheme for qubits with $XX+YY$ coupling that directly and efficiently realizes any two-qubit gate up to single-qubit gates. First, this scheme enables high-fidelity execution of quantum operations and achieves minimum possible gate times. Second, since the scheme spans the entire $\textbf{SU}(4)$ group of two-qubit gates, we can use it to attain the optimal two-qubit gate count for algorithm implementation. These two advantages in synergy give rise to a quantum Complex yet Reduced Instruction Set Computer (CRISC). Though the gate scheme is compact, it supports a comprehensive array of quantum operations. This may seem paradoxical but is realizable due to the fundamental differences between quantum and classical computer architectures. Using our gate scheme, we observe marked improvements across various applications, including generic $n$-qubit gate synthesis, quantum volume, and qubit routing. Furthermore, the proposed scheme also realizes a gate locally equivalent to the commonly used CNOT gate with a gate time of $\frac{\pi}{2g}$, where $g$ is the two-qubit coupling. The AshN scheme is also completely impervious to $ZZ$ error, the main coherent error in transversely coupled systems, as the control parameters implementing the gates can be easily adjusted to take the $ZZ$ term into account.

quant-ph

A Classical Architecture For Digital Quantum Computers

Scaling bottlenecks the making of digital quantum computers, posing challenges from both the quantum and the classical components. We present a classical architecture to cope with a comprehensive list of the latter challenges {\em all at once}, and implement it fully in an end-to-end system by integrating a multi-core RISC-V CPU with our in-house control electronics. Our architecture enables scalable, high-precision control of large quantum processors and accommodates evolving requirements of quantum hardware. A central feature is a microarchitecture executing quantum operations in parallel on arbitrary predefined qubit groups. Another key feature is a reconfigurable quantum instruction set that supports easy qubit re-grouping and instructions extensions. As a demonstration, we implement the widely-studied surface code quantum computing workflow, which is instructive for being demanding on both the controllers and the integrated classical computation. Our design, for the first time, reduces instruction issuing and transmission costs to constants, which do not scale with the number of qubits, without adding any overheads in decoding or dispatching. Rather than relying on specialized hardware for syndrome decoding, our system uses a dedicated multi-core CPU for both qubit control and classical computation, including syndrome decoding. This simplifies the system design and facilitates load-balancing between the quantum and classical components. We implement recent proposals as decoding firmware on a RISC-V system-on-chip (SoC) that parallelizes general inner decoders. By using our in-house Union-Find and PyMatching 2 implementations, we can achieve unprecedented decoding capabilities of up to distances 47 and 67 with the currently available SoCs, under realistic and optimistic assumptions of physical error rate $p=0.001 and p=0.0001, respectively, all in just 1 \textmu s.

quant-ph

Linear Cross Entropy Benchmarking with Clifford Circuits

With the advent of quantum processors exceeding $100$ qubits and the high engineering complexities involved, there is a need for holistically benchmarking the processor to have quality assurance. Linear cross-entropy benchmarking (XEB) has been used extensively for systems with $50$ or more qubits but is fundamentally limited in scale due to the exponentially large computational resources required for classical simulation. In this work we propose conducting linear XEB with Clifford circuits, a scheme we call Clifford XEB. Since Clifford circuits can be simulated in polynomial time, Clifford XEB can be scaled to much larger systems. To validate this claim, we run numerical simulations for particular classes of Clifford circuits with noise and observe exponential decays. When noise levels are low, the decay rates are well-correlated with the noise of each cycle assuming a digital error model. We perform simulations of systems up to 1,225 qubits, where the classical processing task can be easily dealt with by a workstation. Furthermore, using the theoretical guarantees in Chen et al. (arXiv:2203.12703), we prove that Clifford XEB with our proposed Clifford circuits must yield exponential decays under a general error model for sufficiently low errors. Our theoretical results explain some of the phenomena observed in the simulations and shed light on the behavior of general linear XEB experiments.

quant-ph

TrajGen: Generating Realistic and Diverse Trajectories with Reactive and Feasible Agent Behaviors for Autonomous Driving

Realistic and diverse simulation scenarios with reactive and feasible agent behaviors can be used for validation and verification of self-driving system performance without relying on expensive and time-consuming real-world testing. Existing simulators rely on heuristic-based behavior models for background vehicles, which cannot capture the complex interactive behaviors in real-world scenarios. To bridge the gap between simulation and the real world, we propose TrajGen, a two-stage trajectory generation framework, which can capture more realistic behaviors directly from human demonstration. In particular, TrajGen consists of the multi-modal trajectory prediction stage and the reinforcement learning based trajectory modification stage. In the first stage, we propose a novel auxiliary RouteLoss for the trajectory prediction model to generate multi-modal diverse trajectories in the drivable area. In the second stage, reinforcement learning is used to track the predicted trajectories while avoiding collisions, which can improve the feasibility of generated trajectories. In addition, we develop a data-driven simulator I-Sim that can be used to train reinforcement learning models in parallel based on naturalistic driving data. The vehicle model in I-Sim can guarantee that the generated trajectories by TrajGen satisfy vehicle kinematic constraints. Finally, we give comprehensive metrics to evaluate generated trajectories for simulation scenarios, which shows that TrajGen outperforms either trajectory prediction or inverse reinforcement learning in terms of fidelity, reactivity, feasibility, and diversity.

cs.RO

Randomized Benchmarking Beyond Groups

Randomized benchmarking (RB) is the gold standard for experimentally evaluating the quality of quantum operations. The current framework for RB is centered on groups and their representations, but this can be problematic. For example, Clifford circuits need up to $O(n^2)$ gates, and thus Clifford RB cannot scale to larger devices. Attempts to remedy this include new schemes such as linear cross-entropy benchmarking (XEB), cycle benchmarking, and non-uniform RB, but they do not fall within the group-based RB framework. In this work, we formulate the \emph{universal randomized benchmarking (URB) framework} which does away with the group structure and also replaces the recovery gate plus measurement component with a general ``post-processing'' POVM. Not only does this framework cover most of the existing benchmarking schemes, but it also gives the language for and helps inspire the formulation of new schemes. We specifically consider a class of URB schemes called \emph{twirling schemes}. For twirling schemes, the post-processing POVM approximately factorizes into an intermediate channel, inverting maps, and a final measurement. This leads us to study the twirling map corresponding to the gate ensemble specified by the scheme. We prove that if this twirling map is strictly within unit distance of the Haar twirling map in induced diamond norm, the probability of measurement as a function of gate length is a single exponential decay up to small error terms. The core technical tool we use is the matrix perturbation theory of linear operators on quantum channels.

quant-ph

Fluxonium: an alternative qubit platform for high-fidelity operations

Superconducting qubits provide a promising path toward building large-scale quantum computers. The simple and robust transmon qubit has been the leading platform, achieving multiple milestones. However, fault-tolerant quantum computing calls for qubit operations at error rates significantly lower than those exhibited in the state of the art. Consequently, alternative superconducting qubits with better error protection have attracted increasing interest. Among them, fluxonium is a particularly promising candidate, featuring large anharmonicity and long coherence times. Here, we engineer a fluxonium-based quantum processor that integrates high qubit-coherence, fast frequency-tunability, and individual-qubit addressability for reset, readout, and gates. With simple and fast gate schemes, we achieve an average single-qubit gate fidelity of 99.97% and a two-qubit gate fidelity of up to 99.72%. This performance is comparable to the highest values reported in the literature of superconducting circuits. Thus our work, for the first time within the realm of superconducting qubits, reveals an approach toward fault-tolerant quantum computing that is alternative and competitive to the transmon system.

quant-ph

Quantum Instruction Set Design for Performance

A quantum instruction set is where quantum hardware and software meet. We develop new characterization and compilation techniques for non-Clifford gates to accurately evaluate different quantum instruction set designs. We specifically apply them to our fluxonium processor that supports mainstream instruction $\mathrm{iSWAP}$ by calibrating and characterizing its square root $\mathrm{SQiSW}$. We measure a gate fidelity of up to $99.72\%$ with an average of $99.31\%$ and realize Haar random two-qubit gates using $\mathrm{SQiSW}$ with an average fidelity of $96.38\%$. This is an average error reduction of $41\%$ for the former and a $50\%$ reduction for the latter compared to using $\mathrm{iSWAP}$ on the same processor. This shows designing the quantum instruction set consisting of $\mathrm{SQiSW}$ and single-qubit gates on such platforms leads to a performance boost at almost no cost.

quant-ph