SearcharxivSearch

arXiv subjects

Seungwoo Choi

Publications and source records attributed to Seungwoo Choi.

10 recordsLinked to original sources

SpiderLS: Leveraging Full ZX Reduction for Lattice Surgery Compilation

Lattice surgery compilation plays a central role in translating fault-tolerant quantum programs into efficient surface code realizations, where both spatial and temporal resources directly determine the cost of execution. Recent work has demonstrated the benefits of using ZX-diagrams as an intermediate representation for lattice surgery compilation, enabling semantics-preserving transformations that reduce spacetime cost. However, existing compilation restricts ZX reduction to preserve diagram structures that can be directly embedded as lattice surgery junctions. We present SpiderLS, which extends prior approach by leveraging full ZX reduction. To translate the resulting diagram into executable lattice surgery operations, SpiderLS applies a sequence of compiler passes that derives an execution order, generates target code by grouping compatible interactions into multi-target operations, and lowers the target code to Pauli-product measurements. The resulting explicit patch and Pauli-boundary requirements guide logical scheduling and structure-aware spacetime routing. Across representative algorithmic and random workloads, SpiderLS achieves average reductions of 49.2% in spacetime volume and 99.8% in compilation time compared with the prior ZX-based compiler.

quant-ph

A Cyclic Layerwise QAOA Training

The quantum approximate optimization algorithm (QAOA) is a hybrid quantum-classical algorithm for solving combinatorial optimization problems. Multi-angle QAOA (MA-QAOA), which assigns independent parameters to each Hamiltonian operator term, achieves superior approximation performance even with fewer layers than standard QAOA. Unfortunately, this increased expressibility can raise the classical computational cost due to a greater number of parameters. The recently proposed Layerwise MA-QAOA (LMA-QAOA) reduces this overhead by training one layer at a time, but it may suffer from obtaining the precise solution due to the previously fixed parameters. This work addresses two questions for efficient MA-QAOA training: (i) What is the optimal granularity for parameter updates per epoch, and (ii) How can we get precise final cost function results while only partially updating the parameters per epoch? Despite the benefit of reducing the parameters that update per epoch can reduce the classical computation overhead, too fine or coarse a granularity of Hamiltonian update can degrade the MA-QAOA training efficiency. We find that optimizing one complete layer per epoch is an efficient granularity. Moreover, selectively retraining each layer by tracking gradient variations can achieve a final cost function equivalent to the standard MA-QAOA while lowering the parameter update overhead. Based on these insights, we propose Orbit-QAOA, which cyclically revisits layers and selectively freezes stabilized parameters. Across diverse graph benchmarks, Orbit-QAOA reduces training steps by up to 81.8%, reduces approximation ratio error by up to 72x compared to the unified stop condition-applied enhanced LMA-QAOA, and achieves equivalent approximation performance compared to the standard MA-QAOA.

quant-ph

Mantra: Rewriting Quantum Programs to Minimize Trap-Movements for Zoned Rydberg Atom Arrays

A zoned neutral atom architecture achieves exceptional fidelity by segregating the execution spaces of 1- and 2-qubit gates, being a promising candidate for high-accuracy quantum systems. Unfortunately, naively applying programs designed for static qubit topologies to zoned architectures may result in most execution time being consumed by inter-zone travels of atoms. To address this, we introduce Mantra (Minimizing trAp movemeNts for aTom aRray Architectures), which rewrites quantum programs to reduce the interleaving of single- and two-qubit gates. Mantra incorporates three strategies: (i) a fountain-shaped controlled-Z (CZ) chain, (ii) ZZ-interaction protocol without a 1-qubit gate, and (iii) preemptive gate scheduling. Mantra reduces inter-zone movements by 68%, physical gate counts by 35%, and improves circuit fidelities by 17% compared to the standard executions.

quant-ph

PIMutation: Exploring the Potential of PIM Architecture for Quantum Circuit Simulation

Quantum circuit simulations are essential for the verification of quantum algorithms on behalf of real quantum devices. However, the memory requirements for such simulations grow exponentially with the number of qubits involved in quantum programs. Moreover, a substantial number of computations in quantum circuit simulations cause low locality data accesses, as they require extensive computations across the entire table of the full state vector. These characteristics lead to significant latency and energy overheads during data transfers between the CPU and main memory. Processing-in-Memory (PIM), which integrates computational logic near DRAM banks, could present a promising solution to address these challenges. In this paper, we introduce PIMutation (PIM framework for qUanTum circuit simulATION) for achieving fast and energy-efficient quantum circuit simulation. PIMutation is the first attempt to leverage UPMEM, a publicly available PIM-integrated DIMM, to implement quantum circuit simulations. PIMutation incorporates three optimization strategies to overcome the overhead of quantum circuit simulation using the real PIM system: (i) gate merging, (ii) row swapping, and (iii) vector partitioning. Our evaluations show that PIMutation achieves an average speedup of 2.99x and 16.51x with a reduction of energy of 25.23% and 75.29% over the QuEST simulator on CPU in 16- and 32-qubit benchmarks, respectively.

quant-ph

Balancing Thermal Relaxation Deviations of Near-Future Quantum Computing Results via Bit-Inverted Programs

One of the predominant causes of program distortion in the real quantum computing system may be attributed to the probability deviation caused by thermal relaxation. We introduce Barber (Balancing reAdout Results using Bit-invErted ciRcuits), a method designed to counteract the asymmetric thermal relaxation deviation and improve the reliability of near-term quantum programs. Barber collaborates with a bit-inverted quantum circuit, where the excited quantum state of qubits is assigned to the $\lvert 0 \rangle$ and the unexcited state to the $\lvert 1 \rangle$. In doing so, bit-inverted quantum circuits can experience thermal relaxation in the opposite direction compared to standard quantum circuits. Barber can effectively suppress the thermal relaxation deviation in program's readout results by selectively merging distributions from the standard and bit-inverted circuits.

quant-ph

Distribution-Adaptive Dynamic Shot Optimization for Variational Quantum Algorithms

Variational quantum algorithms (VQAs) have attracted remarkable interest over the past few years because of their potential computational advantages on near-term quantum devices. They leverage a hybrid approach that integrates classical and quantum computing resources to solve high-dimensional problems that are challenging for classical approaches alone. In the training process of variational circuits, constructing an accurate probability distribution for each epoch is not always necessary, creating opportunities to reduce computational costs through shot reduction. However, existing shot-allocation methods that capitalize on this potential often lack adaptive feedback or are tied to specific classical optimizers, which limits their applicability to common VQAs and broader optimization techniques. Our observations indicate that the information entropy of a quantum circuit's output distribution exhibits an approximately exponential relationship with the number of shots needed to achieve a target Hellinger distance. In this work, we propose a distribution-adaptive dynamic shot (DDS) framework that efficiently adjusts the number of shots per iteration in VQAs using the entropy distribution from the prior training epoch. Our results demonstrate that the DDS framework sustains inference accuracy while achieving a ~50% reduction in average shot count compared to fixed-shot training, and ~60% higher accuracy than recently proposed tiered shot allocation methods. Furthermore, in noisy simulations that reflect the error rates of actual IBM quantum systems, DDS achieves approximately a ~30% reduction in the total number of shots compared to the fixed-shot method with minimal degradation in accuracy, and offers about ~70% higher computational accuracy than tiered shot allocation methods.

quant-ph

Reliable Decision from Multiple Subtasks through Threshold Optimization: Content Moderation in the Wild

Social media platforms struggle to protect users from harmful content through content moderation. These platforms have recently leveraged machine learning models to cope with the vast amount of user-generated content daily. Since moderation policies vary depending on countries and types of products, it is common to train and deploy the models per policy. However, this approach is highly inefficient, especially when the policies change, requiring dataset re-labeling and model re-training on the shifted data distribution. To alleviate this cost inefficiency, social media platforms often employ third-party content moderation services that provide prediction scores of multiple subtasks, such as predicting the existence of underage personnel, rude gestures, or weapons, instead of directly providing final moderation decisions. However, making a reliable automated moderation decision from the prediction scores of the multiple subtasks for a specific target policy has not been widely explored yet. In this study, we formulate real-world scenarios of content moderation and introduce a simple yet effective threshold optimization method that searches the optimal thresholds of the multiple subtasks to make a reliable moderation decision in a cost-effective way. Extensive experiments demonstrate that our approach shows better performance in content moderation compared to existing threshold optimization methods and heuristics.

cs.LG

Attentron: Few-Shot Text-to-Speech Utilizing Attention-Based Variable-Length Embedding

On account of growing demands for personalization, the need for a so-called few-shot TTS system that clones speakers with only a few data is emerging. To address this issue, we propose Attentron, a few-shot TTS model that clones voices of speakers unseen during training. It introduces two special encoders, each serving different purposes. A fine-grained encoder extracts variable-length style information via an attention mechanism, and a coarse-grained encoder greatly stabilizes the speech synthesis, circumventing unintelligible gibberish even for synthesizing speech of unseen speakers. In addition, the model can scale out to an arbitrary number of reference audios to improve the quality of the synthesized speech. According to our experiments, including a human evaluation, the proposed model significantly outperforms state-of-the-art models when generating speech for unseen speakers in terms of speaker similarity and quality.

eess.AS

Temporal Convolution for Real-time Keyword Spotting on Mobile Devices

Keyword spotting (KWS) plays a critical role in enabling speech-based user interactions on smart devices. Recent developments in the field of deep learning have led to wide adoption of convolutional neural networks (CNNs) in KWS systems due to their exceptional accuracy and robustness. The main challenge faced by KWS systems is the trade-off between high accuracy and low latency. Unfortunately, there has been little quantitative analysis of the actual latency of KWS models on mobile devices. This is especially concerning since conventional convolution-based KWS approaches are known to require a large number of operations to attain an adequate level of performance. In this paper, we propose a temporal convolution for real-time KWS on mobile devices. Unlike most of the 2D convolution-based KWS approaches that require a deep architecture to fully capture both low- and high-frequency domains, we exploit temporal convolutions with a compact ResNet architecture. In Google Speech Command Dataset, we achieve more than \textbf{385x} speedup on Google Pixel 1 and surpass the accuracy compared to the state-of-the-art model. In addition, we release the implementation of the proposed and the baseline models including an end-to-end pipeline for training models and evaluating them on mobile devices.

cs.SD

Towards Real-Time Automatic Portrait Matting on Mobile Devices

We tackle the problem of automatic portrait matting on mobile devices. The proposed model is aimed at attaining real-time inference on mobile devices with minimal degradation of model performance. Our model MMNet, based on multi-branch dilated convolution with linear bottleneck blocks, outperforms the state-of-the-art model and is orders of magnitude faster. The model can be accelerated four times to attain 30 FPS on Xiaomi Mi 5 device with moderate increase in the gradient error. Under the same conditions, our model has an order of magnitude less number of parameters and is faster than Mobile DeepLabv3 while maintaining comparable performance. The accompanied implementation can be found at \url{https://github.com/hyperconnect/MMNet}.

cs.CV