SearcharxivSearch

arXiv subjects

Pei Yuan

Publications and source records attributed to Pei Yuan.

15 recordsLinked to original sources

Quantum Circuit for General Unitary: Improved T-count via Block Flattening and Dilation

Synthesizing arbitrary $n$-qubit unitaries using as few non-Clifford gates as possible is a central problem in fault-tolerant quantum compilation. We present a Clifford+$T$ quantum circuit construction that approximately implements any classically specified unitary to within error $\epsilon$ and achieves a worst-case $T$-count with leading exponential scaling of $2^{5n/4}$ whenever $\log(1/\epsilon)=\operatorname{poly}(n)$. This improves upon the best previous $2^{4n/3}$ scaling. The key innovation lies in treating the target unitary as a single block-encoded object rather than a long product of simpler operations. A technique of block flattening controls the normalization while preserving an efficient implementation of the block encoding; subsequently, quantum singular value transformation maps its common singular value to one, thereby recovering the target unitary.

quant-ph

Efficient Depth--Ancilla Tradeoffs for Hamming Weight Computation and Symmetric Boolean Functions

Hamming weight computation maps an $n$-bit input to the number of ones it contains. It is a basic subroutine in quantum computing, and the core building block for symmetric Boolean functions, whose value depends only on the Hamming weight of the input. Moreover, symmetric Boolean functions are among the most common primitives in quantum computing. Efficient circuits for both problems are therefore important for the efficiency of many quantum algorithms. We study the depth-ancilla tradeoffs of Hamming weight computation under two qubit connectivity models, all-to-all and two-dimensional nearest-neighbor square grid (2D), in both the standard and dynamic circuit models. In the standard all-to-all model, we obtain depth $O(\log n)$ with a sublinear number of ancillas. In the standard 2D model, we give a circuit of depth $O(\sqrt n)$ with $O(\log^2 n)$ ancillas, and a matching lower bound showing that $\Theta(\sqrt n)$ is optimal. In both dynamic models, we obtain constant-depth circuits with $O(n^{1+\varepsilon}\operatorname{polylog}\,n)$ ancillary qubits for every fixed $\varepsilon>0$. All constructions give a smooth depth-ancilla tradeoff, and they also extend to arbitrary symmetric Boolean functions.

quant-ph

Optimal T Counts under Sparsity: from QROM to State Preparation and Block Encoding

Many quantum algorithms require coherent access to classical data, often modeled by quantum read-only memory (QROM). We initiate the study of the $T$ count of sparse QROM, in which only $s$ of the $2^n$ addresses store nonzero data. We prove asymptotically optimal $T$-count bounds $\Theta(\sqrt{sm} + \sqrt{sn})$ with square-root dependence on the support size $s$ and message length $m$. Our upper bounds use a multilevel hashing scheme, while our lower bounds reduce sparse QROM to state preparation and use counting arguments for adaptive Clifford+$T$ circuits. The lower bounds thus hold even when mid-circuit measurements and classically controlled operations are allowed. As applications, we obtain matching $T$-count bounds $\Theta(\sqrt{sn} + \sqrt{s\log(1/\varepsilon)} + \log(1/\varepsilon))$ for $s$-sparse state preparation and $\Theta( \sqrt{2^n sn} + \sqrt{2^n s\log(s/\varepsilon_{\mathrm{BE}})} + \log(s/\varepsilon_{\mathrm{BE}}))$ for block encoding of $s$-sparse matrices, where $\varepsilon$ and $\varepsilon_{\mathrm{BE}}$ are the precision of state preparation and block encoding, respectively.

quant-ph

The Dynamical Lie Algebra of QAOA-MaxCut on the Complete Graph

We give an analytical expression for the dynamical Lie algebra corresponding to the QAOA-MaxCut problem on complete graphs, and show that the variance of the associated loss function scales linearly in the number of qubits. This solves an open problem from [ASYZ26] and confirms that such systems do not exhibit barren plateaus. The proof is based on projecting the dynamical Lie algebra generators onto subspaces given by the Schur-Weyl duality between irreducible representations of the unitary and symmetric groups.

quant-ph

On the dynamical Lie algebras of quantum approximate optimization algorithms

Dynamical Lie algebras (DLAs) have emerged as a valuable tool in the study of parameterized quantum circuits, helping to characterize both their expressiveness and trainability. In particular, the absence or presence of barren plateaus (BPs) -- flat regions in parameter space that prevent the efficient training of variational quantum algorithms -- has recently been shown to be intimately related to quantities derived from the associated DLA. In this work, we investigate DLAs for the quantum approximate optimization algorithm (QAOA), one of the most studied variational quantum algorithms for solving graph MaxCut and other combinatorial optimization problems. While DLAs for QAOA circuits have been studied before, existing results have either been based on numerical evidence, or else correspond to DLA generators specifically chosen to be universal for quantum computation on a subspace of states. We initiate an analytical study of barren plateaus and other statistics of QAOA algorithms, and give bounds on the dimensions of the corresponding DLAs and their centers for general graphs. We then focus on the $n$-vertex cycle and complete graphs. For the cycle graph we give an explicit basis, identify its decomposition into the direct sum of a $2$-dimensional center and a semisimple component isomorphic to $n-1$ copies of $su(2)$. We give an explicit basis for this isomorphism, and a closed-form expression for the variance of the cost function, proving the absence of BPs. For the complete graph we prove that the dimension of the DLA is $O(n^3)$ and give an explicit basis for the DLA.

quant-ph

QAOA-MaxCut has barren plateaus for almost all graphs

The QAOA has been the subject of intense study over recent years, yet the corresponding Dynamical Lie Algebra (DLA)--a key indicator of the expressivity and trainability of VQAs--remains poorly understood beyond highly symmetric instances. An exponentially scaling DLA dimension is associated with the presence of so-called barren plateaus (BP) in the optimization landscape, which renders training intractable. In this work, we investigate the DLA of QAOA applied to the canonical MaxCut, for both weighted and unweighted graphs. For weighted graphs, we show that when the weights are drawn from a continuous distribution, the DLA dimension grows as $Θ(4^n)$ almost surely for all connected graphs except paths and cycles. In the more common unweighted setting, we show that asymptotically all but an exponentially vanishing fraction of graphs have $Θ(4^n)$ large DLA dimension. The entire simple Lie algebra decomposition of the corresponding DLAs is also identified, from which we prove that the variance of the loss function is $O(1/2^n)$, implying that QAOA on these weighted and unweighted graphs all suffers from BP. Moreover, we give explicit constructions for families of graphs whose DLAs have exponential dimension, including cases whose MaxCut is in $\mathsf P$. Our proof of the unweighted case is based on a number of splitting lemmas and DLA-freeness conditions that allow one to convert prohibitively complicated Lie algebraic problems into amenable graph theoretic problems. These form the basis for a new algorithm that computes such DLAs orders of magnitude faster than previous methods, reducing runtimes from days to seconds on standard hardware. We apply this algorithm to MQLib, a classical MaxCut benchmark suite covering over 3,500 instances with up to 53,130 vertices, and find that, ignoring edge weights, at least 75% of the instances possess a DLA of dimension at least $2^{128}$.

quant-ph

On generating direct powers of dynamical Lie algebras

The expressibility and trainability of parameterized quantum circuits has been shown to be intimately related to their associated dynamical Lie algebras (DLAs). From a quantum algorithm design perspective, given a set $A$ of DLA generators, two natural questions arise: (i) what is the DLA $\mathfrak{g}_{A}$ generated by ${A}$; and (ii) how does modifying the generator set lead to changes in the resulting DLA. While the first question has been the subject of significant attention, much less has been done regarding the second. In this work we focus on the second question, and show how modifying ${A}$ can result in a generator set ${A}'$ such that $\mathfrak{g}_{{A}'}\cong \bigoplus_{j=1}^{K}\mathfrak{g}_{A}$, for some $K \ge 1$. In other words, one generates the direct sum of $K$ copies of the original DLA. In particular, we give qubit- and parameter-efficient ways of achieving this, using only $\log K$ additional qubits, and only a constant factor increase in the number of DLA generators. For cyclic DLAs, which include Pauli DLAs and QAOA-MaxCut DLAs as special cases, this can be done with $\log K $ additional qubits and the same number of DLA generators as ${A}$.

quant-ph

Full Characterization of the Depth Overhead for Quantum Circuit Compilation with Arbitrary Qubit Connectivity Constraint

In some physical implementations of quantum computers, 2-qubit operations can be applied only on certain pairs of qubits. Compilation of a quantum circuit into one compliant to such qubit connectivity constraint results in an increase of circuit depth. Various compilation algorithms were studied, yet what this depth overhead is remains elusive. In this paper, we fully characterize the depth overhead by the routing number of the underlying constraint graph, a graph-theoretic measure which has been studied for 3 decades. We also give reduction algorithms between different graphs, which allow compilation for one graph to be transferred to one for another. These results, when combined with existing routing algorithms, give asymptotically optimal compilation for all commonly seen connectivity graphs in quantum computing.

quant-ph

Depth-Efficient Quantum Circuit Synthesis for Deterministic Dicke State Preparation

The $n$-qubit $k$-weight Dicke states $|D^n_k\rangle$, defined as the uniform superposition of all computational basis states with exactly $k$ qubits in state $|1\rangle$, form a basis of the symmetric subspace and represent an important class of entangled quantum states with broad applications in quantum computing. We propose deterministic quantum circuits for Dicke state preparation under two commonly seen qubit connectivity constraints: 1. All-to-all qubit connectivity: our circuit has depth $O(\log(k)\log(n/k)+k)$, which improves the previous best bound of $O(k\log(n/k))$. 2. Grid qubit connectivity ($(n_1\times n_2)$-grid, $n_1\le n_2$): (a) For $k\ge n_2/n_1$, we design a circuit with depth $O(k\log(n/k)+n_2)$, surpassing the prior $O(\sqrt{nk})$ bound. (b) For $k< n_2/n_1$, we design an optimal-depth circuit with depth $O(n_2)$. Furthermore, we establish the depth lower bounds of $Ω(\log(n))$ for all-to-all qubit connectivity and $Ω(n_2)$ for $(n_1\times n_2)$-grid connectivity constraints, demonstrating the near-optimality of our constructions.

quant-ph

Does qubit connectivity impact quantum circuit complexity?

Some physical implementation schemes of quantum computing can apply two-qubit gates only on certain pairs of qubits. These connectivity constraints are commonly viewed as a significant disadvantage. For example, compiling an unrestricted $n$-qubit quantum circuit to one with poor qubit connectivity, such as a 1D chain, usually results in a blowup of depth by $O(n^2)$ and size by $O(n)$. It is appealing to conjecture that this overhead is unavoidable -- a random circuit on $n$ qubits has $Θ(n)$ two-qubit gates in each layer and a constant fraction of them act on qubits separated by distance $Θ(n)$. While it is known that almost all $n$-qubit unitary operations need quantum circuits of $Ω(4^n/n)$ depth and $Ω(4^n)$ size to realize with all-to-all qubit connectivity, in this paper, we show that all $n$-qubit unitary operations can be implemented by quantum circuits of $O(4^n/n)$ depth and $O(4^n)$ size even under {1D chain} qubit connectivity constraint. We extend this result and investigate qubit connectivity in three directions. First, we consider more general connectivity graphs and show that the circuit size can always be made $O(4^n)$ as long as the graph is connected. For circuit depth, we study $d$-dimensional grids, complete $d$-ary trees and expander graphs, and show results similar to the 1D chain. Second, we consider the case when ancillary qubits are available. We show that, with ancilla, the circuit depth can be made polynomial, and the space-depth trade-off is not impaired by connectivity constraints unless we have exponentially many ancillary qubits. Third, we obtain nearly optimal results on special families of unitaries, including diagonal unitaries, 2-by-2 block diagonal unitaries, and Quantum State Preparation (QSP) unitaries, the last being a fundamental task used in many quantum algorithms for machine learning and linear algebra.

quant-ph

Optimal (controlled) quantum state preparation and improved unitary synthesis by quantum circuits with any number of ancillary qubits

As a cornerstone for many quantum linear algebraic and quantum machine learning algorithms, controlled quantum state preparation (CQSP) aims to provide the transformation of $|i\rangle |0^n\rangle \to |i\rangle |ψ_i\rangle $ for all $i\in \{0,1\}^k$ for the given $n$-qubit states $|ψ_i\rangle$. In this paper, we construct a quantum circuit for implementing CQSP, with depth $O\left(n+k+\frac{2^{n+k}}{n+k+m}\right)$ and size $O\left(2^{n+k}\right)$ for any given number $m$ of ancillary qubits. These bounds, which can also be viewed as a time-space tradeoff for the transformation, are \optimal for any integer parameters $m,k\ge 0$ and $n\ge 1$. When $k=0$, the problem becomes the canonical quantum state preparation (QSP) problem with ancillary qubits, which asks for efficient implementations of the transformation $|0^n\rangle|0^m\rangle \to |ψ\rangle |0^m\rangle$. This problem has many applications with many investigations, yet its circuit complexity remains open. Our construction completely solves this problem, pinning down its depth complexity to $Θ(n+2^{n}/(n+m))$ and its size complexity to $Θ(2^{n})$ for any $m$. Another fundamental problem, unitary synthesis, asks to implement a general $n$-qubit unitary by a quantum circuit. Previous work shows a lower bound of $Ω(n+4^n/(n+m))$ and an upper bound of $O(n2^n)$ for $m=Ω(2^n/n)$ ancillary qubits. In this paper, we quadratically shrink this gap by presenting a quantum circuit of the depth of $O\left(n2^{n/2}+\frac{n^{1/2}2^{3n/2}}{m^{1/2}}\right)$.

quant-ph

Asymptotically Optimal Circuit Depth for Quantum State Preparation and General Unitary Synthesis

The Quantum State Preparation problem aims to prepare an $n$-qubit quantum state $|ψ_v\rangle =\sum_{k=0}^{2^n-1}v_k|k\rangle$ from the initial state $|0\rangle^{\otimes n}$, for a given unit vector $v=(v_0,v_1,v_2,\ldots,v_{2^n-1})^T\in \mathbb{C}^{2^n}$ with $\|v\|_2 = 1$. The problem is of fundamental importance in quantum algorithm design, Hamiltonian simulation and quantum machine learning, yet its circuit depth and size complexity remain open when ancillary qubits are available. In this paper, we study efficient constructions of quantum circuits with $m$ ancillary qubits that can prepare $|ψ_v\rangle$ in depth $\tilde O\left(\frac{2^n}{m+n}+n\right)$and size $O(2^n)$, achieving the optimal value for both measures simultaneously. These results also imply a depth complexity of $Θ(4^n/(m+n))$ for quantum circuits implementing a general $n$-qubit unitary using $m = O(2^n/n)$ ancillary qubits. This resolves the depth complexity for circuits without ancillary qubits, and for circuits with exponentially many ancillary qubits, this gives a quadratic saving from $O(4^n)$ to $\tilde Θ(2^n)$. Our circuits are deterministic, prepare the state and carry out the unitary precisely, utilize the ancillary qubits tightly and the depths are optimal in a wide range of parameter regime. The results can be viewed as (optimal) time-space tradeoff bounds, which is not only theoretically interesting, but also practically relevant in the current trend that the number of qubits starts to take off, by showing a way to use a large number of qubits to compensate the short qubit lifetime.

quant-ph

Strong Quantum Nonlocality without Entanglement in Multipartite Quantum Systems

In this paper, we generalize the concept of strong quantum nonlocality from two aspects. Firstly in $\mathbb{C}^d\otimes\mathbb{C}^d\otimes\mathbb{C}^d$ quantum system, we present a construction of strongly nonlocal quantum states containing $6(d-1)^2$ orthogonal product states, which is one order of magnitude less than the number of basis states $d^3$. Secondly, we give the explicit form of strongly nonlocal orthogonal product basis in $\mathbb{C}^3\otimes \mathbb{C}^3\otimes \mathbb{C}^3\otimes \mathbb{C}^3$ quantum system, where four is the largest known number of subsystems in which there exists strong quantum nonlocality up to now. Both the two results positively answer the open problems in [Halder, \textit{et al.}, PRL, 122, 040403 (2019)], that is, there do exist and even smaller number of quantum states can demonstrate strong quantum nonlocality without entanglement.

quant-ph

A Quantum-inspired Classical Algorithm for Separable Non-negative Matrix Factorization

Non-negative Matrix Factorization (NMF) asks to decompose a (entry-wise) non-negative matrix into the product of two smaller-sized nonnegative matrices, which has been shown intractable in general. In order to overcome this issue, the separability assumption is introduced which assumes all data points are in a conical hull. This assumption makes NMF tractable and is widely used in text analysis and image processing, but still impractical for huge-scale datasets. In this paper, inspired by recent development on dequantizing techniques, we propose a new classical algorithm for separable NMF problem. Our new algorithm runs in polynomial time in the rank and logarithmic in the size of input matrices, which achieves an exponential speedup in the low-rank setting.

cs.DS

Exact quantum query complexity of weight decision problems via Chebyshev polynomials

The weight decision problem, which requires to determine the Hamming weight of a given binary string, is a natural and important problem, with applications in cryptanalysis, coding theory, fault-tolerant circuit design and so on. In particular, both Deutsch-Jozsa problem and Grover search problem can be interpreted as special cases of weight decision problems. In this work, we investigate the exact quantum query complexity of weight decision problems, where the quantum algorithm must always output the correct answer. More specifically we consider a partial Boolean function which distinguishes whether the Hamming weight of the length-$n$ input is $k$ or it is $l$. Our contribution includes both upper bounds and lower bounds for the precise number of queries. Furthermore, for most choices of $(\frac{k}{n},\frac{l}{n})$ and sufficiently large $n$, the gap between our upper and lower bounds is no more than one. To get the results, we first build the connection between Chebyshev polynomials and our problem, then determine all the boundary cases of $(\frac{k}{n},\frac{l}{n})$ with matching upper and lower bounds, and finally we generalize to other cases via a new \emph{quantum padding} technique. This quantum padding technique can be of independent interest in designing other quantum algorithms.

quant-ph