SearcharxivSearch

arXiv subjects

Xun Gao

Publications and source records attributed to Xun Gao.

At least 19 recordsLinked to original sources

Concurrence of Symmetry Breaking and Nonlocality Phase Transitions in Diffusion Models

Diffusion models undergo a phase transition in a critical time window during generation dynamics, with two complementary diagnoses of criticality. The symmetry breaking picture views the critical window as when trajectories bifurcate into different semantic minima of the energy landscape, whereas the nonlocality picture views the critical window as when local denoising fails. We study whether two notions of such phase transitions are concurrent in modern diffusion transformers. By evaluating the dynamics and outcomes of the generation trajectory, we observe a near-simultaneous occurrence of the non-locality and symmetry breaking critical times. Our work is the first to unify the two notions of phase transitions in practice: it provides a concrete diagnostic for when and why diffusion models rely on conditioning and global denoising, enabling principled evaluation of model efficiency and guiding the design of architectures and sampling schemes that avoid unnecessary computation.

cs.LG

Learning and Generating Mixed States Prepared by Shallow Channel Circuits

Learning quantum states from measurement data is a central problem in quantum information and computational complexity. In this work, we study the problem of learning to generate mixed states on a finite-dimensional lattice. Motivated by recent developments in mixed state phases of matter, we focus on arbitrary states in the trivial phase. A state belongs to the trivial phase if there exists a shallow preparation channel circuit under which local reversibility is preserved throughout the preparation. We prove that any mixed state in this class can be efficiently learned from measurement access alone. Specifically, given copies of an unknown trivial phase mixed state, our algorithm outputs a shallow local channel circuit that approximately generates this state in trace distance. The sample complexity and runtime are polynomial (or quasi-polynomial) in the number of qubits, assuming constant (or polylogarithmic) circuit depth and gate locality. Importantly, the learner is not given the original preparation circuit and relies only on its existence. Our results provide a structural foundation for quantum generative models based on shallow channel circuits. In the classical limit, our framework also inspires an efficient algorithm for classical diffusion models using only a polynomial overhead of training and generation.

quant-ph

Time complexity in preparing metrologically useful quantum states

We investigate the fundamental time complexity, as constrained by Lieb-Robinson bounds, for preparing entangled states useful in quantum metrology. We relate the minimum time to the Quantum Fisher Information ($F_Q$) for a system of $N$ quantum spins on a $d$-dimensional lattice with $1/r^\alpha$ interactions with $r$ being the distance between two interacting spins. We focus on states with $F_Q \sim N^{1+\gamma}$ where $\gamma \in (0,1]$, i.e., scaling from the standard quantum limit to the Heisenberg limit. For short-range interactions ($\alpha > 2d+1$), we prove the minimum time $t$ scales as $t \gtrsim L^\gamma$, where $L \sim N^{1/d}$. For long-range interactions, we find a hierarchy of possible speedups: $t \gtrsim L^{\gamma(\alpha-2d)}$ for $2d < \alpha < 2d+1$, $t \gtrsim \log L$ for $(2-\gamma)d < \alpha < 2d$, and $t$ may even vanish algebraically in $1/L$ for $\alpha < (2-\gamma)d$. These bounds extend to the minimum circuit depth required for state preparation, assuming two-qubit gate speeds scale as $1/r^\alpha$. We further show that these bounds are saturable, up to sub-polynomial corrections, for all $\alpha$ at the Heisenberg limit ($\gamma=1$) and for $\alpha > (2-\gamma)d$ when $\gamma<1$. Our results establish a benchmark for the time-optimality of protocols that prepare metrologically useful quantum states.

quant-ph

High-Distance Error-Correcting Codes for Fermion-to-Qubit Mappings in 2D and 3D

Quantum simulation of fermionic systems is a leading application of quantum computers. One promising approach is to represent fermions with qubits via fermion-to-qubit mappings. In this work, we present high-distance fermion-to-qubit stabilizer codes for simulating 2D and 3D fermionic systems. These codes achieve arbitrarily large code distances while keeping stabilizer weights constant. They also preserve locality by mapping local fermionic operators to local qubit operators at any fixed distance. Notably, our 3D construction is the first to simultaneously achieve high distance, constant stabilizer weights, and locality preservation. Our construction is based on concatenating a small-distance 2D or 3D fermion-to-qubit code with a high-distance fermionic color code. Together, these features provide a robust and scalable pathway to quantum simulation of fermionic systems.

quant-ph

Local Diffusion Models and Phases of Data Distributions

As a class of generative artificial intelligence frameworks inspired by statistical physics, diffusion models have shown extraordinary performance in synthesizing complicated data distributions through a denoising process gradually guided by score functions. Real-life data, like images, is often spatially structured in low-dimensional spaces. However, ordinary diffusion models ignore this local structure and learn spatially global score functions, which are often computationally expensive. In this work, motivated by recent advances in non-equilibrium statistical physics, we develop a generic framework for defining phases of data distributions and use it to analyze the locality requirements of denoisers in diffusion models. We define two distributions as belonging to the same data distribution phase if they can be mutually connected via spatially local operations such as local denoisers, along the same evolution path as the diffusion. We demonstrate that the reverse denoising process consists of an early trivial phase and a late data phase, sandwiching a rapid phase transition where local denoisers must fail. We further demonstrate that the performance of local denoisers is closely tied to spatial Markovianity, which provides an operational criterion for diagnosing such phase transitions. We validate this criterion through numerical experiments on real-world datasets. Our work suggests guidance for simpler and more efficient architectures of diffusion models: far from the phase transition point, we can use small local neural networks to compute the score function; global neural networks are only necessary around the narrow time interval of phase transitions. This result also opens up new directions for studying phases of data distributions, the broader science of generative artificial intelligence, and guiding the design of neural networks inspired by physics concepts.

cs.LG

DiLoCoX: A Low-Communication Large-Scale Training Framework for Decentralized Cluster

The distributed training of foundation models, particularly large language models (LLMs), demands a high level of communication. Consequently, it is highly dependent on a centralized cluster with fast and reliable interconnects. Can we conduct training on slow networks and thereby unleash the power of decentralized clusters when dealing with models exceeding 100 billion parameters? In this paper, we propose DiLoCoX, a low-communication large-scale decentralized cluster training framework. It combines Pipeline Parallelism with Dual Optimizer Policy, One-Step-Delay Overlap of Communication and Local Training, and an Adaptive Gradient Compression Scheme. This combination significantly improves the scale of parameters and the speed of model pre-training. We justify the benefits of one-step-delay overlap of communication and local training, as well as the adaptive gradient compression scheme, through a theoretical analysis of convergence. Empirically, we demonstrate that DiLoCoX is capable of pre-training a 107B foundation model over a 1Gbps network. Compared to vanilla AllReduce, DiLoCoX can achieve a 357x speedup in distributed training while maintaining negligible degradation in model convergence. To the best of our knowledge, this is the first decentralized training framework successfully applied to models with over 100 billion parameters.

cs.LG

2D Quon Language: Unifying Framework for Cliffords, Matchgates, and Beyond

Simulating generic quantum states and dynamics is practically intractable using classical computers. However, certain special classes -- namely Clifford and matchgate circuits -- permit efficient computation. They provide invaluable tools for studying many-body physics, quantum chemistry, and quantum computation. While both play foundational roles across multiple disciplines, the origins of their tractability seem disparate, and their relationship remain unclear. A deeper understanding of such tractable classes could expand their scope and enable a wide range of new applications. In this work, we make progress toward the unified understanding of the Clifford and matchgate -- these two classes are, in fact, distinct special cases of a single underlying structure. Specifically, we introduce the 2D Quon language, which combines Majorana worldlines with their underlying spacetime topology to diagrammatically represent quantum processes and tensor networks. In full generality, the 2D Quon language is universal -- capable of representing arbitrary quantum states, dynamics, or tensor networks -- yet they become especially powerful in describing Clifford and matchgate classes. Each class can be efficiently characterized in a visually recognizable manner using the Quon framework. This capability naturally gives rise to several families of efficiently computable tensor networks introduced in this work: punctured matchgates, hybrid Clifford-matchgate-MPS, and ansatze generated from factories of tractable networks. All of these exhibit high non-Cliffordness, high non-matchgateness, and large bipartite entanglement entropy. We discuss a range of applications of our approach, from recovering well-known results such as the Kramers-Wannier duality and the star-triangle relation of the Ising model, to enabling variational optimization with novel ansatz states.

quant-ph

High coherence fluxonium manufactured with a wafer-scale uniformity process

Fluxonium qubits are recognized for their high coherence times and high operation fidelities, attributed to their unique design incorporating a superinductor, which is typically implemented using an array of over 100 Josephson junctions. However, this complexity poses significant fabrication challenges, particularly in achieving high yield and junction uniformity with traditional methods. Here, we introduce an overlap process for Josephson junction fabrication that achieves nearly 100% yield and maintains uniformity across a 2-inch wafer with less than 5% variation for the phase slip junction and less than 2% for the entire junction array. We use a compact junction array design that achieves state-of-the-art dielectric loss tangents and flux noise levels, as confirmed by multiple devices. This enables fluxonium qubits to reach energy relaxation times exceeding 1 millisecond at the flux frustration point. This work paves the way for scalable high coherence fluxonium processors using CMOS-compatible processes, marking a significant step towards practical quantum computing.

quant-ph

FreeInv: Free Lunch for Improving DDIM Inversion

Naive DDIM inversion process usually suffers from a trajectory deviation issue, i.e., the latent trajectory during reconstruction deviates from the one during inversion. To alleviate this issue, previous methods either learn to mitigate the deviation or design cumbersome compensation strategy to reduce the mismatch error, exhibiting substantial time and computation cost. In this work, we present a nearly free-lunch method (named FreeInv) to address the issue more effectively and efficiently. In FreeInv, we randomly transform the latent representation and keep the transformation the same between the corresponding inversion and reconstruction time-step. It is motivated from a statistical perspective that an ensemble of DDIM inversion processes for multiple trajectories yields a smaller trajectory mismatch error on expectation. Moreover, through theoretical analysis and empirical study, we show that FreeInv performs an efficient ensemble of multiple trajectories. FreeInv can be freely integrated into existing inversion-based image and video editing techniques. Especially for inverting video sequences, it brings more significant fidelity and efficiency improvements. Comprehensive quantitative and qualitative evaluation on PIE benchmark and DAVIS dataset shows that FreeInv remarkably outperforms conventional DDIM inversion, and is competitive among previous state-of-the-art inversion methods, with superior computation efficiency.

cs.CV

Fault-tolerant compiling of classically hard IQP circuits on hypercubes

Realizing computationally complex quantum circuits in the presence of noise and imperfections is a challenging task. While fault-tolerant quantum computing provides a route to reducing noise, it requires a large overhead for generic algorithms. Here, we develop and analyze a hardware-efficient, fault-tolerant approach to realizing complex sampling circuits. We co-design the circuits with the appropriate quantum error correcting codes for efficient implementation in a reconfigurable neutral atom array architecture, constituting what we call a fault-tolerant compilation of the sampling algorithm. Specifically, we consider a family of $[[2^D , D, 2]]$ quantum error detecting codes whose transversal and permutation gate set can realize arbitrary degree-$D$ instantaneous quantum polynomial (IQP) circuits. Using native operations of the code and the atom array hardware, we compile a fault-tolerant and fast-scrambling family of such IQP circuits in a hypercube geometry, realized recently in the experiments by Bluvstein et al. [Nature 626, 7997 (2024)]. We develop a theory of second-moment properties of degree-$D$ IQP circuits for analyzing hardness and verification of random sampling by mapping to a statistical mechanics model. We provide evidence that sampling from hypercube IQP circuits is classically hard to simulate and analyze the linear cross-entropy benchmark (XEB) in comparison to the average fidelity. To realize a fully scalable approach, we first show that Bell sampling from degree-$4$ IQP circuits is classically intractable and can be efficiently validated. We further devise new families of $[[O(d^D),D,d]]$ color codes of increasing distance $d$, permitting exponential error suppression for transversal IQP sampling. Our results highlight fault-tolerant compiling as a powerful tool in co-designing algorithms with specific error-correcting codes and realistic hardware.

quant-ph

Velocity-comb modulation transfer spectroscopy

Sub-Doppler laser spectroscopy is a crucial technique for laser frequency stabilization, playing a significant role in atomic physics, precision measurement, and quantum communication. However, recent efforts to improve frequency stability appear to have reached a bottleneck, as they primarily focus on external technical approaches while neglecting the fundamental issue of low atomic utilization (< 1%), caused by only near-zero transverse velocity atoms involved in the transition. Here, we propose a velocity-comb modulation transfer spectroscopy (MTS) solution that takes advantage of the velocity-selective resonance effect of multi-frequency comb lasers to enhance the utilization of non-zero-velocity atoms. In the probe-pump configuration, each pair of counter-propagating lasers interacts with atoms from different transverse velocity-comb groups, independently contributing to the spectral amplitude and signal-to-noise ratio. Preliminary proof-of-principle results show that the frequency stability of the triple-frequency laser is optimized by nearly a factor of \sqrt{3} compared to the single-frequency laser, consistent with theoretical expectations. With more frequency comb components, MTS-stabilized lasers are expected to achieve order-of-magnitude breakthroughs in frequency stability, taking an important step toward next-generation compact optical clocks. This unique method can also be widely applied to any quantum system with a wide velocity distribution, inspiring innovative advances in numerous fields with a fresh perspective.

physics.atom-ph

Measuring Non-Gaussian Magic in Fermions: Convolution, Entropy, and the Violation of Wick's Theorem and the Matchgate Identity

Classically hard to simulate quantum states, or "magic states", are prerequisites to quantum advantage, highlighting an apparent separation between classically and quantumly tractable problems. Classically simulable states such as Clifford circuits on stabilizer states, free bosonic states, free fermions, and matchgate circuits are all in some sense Gaussian. While free bosons and fermions arise from quadratic Hamiltonians, recent works have demonstrated that bosonic and qudit systems converge to Gaussians and stabilizers under convolution. In this work, we similarly identify convolution for fermions and find efficient measures of non-Gaussian magic in pure fermionic states. We demonstrate that three natural notions for the Gaussification of a state, (1) the Gaussian state with the same covariance matrix, (2) the fixed point of convolution, and (3) the closest Gaussian in relative entropy, coincide by proving a central limit theorem for fermionic systems. We then utilize the violation of Wick's theorem and the matchgate identity to quantify non-Gaussian magic in addition to a SWAP test.

quant-ph

Generalization Error in Quantum Machine Learning in the Presence of Sampling Noise

Tackling output sampling noise due to finite shots of quantum measurement is an unavoidable challenge when extracting information in machine learning with physical systems. A technique called Eigentask Learning was developed recently as a framework for learning with infinite input training data in the presence of output sampling noise. In the work of Eigentask Learning, numerical evidence was presented that extracting low-noise contributions of features can practically improve performance for machine learning tasks, displaying robustness to overfitting and increasing generalization accuracy. However, it remains unsolved to quantitatively characterize generalization errors in situations where the training dataset is finite, while output sampling noise still exists. In this study, we use methodologies from statistical mechanics to calculate the training and generalization errors of a generic quantum machine learning system when the input training dataset and output measurement sampling shots are both finite. Our analytical findings, supported by numerical validation, offer solid justification that Eigentask Learning provides optimal learning in the sense of minimizing generalization errors.

quant-ph

A polynomial-time classical algorithm for noisy quantum circuits

We provide a polynomial-time classical algorithm for noisy quantum circuits. The algorithm computes the expectation value of any observable for any circuit, with a small average error over input states drawn from an ensemble (e.g. the computational basis). Our approach is based upon the intuition that noise exponentially damps non-local correlations relative to local correlations. This enables one to classically simulate a noisy quantum circuit by only keeping track of the dynamics of local quantum information. Our algorithm also enables sampling from the output distribution of a circuit in quasi-polynomial time, so long as the distribution anti-concentrates. A number of practical implications are discussed, including a fundamental limit on the efficacy of noise mitigation strategies: for constant noise rates, any quantum circuit for which error mitigation is efficient on most input states, is also classically simulable on most input states.

quant-ph

Large-scale quantum reservoir learning with an analog quantum computer

Quantum machine learning has gained considerable attention as quantum technology advances, presenting a promising approach for efficiently learning complex data patterns. Despite this promise, most contemporary quantum methods require significant resources for variational parameter optimization and face issues with vanishing gradients, leading to experiments that are either limited in scale or lack potential for quantum advantage. To address this, we develop a general-purpose, gradient-free, and scalable quantum reservoir learning algorithm that harnesses the quantum dynamics of neutral-atom analog quantum computers to process data. We experimentally implement the algorithm, achieving competitive performance across various categories of machine learning tasks, including binary and multi-class classification, as well as timeseries prediction. Effective and improving learning is observed with increasing system sizes of up to 108 qubits, demonstrating the largest quantum machine learning experiment to date. We further observe comparative quantum kernel advantage in learning tasks by constructing synthetic datasets based on the geometric differences between generated quantum and classical data kernels. Our findings demonstrate the potential of utilizing classically intractable quantum correlations for effective machine learning. We expect these results to stimulate further extensions to different quantum hardware and machine learning paradigms, including early fault-tolerant hardware and generative machine learning tasks.

quant-ph

Ultrafast adiabatic passages in ultrastrongly coupled light-matter systems

We have obtained the solutions of the multimode quantum Rabi model when all modes have identical frequencies $ω$, including dark states $|ϕ_K\rangle$ with at least $K$ $(K=1,2,3,\ldots)$ photons. Extended to the multiqubit case, they lie close to another dark state $\vert ψ\rangle$ with at most one photon in the spectrum. Taking advantages of such solutions, we find a linear and symmetry-protected adiabatic passage through $\vert ψ\rangle$ to fast generate arbitrary single-photon $M$-mode $W$ states $\vert W_M\rangle$ with exactly the same speed. The effective minimum energy gap during the adiabatic evolution is further enlarged to $0.63ω$ when Stark shifts are included, such that arbitrary $\vert W_M\rangle$ can be ultrafast generated in $1.55\times 2πω^{-1}$ with fidelity $99\%$, indepedent of $M$. This work reveals the existence of linear ultrafast adiabatic passages in light-matter systems.

quant-ph

Expressive Quantum Perceptrons for Quantum Neuromorphic Computing

Quantum neuromorphic computing (QNC) is a sub-field of quantum machine learning (QML) that capitalizes on inherent system dynamics. As a result, QNC can run on contemporary, noisy quantum hardware and is poised to realize challenging algorithms in the near term. One key issue in QNC is the characterization of the requisite dynamics for ensuring expressive quantum neuromorphic computation. We address this issue by proposing a building block for QNC architectures, what we call quantum perceptrons (QPs). Our proposed QPs compute based on the analog dynamics of interacting qubits with tunable coupling constants. We show that QPs are, with restricted resources, a quantum equivalent to the classical perceptron, a simple mathematical model for a neuron that is the building block of various machine learning architectures. Moreover, we show that QPs are theoretically capable of producing any unitary operation. Thus, QPs are computationally more expressive than their classical counterparts. As a result, QNC architectures built our of QPs are, theoretically, universal. We introduce a technique for mitigating barren plateaus in QPs called entanglement thinning. We demonstrate QPs' effectiveness by applying them to numerous QML problems, including calculating the inner products between quantum states, energy measurements, and time-reversal. Finally, we discuss potential implementations of QPs and how they can be used to build more complex QNC architectures.

quant-ph

Arbitrary Polynomial Separations in Trainable Quantum Machine Learning

Recent theoretical results in quantum machine learning have demonstrated a general trade-off between the expressive power of quantum neural networks (QNNs) and their trainability; as a corollary of these results, practical exponential separations in expressive power over classical machine learning models are believed to be infeasible as such QNNs take a time to train that is exponential in the model size. We here circumvent these negative results by constructing a hierarchy of efficiently trainable QNNs that exhibit unconditionally provable, polynomial memory separations of arbitrary constant degree over classical neural networks -- including state-of-the-art models, such as Transformers -- in performing a classical sequence modeling task. This construction is also computationally efficient, as each unit cell of the introduced class of QNNs only has constant gate complexity. We show that contextuality -- informally, a quantitative notion of semantic ambiguity -- is the source of the expressivity separation, suggesting that other learning tasks with this property may be a natural setting for the use of quantum learning algorithms.

quant-ph