SearcharxivSearch

arXiv subjects

Zhipeng Liang

Publications and source records attributed to Zhipeng Liang.

At least 19 recordsLinked to original sources

The four-dimensional Chamon code

Fracton models have attracted considerable interest as candidates for quantum memories because of their unconventional ground-state degeneracy (GSD) and restricted-mobility excitations. The four-dimensional (4D) Chamon code introduced in our previous work is constructed via the 4D XYZ product of two two-dimensional (2D) toric codes. Its GSD grows exponentially with the system size, similar to that of the three-dimensional (3D) Chamon code, suggesting that it may be regarded as a 4D generalization of the 3D Chamon code. However, the excitation properties of the 4D Chamon code have not been studied in depth, and a high-performance decoding strategy is still lacking. In this work, we first establish the correspondence between the algebraic structure of the 4D Chamon code and the 4D lattice, thereby characterizing the geometric distributions of qubits and stabilizers. Second, we show that the 4D Chamon code supports three types of restricted-mobility excitations analogous to those of the 3D Chamon code, further supporting its interpretation as a 4D generalization of the 3D Chamon code. Finally, we uncover two structural properties relevant to decoding: a hyperplane symmetry and a projection-induced 2D toric-code structure. By exploiting these properties, we develop a two-layer decoding strategy that decomposes the original decoding problem into multiple independent and parallelizable subproblems. Numerical simulations show that the proposed decoder substantially outperforms BP-OSD in decoding accuracy, demonstrating the benefit of incorporating the intrinsic geometric and algebraic structures of the code into decoder design.

quant-ph

MoWorld: A Flash World Model

The future of World Models depends not only on scaling model capability, but also on scaling practicality and inference efficiency. High-frame-rate inference enables responsive perception, planning, and control in real-world autonomous systems. To this end, we present MoWorld, a cost-effective yet high-performance Flash World Model with an end-to-end framework spanning data generation, pre-training, distillation, and efficient inference, enabling up to 50 FPS real-time interaction with cinematic visual quality without the need of high-end GPUs. To enable large-scale real-world deployment, MoWorld jointly optimizes model capability and cost throughout the entire development pipeline. Specifically, unlike existing approaches that primarily rely on large-scale video corpora, MoWorld is built upon a scalable 3D-native data engine accumulated from our large-scale 3D vision and generative modeling pipeline, enabling the efficient construction of geometrically consistent training data across diverse real-world and synthetic environments. Based on this foundation, a curriculum cross-frame pre-training strategy for stable and scalable World Model learning, an efficient denoising-step distillation algorithm to reduce diffusion training cost, and a mixed-precision parallel inference framework for low-cost real-time deployment. MoWorld is the first real-time interactive World Model built on the Neural Processing Unit (NPU) and can achieves up to 50 FPS in such the devices, enabling practical and efficient deployment at scale. Comprehensive evaluations demonstrate that MoWorld achieves leading performance; notably, its average inference cost is only 30\%-50\% of that of existing World Models, providing a practical foundation for large-scale real-world applications of World Models. We also demonstrate diverse applications of MoWorld.

cs.CV

An exploration of the noise sensitivity of the Shor's algorithm

Quantum algorithms face significant challenges due to qubit susceptibility to environmental noise, and quantum error correction typically requires prohibitive resource overhead. This paper proposes that quantum algorithms may possess inherent noise resilience characteristics that could reduce implementation barriers. We investigate Shor's algorithm by applying circuit-level noise models directly to the original algorithm circuit. Our findings reveal that Shor's algorithm demonstrates superior fault tolerance under Z noise compared to X and Y noise. Focusing on the modular exponentiation circuit which is the core component of the algorithm, we conduct fault-tolerant position statistics on circuits with bit lengths from 4 to 9. The results show that under Z noise, fault-tolerant positions grow with the same quartic polynomial order as potential error positions as the problem scale increases. In contrast, fault tolerance under X and Y noise exhibits a strong dependence on the composite number N and the parameter a. Based on these findings, we develop an extrapolation method predicting that the minimum probability of a correct output of the modular exponentiation circuit to factor 2048 bit integers under biased noise is approximately 1.417*{10}^{-17}.

quant-ph

Quantum XYZ cyclic codes for biased noise

In some quantum computing architectures, Pauli noise is highly biased. Tailoring Quantum error-correcting codes to the biased noise may benefit reducing the physical qubit overhead without reducing the logical error rate. In this paper, we propose a family of quantum XYZ cyclic codes, which are the only one family of quantum cyclic codes with code distance increasing with code length to our best knowledge and have good error-correcting performance against biased noise. Our simulation results show that the quantum XYZ cyclic codes have $50\%$ code-capacity thresholds for all three types of pure Pauli noise and around $13\%$ code-capacity threshold for depolarizing noise. In the finite-bias regime, when the noise is biased towards Pauli $Z$ errors with noise bias ratios $\eta_Z=1000$, the corresponding code-capacity threshold is around $49\%$. Besides, we show that to reach the same code distance, the physical qubit overhead of XYZ cyclic code is much less than that of the XZZX surface code.

quant-ph

Single-Trajectory Distributionally Robust Reinforcement Learning

To mitigate the limitation that the classical reinforcement learning (RL) framework heavily relies on identical training and test environments, Distributionally Robust RL (DRRL) has been proposed to enhance performance across a range of environments, possibly including unknown test environments. As a price for robustness gain, DRRL involves optimizing over a set of distributions, which is inherently more challenging than optimizing over a fixed distribution in the non-robust case. Existing DRRL algorithms are either model-based or fail to learn from a single sample trajectory. In this paper, we design a first fully model-free DRRL algorithm, called distributionally robust Q-learning with single trajectory (DRQ). We delicately design a multi-timescale framework to fully utilize each incrementally arriving sample and directly learn the optimal distributionally robust policy without modelling the environment, thus the algorithm can be trained along a single trajectory in a model-free fashion. Despite the algorithm's complexity, we provide asymptotic convergence guarantees by generalizing classical stochastic approximation tools. Comprehensive experimental results demonstrate the superior robustness and sample complexity of our proposed algorithm, compared to non-robust methods and other robust RL algorithms.

stat.ML

High-dimensional quantum XYZ product codes for biased noise

Three-dimensional (3D) quantum XYZ product can construct a class of non-CSS quantum codes by using three classical codes. However, there has been limited study on their error-correcting performance so far and whether this code construction can be generalized to higher dimension is an open question. In this paper, we first study the error-correcting performance of the 3D Chamon code, which is an instance of the 3D XYZ product of three repetition codes. Second, we show that the 3D XYZ product can be generalized to four dimension and propose four-dimensional (4D) XYZ product code construction, which constructs a class of non-CSS quantum codes by using either four classical codes or two CSS quantum codes. Compared with the 4D homological product, we show that the 4D XYZ product can construct non-CSS codes with higher code dimension or code distance. Third, we consider two instances of the 4D XYZ product, to which we refer as the 4D Chamon code and the 4D XYZ product concatenated code, respectively. Our simulation results show that, the 4D XYZ product can construct non-CSS codes with better error-correcting performance against Pauli-$Z$-biased noise than CSS codes constructed by the 4D homological product. Finally, we present the geometric arrangement of the 4D Chamon code within a 4D cubic lattice, demonstrating that it possesses two key characteristics of fracton models, which strongly suggest that it is a novel 4D fracton model.

quant-ph

Improved Belief Propagation Decoding Algorithms for Surface Codes

Quantum error correction is crucial for universal fault-tolerant quantum computing. Highly accurate and low-time-complexity decoding algorithms play an indispensable role in ensuring quantum error correction works effectively. Among existing decoding algorithms, belief propagation (BP) is notable for its nearly linear time complexity and general applicability to stabilizer codes. However, BP's decoding accuracy without post-processing is unsatisfactory in most situations. This article focuses on improving the decoding accuracy of BP over GF(4) for surface codes. Inspired by machine learning optimization techniques, we first propose Momentum-BP and AdaGrad-BP to reduce oscillations in message updating, breaking the trapping sets of surface codes. We further propose EWAInit-BP, which adaptively updates initial probabilities and provides a 1 to 3 orders of magnitude improvement over traditional BP for planar surface code, toric code, and XZZX surface code without any post-processing method, showing high decoding accuracy even under parallel scheduling. The theoretical $O(1)$ time complexity under parallel implementation and high accuracy of EWAInit-BP make it a promising candidate for high-precision real-time decoders.

quant-ph

Recursive expansion of Tanner graph: a method to construct stabilizer codes with high coding rate

Quantum stabilizer codes face the problem of low coding rate. In this article, following the idea of recursively expanding Tanner graph proposed in our previous work, we try to construct new stabilizer codes with high coding rate, and propose XZ-type Tanner-graph-recursive-expansion (XZ-TGRE) code and Tanner-graph-recursive-expansion hypergraph product (TGRE-HP) code. XZ-TGRE code have zero asymptotic coding rate, but its coding rate tends to zero extremely slowly with the growth of code length. Under the same code length, its coding rate is much higher than that of surface code. The coding rate of TGRE-HP is the constant 0.2, which is the highest constant coding rate of stabilizer codes to our best knowledge. We prove that the code distance of XZ-TGRE code scales as $O(log(N))$, and that of TGRE-HP code scales as $O(\log \sqrt{N})$, where $N$ is the code length. Moreover, the code capacity noise threshold of XZ-TGRE code is around 0.078, and that of TGRE-HP code is around 0.096. This articles shows that the idea of recursively expanding Tanner graph might have potential to construct quantum codes with good performance.

quant-ph

Hypergraph product code with 0.2 constant coding rate and high code capacity noise threshold

The low coding rate of quantum stabilizer codes results in formidable physical qubit overhead when realizing quantum error correcting in engineering. In this letter, we propose a new class of hypergraph-product code called TGRE-hypergraph-product code. This code has constant coding rate 0.2, which is the highest constant coding rate of quantum stabilizer codes to our best knowledge. We perform simulations to test the error correcting capability TGRE-hypergraph-product code and find their code capacity noise threshold in depolarizing noise channel is around 0.096.

quant-ph

Quantum polar stabilizer codes based on polarization of pure quantum channel don't work for quantum computing

Inspired by classical polar codes, whose coding rate can asymptotically achieve the Shannon capacity, researchers are trying to find its analogue in quantum information field, which are called quantum polar codes. However, no one has designed a quantum polar coding scheme which applies to quantum computing yet. There are two intuitions in previous research. The first is that directly converting classical polar coding circuits to quantum ones will produce polarization phenomenon of pure quantum channel, which has been proved in our previous work. The second is that based on this quantum polarization phenomenon one can design a quantum polar coding scheme that applies to quantum computing. There are several previous work following the second intuition, none of which has been verified by experiments. In this paper, we follow the second intuition and propose a more reasonable quantum polar stabilizer code construction algorithm than any previous ones by using the theory of stabilizer codes. Unfortunately, simulation experiments show that even the stabilizer codes obtained from this more reasonable construction algorithm don't work, which implies that the second intuition leads to a dead end. Based on the analysis on why the second intuition don't work, we provide a possible future direction of designing quantum stabilizer codes with high coding rate by borrowing the idea of classical polar codes. following this direction, we find a class of quantum stabilizer codes with coding rate 0.5 for pure Pauli X, Z and Y noise.

quant-ph

Determining the upper bound of code distance of quantum stabilizer codes through Monte Carlo method based on fully decoupled belief propagation

Code distance is an important parameter for quantum stabilizer codes (QSCs). Directly precisely computing it is an NP-complete problem. However, the upper bound of code distance can be computed by some efficient methods. In this paper, employing the idea of Monte Carlo method, we propose the algorithm of determining the upper bound of code distance of QSCs based on fully decoupled belief propagation. Our algorithm shows high precision - the upper bound of code distance determined by the algorithm of a variety of QSCs whose code distance is known is consistent with actual code distance. Besides, we explore the upper bound of logical X operators of Z-type Tanner-graph-recursive-expansion (Z-TGRE) code and Chamon code, which is a kind of XYZ product code constructed by three repetition codes. The former is consistent with the theoretical analysis, and the latter implies the code distance of XYZ product codes can very likely achieve $O(N^{2/3})$, which supports the conjecture of Leverrier et al..

quant-ph

Improved belief propagation decoding algorithm based on decoupling representation of Pauli operators for quantum LDPC codes

We propose a new method called decoupling representation to represent Pauli operators as vectors over $GF(2)$, based on which we propose partially decoupled belief propagation and fully decoupled belief propagation decoding algorithm for quantum low density parity-check codes. These two algorithms have the capability to deal with the correlations between the $X$ part and the $Z$ part of the vectors in symplectic representation, which are introduced by Pauli $Y$ errors. Hence, they can not only apply to CSS codes, but also to non-CSS codes. Under the assumption that there is no measurement error, compared with traditional belief propagation algorithm in symplectic representation over $GF(2)$, within the same number of iterations, the decoding accuracy of partially decoupled belief propagation and fully decoupled belief propagation algorithm is significantly improved in pure $Y$ noise and depolarizing noise, which supports that decoding algorithms of quantum error correcting codes might have better performance in decoupling representation than in symplectic representation. The impressive performance of fully decoupled belief propagation algorithm might promote the realization of quantum error correcting codes in engineering.

quant-ph

Reweighted Mixup for Subpopulation Shift

Subpopulation shift exists widely in many real-world applications, which refers to the training and test distributions that contain the same subpopulation groups but with different subpopulation proportions. Ignoring subpopulation shifts may lead to significant performance degradation and fairness concerns. Importance reweighting is a classical and effective way to handle the subpopulation shift. However, recent studies have recognized that most of these approaches fail to improve the performance especially when applied to over-parameterized neural networks which are capable of fitting any training samples. In this work, we propose a simple yet practical framework, called reweighted mixup (RMIX), to mitigate the overfitting issue in over-parameterized models by conducting importance weighting on the ''mixed'' samples. Benefiting from leveraging reweighting in mixup, RMIX allows the model to explore the vicinal space of minority samples more, thereby obtaining more robust model against subpopulation shift. When the subpopulation memberships are unknown, the training-trajectories-based uncertainty estimation is equipped in the proposed RMIX to flexibly characterize the subpopulation distribution. We also provide insightful theoretical analysis to verify that RMIX achieves better generalization bounds over prior works. Further, we conduct extensive empirical studies across a wide range of tasks to validate the effectiveness of the proposed method.

cs.LG

Distributionally Robust Offline Reinforcement Learning with Linear Function Approximation

Among the reasons hindering reinforcement learning (RL) applications to real-world problems, two factors are critical: limited data and the mismatch between the testing environment (real environment in which the policy is deployed) and the training environment (e.g., a simulator). This paper attempts to address these issues simultaneously with distributionally robust offline RL, where we learn a distributionally robust policy using historical data obtained from the source environment by optimizing against a worst-case perturbation thereof. In particular, we move beyond tabular settings and consider linear function approximation. More specifically, we consider two settings, one where the dataset is well-explored and the other where the dataset has sufficient coverage of the optimal policy. We propose two algorithms~-- one for each of the two settings~-- that achieve error bounds $\tilde{O}(d^{1/2}/N^{1/2})$ and $\tilde{O}(d^{3/2}/N^{1/2})$ respectively, where $d$ is the dimension in the linear function approximation and $N$ is the number of trajectories in the dataset. To the best of our knowledge, they provide the first non-asymptotic results of the sample complexity in this setting. Diverse experiments are conducted to demonstrate our theoretical findings, showing the superiority of our algorithm against the non-robust one.

cs.LG

UMIX: Improving Importance Weighting for Subpopulation Shift via Uncertainty-Aware Mixup

Subpopulation shift widely exists in many real-world machine learning applications, referring to the training and test distributions containing the same subpopulation groups but varying in subpopulation frequencies. Importance reweighting is a normal way to handle the subpopulation shift issue by imposing constant or adaptive sampling weights on each sample in the training dataset. However, some recent studies have recognized that most of these approaches fail to improve the performance over empirical risk minimization especially when applied to over-parameterized neural networks. In this work, we propose a simple yet practical framework, called uncertainty-aware mixup (UMIX), to mitigate the overfitting issue in over-parameterized models by reweighting the ''mixed'' samples according to the sample uncertainty. The training-trajectories-based uncertainty estimation is equipped in the proposed UMIX for each sample to flexibly characterize the subpopulation distribution. We also provide insightful theoretical analysis to verify that UMIX achieves better generalization bounds over prior works. Further, we conduct extensive empirical studies across a wide range of tasks to validate the effectiveness of our method both qualitatively and quantitatively. Code is available at https://github.com/TencentAILabHealthcare/UMIX.

cs.LG

Vertical Federated Linear Contextual Bandits

In this paper, we investigate a novel problem of building contextual bandits in the vertical federated setting, i.e., contextual information is vertically distributed over different departments. This problem remains largely unexplored in the research community. To this end, we carefully design a customized encryption scheme named orthogonal matrix-based mask mechanism(O3M) for encrypting local contextual information while avoiding expensive conventional cryptographic techniques. We further apply the mechanism to two commonly-used bandit algorithms, LinUCB and LinTS, and instantiate two practical protocols for online recommendation under the vertical federated setting. The proposed protocols can perfectly recover the service quality of centralized bandit algorithms while achieving a satisfactory runtime efficiency, which is theoretically proved and analyzed in this paper. By conducting extensive experiments on both synthetic and real-world datasets, we show the superiority of the proposed method in terms of privacy protection and recommendation performance.

cs.LG

On Private Online Convex Optimization: Optimal Algorithms in $\ell_p$-Geometry and High Dimensional Contextual Bandits

Differentially private (DP) stochastic convex optimization (SCO) is ubiquitous in trustworthy machine learning algorithm design. This paper studies the DP-SCO problem with streaming data sampled from a distribution and arrives sequentially. We also consider the continual release model where parameters related to private information are updated and released upon each new data, often known as the online algorithms. Despite that numerous algorithms have been developed to achieve the optimal excess risks in different $\ell_p$ norm geometries, yet none of the existing ones can be adapted to the streaming and continual release setting. To address such a challenge as the online convex optimization with privacy protection, we propose a private variant of online Frank-Wolfe algorithm with recursive gradients for variance reduction to update and reveal the parameters upon each data. Combined with the adaptive differential privacy analysis, our online algorithm achieves in linear time the optimal excess risk when $1<p\leq 2$ and the state-of-the-art excess risk meeting the non-private lower ones when $2<p\leq\infty$. Our algorithm can also be extended to the case $p=1$ to achieve nearly dimension-independent excess risk. While previous variance reduction results on recursive gradient have theoretical guarantee only in the independent and identically distributed sample setting, we establish such a guarantee in a non-stationary setting. To demonstrate the virtues of our method, we design the first DP algorithm for high-dimensional generalized linear bandits with logarithmic regret. Comparative experiments with a variety of DP-SCO and DP-Bandit algorithms exhibit the efficacy and utility of the proposed algorithms.

cs.LG

DRFLM: Distributionally Robust Federated Learning with Inter-client Noise via Local Mixup

Recently, federated learning has emerged as a promising approach for training a global model using data from multiple organizations without leaking their raw data. Nevertheless, directly applying federated learning to real-world tasks faces two challenges: (1) heterogeneity in the data among different organizations; and (2) data noises inside individual organizations. In this paper, we propose a general framework to solve the above two challenges simultaneously. Specifically, we propose using distributionally robust optimization to mitigate the negative effects caused by data heterogeneity paradigm to sample clients based on a learnable distribution at each iteration. Additionally, we observe that this optimization paradigm is easily affected by data noises inside local clients, which has a significant performance degradation in terms of global model prediction accuracy. To solve this problem, we propose to incorporate mixup techniques into the local training process of federated learning. We further provide comprehensive theoretical analysis including robustness analysis, convergence analysis, and generalization ability. Furthermore, we conduct empirical studies across different drug discovery tasks, such as ADMET property prediction and drug-target affinity prediction.

cs.LG