SearcharxivSearch

arXiv subjects

Min Ye

Publications and source records attributed to Min Ye.

At least 19 recordsLinked to original sources

Real-time decoder for a MegaQuOp quantum computer using a single CPU

As quantum computers advance toward the regime of MegaQuOp machines executing millions of gates, a decoding system capable of real-time error correction in such a device will be crucial. Recent efforts have been focused on decoding an error-corrected memory or a small number of logical operations. Here we demonstrate an end-to-end real-time decoding stack for a universal fault-tolerant trapped-ion quantum computer architecture capable of decoding real workloads with millions of logical gates over hundreds of logical qubits. The complete pipeline, including on the fly detector error model generation, decoding of all logical qubits, logical operations, and magic-state factories, runs on a single CPU. We benchmark the decoder on practically relevant quantum applications spanning up to 408 logical qubits, and up to one million $T$ gates. Assuming a trapped-ion architecture with 1 to 5 ms cycle time, the decoding delay stretches the computation by less than $0.3\%$ at $p_{\mathrm{CNOT}}=10^{-4}$ and less than $12\%$ at $p_{\mathrm{CNOT}}=5\times 10^{-4}$ for all workloads studied. These results demonstrate real-time decoding at MegaQuOp scale on a single conventional CPU.

quant-ph

Queue-Aware Graph Reinforcement Learning for UAV-ISAC-Assisted Maritime Data Collection

This paper studies high-altitude platform (HAP)-assisted sparse cooperative integrated sensing and communication (ISAC) for UAV-enabled ocean monitoring. A fleet of rotary-wing UAVs senses drifting buoys, collects their monitoring data, and reports local posterior estimates to a HAP that performs fusion and sparse cooperation control. The model explicitly accounts for a spatially correlated sea-patch field, patch-aware buoy dynamics, RCS- and clutter-aware echo sensing, fused posterior Cram\'er-Rao bounds (PCRBs), and propulsion-energy-limited UAV mobility. The long-horizon objective is cast as a queue-weighted buffered-collection Markov decision process rather than instantaneous throughput, where each buoy maintains a backlog of buffered observations. The resulting long-horizon design is formulated as a mixed discrete-continuous problem with sensing, communication, mobility, safety, buffered-collection, and onboard-energy constraints. To address the combinatorial association component without replacing learning by a deterministic optimizer, we propose a structured feasible-association graph-MARL framework. A heterogeneous graph encoder produces candidate-edge logits, and a masked sequential b-matching policy samples legal UAV-buoy associations while exactly satisfying UAV-load and buoy-cluster constraints. A MAPPO-style training procedure, an independent queue-state value critic, and a consistency-verification protocol are then specified to support reproducible training. Simulation results on congested maritime scenarios show that the proposed policy improves the cumulative queue-weighted collection utility by about 106\% over the rate-driven deterministic decoder, maintains a large margin across sea-state sweeps and medium-to-heavy traffic loads, and transfers to larger networks without fine-tuning.

eess.SY

Sensing-Assisted Predictive Beamforming for UAV-Enabled Ocean Monitoring Networks

This paper investigates a sensing-assisted predictive beamforming framework for UAV--buoy maritime monitoring by explicitly accounting for wave-induced buoy dynamics and residual sea clutter. A frame-based UAV mission workflow is first established, where the UAV transmits integrated sensing and communication signals to acquire buoy echoes and to support subsequent uplink beam alignment. To characterize short-horizon buoy motion, a correlated-acceleration state-space model is developed by combining a Singer process for wave-driven excitation with a slowly varying current-drift term. Given the resulting nonlinear reflection, Doppler, and delay measurements, the posterior Fisher information matrix and the corresponding posterior Cram\'er--Rao bound (PCRB) are derived, and the predicted horizontal-position PCRB is adopted as the sensing metric. A per-frame worst-buoy design is then formulated to jointly optimize sensing power allocation and UAV position under uplink-rate, UAV-power, and mobility constraints. By exploiting a Schur-complement reformulation and a lagged successive convex approximation, the resulting subproblem is converted into a convex conic program with tractable complexity. Simulation results show that the proposed scheme maintains robust prediction and communication performance under denser buoy deployments and harsher sea conditions, and outperforms several baseline designs. In particular, the pronounced root mean square error (RMSE) degradation of the communication-only benchmark confirms that sensing-assisted state refinement is essential for accurate predictive beamforming in dynamic maritime environments. Compared with a full first-order Taylor expansion method, it achieves a more attractive performance--complexity tradeoff for online deployment.

eess.SP

Fault-Tolerant Quantum Computing with Trapped Ions: The Walking Cat Architecture

We propose a fault-tolerant quantum computer architecture for trapped-ion devices, which we call the walking cat architecture. Our blueprint includes a compiler, a detailed description of all the quantum error-correction protocols, a micro-architecture, a sufficiently fast decoder, and thorough simulations. The backbone of the architecture is a cat factory, producing cat states distributed throughout the machine, which are consumed to perform logical operations. The walking cat architecture is based entirely on a modern quantum error-correction approach called low-density parity-check (LDPC) codes. We identify promising instances of the walking cat architecture, such as (1) a simple architecture based on a single LDPC code, (2) a fast architecture based on fast logical gates relying on a [[70, 6, 9]] code, equipped with Clifford-frame tracking for any 6-qubit Clifford gate, and (3) a dense architecture based on a [[102, 22, 9]]] code encoding 22 logical qubits per memory block. Our dense architecture provides a design with 110 logical qubits executing about one million T gates per day using only 2,514 physical qubits. We estimate that the quantum Hamiltonian simulation of a Heisenberg model on 100 sites can be executed within one month with 10,000 physical qubits, including all shots required to achieve chemical accuracy, suggesting that such a device could enter the regime of classically intractable physics simulations. Our design relies on hardware components that have been experimentally demonstrated on small devices. We emphasize simplicity over hypothetical performance to facilitate the practical realization of this machine. Based on this approach, we believe that a fault-tolerant quantum computer with hundreds of logical qubits capable of running millions of logical gates can be built in the near term, providing a platform to explore a broad range of applications.

quant-ph

Beam search decoder for quantum LDPC codes

We propose a decoder for quantum low density parity check (LDPC) codes based on a beam search heuristic guided by belief propagation (BP). Our beam search decoder applies to all quantum LDPC codes and achieves different speed-accuracy tradeoffs by tuning its parameters such as the beam width. We perform numerical simulations under circuit level noise for the $[[144, 12, 12]]$ bivariate bicycle (BB) code at noise rate $p=10^{-3}$ to estimate the logical error rate and the 99.9 percentile runtime and we compare with the BP-OSD decoder which has been the default quantum LDPC decoder for the past six years. A variant of our beam search decoder with a beam width of 64 achieves a $17\times$ reduction in logical error rate. With a beam width of 8, we reach the same logical error rate as BP-OSD with a $26.2\times$ reduction in the 99.9 percentile runtime. We identify the beam search decoder with beam width of 32 as a promising candidate for trapped ion architectures because it achieves a $5.6\times$ reduction in logical error rate with a 99.9 percentile runtime per syndrome extraction round below 1ms at $p=5 \times10^{-4}$. Remarkably, this is achieved in software on a single core, without any parallelization or specialized hardware (FPGA, ASIC), suggesting one might only need three 32-core CPUs to decode a trapped ion quantum computer with 1000 logical qubits.

quant-ph

Correction of chain losses in trapped ion quantum computers

Neutral atom quantum computers and to a lesser extent trapped ions may suffer from atom loss. In this work, we investigate the impact of atom loss in long chains of trapped ions. Even though this is a relatively rare event, ion loss in long chains must be addressed because it destabilizes the entire chain resulting in the loss of all the qubits of the chain. We propose a solution to the chain loss problem based on (1) a quantum error correction code distributed over multiple long chains, (2) beacon qubits within each long chain to detect the loss of a chain, and (3) a decoder adapted to correct a combination of circuit faults and erasures after beacon qubits convert chain losses into erasures. We verify the chain loss correction capability of our scheme through circuit level simulations with a distributed $[[72,12,6]]$ BB code with beacon qubits.

quant-ph

Attending on Multilevel Structure of Proteins enables Accurate Prediction of Cold-Start Drug-Target Interactions

Cold-start drug-target interaction (DTI) prediction focuses on interaction between novel drugs and proteins. Previous methods typically learn transferable interaction patterns between structures of drug and proteins to tackle it. However, insight from proteomics suggest that protein have multi-level structures and they all influence the DTI. Existing works usually represent protein with only primary structures, limiting their ability to capture interactions involving higher-level structures. Inspired by this insight, we propose ColdDTI, a framework attending on protein multi-level structure for cold-start DTI prediction. We employ hierarchical attention mechanism to mine interaction between multi-level protein structures (from primary to quaternary) and drug structures at both local and global granularities. Then, we leverage mined interactions to fuse structure representations of different levels for final prediction. Our design captures biologically transferable priors, avoiding the risk of overfitting caused by excessive reliance on representation learning. Experiments on benchmark datasets demonstrate that ColdDTI consistently outperforms previous methods in cold-start settings.

cs.LG

Distributed fault-tolerant quantum memories over a 2xL array of qubit modules

We propose an architecture for a quantum memory distributed over a $2 \times L$ array of modules equipped with a cyclic shift implemented via flying qubits. The logical information is distributed across the first row of $L$ modules and quantum error correction is executed using ancilla modules on the second row equipped with a cyclic shift. This work proves that quantum LDPC codes such as BB codes can maintain their performance in a distributed setting while using solely one simple connector: a cyclic shift. We propose two strategies to perform quantum error correction on a $2 \times L$ module array: (i) The cyclic layout which applies to any stabilizer codes, whereas previous results for qubit arrays are limited to CSS codes. (ii) The sparse cyclic layout, specific to bivariate bicycle (BB) codes. For the $[[144,12,12]]$ BB code, using the sparse cyclic layout we obtain a quantum memory with $12$ logical qubits distributed over $12$ modules, containing $12$ physical qubits each. We propose physical implementations of this architecture using flying qubits, that can be faithfully transported, and include qubits encoded in ions, neutral atoms, electrons or photons. We performed numerical simulations when modules are long ion chains and when modules are single-qubit arrays of ions showing that the distributed BB code achieves a logical error rate below $2 \cdot 10^{-6}$ when the physical error rate is $10^{-3}$.

quant-ph

Quantum error correction for long chains of trapped ions

We propose a model for quantum computing with long chains of trapped ions and we design quantum error correction schemes for this model. The main components of a quantum error correction scheme are the quantum code and a quantum circuit called the syndrome extraction circuit, which is executed to perform error correction with this code. In this work, we design syndrome extraction circuits tailored to our ion chain model, a syndrome extraction tuning protocol to optimize these circuits, and we construct new quantum codes that outperform the state-of-the-art for chains of about $50$ qubits. To establish a baseline under the ion chain model, we simulate the performance of surface codes and bivariate bicycle (BB) codes equipped with our optimized syndrome extraction circuits. Then, we propose a new variant of BB codes defined by weight-five measurements, that we refer to as BB5 codes and we identify BB5 codes that achieve a better minimum distance than any BB codes with the same number of logical qubits and data qubits, such as a $[[48, 4, 7]]$ BB5 code. For a physical error rate of $10^{-3}$, the $[[48, 4, 7]]$ BB5 code achieves a logical error rate per logical qubit of $5 \cdot 10^{-5}$, which is four times smaller than the best BB code in our baseline family. It also achieves the same logical error rate per logical qubit as the distance-7 surface code but using four times fewer physical qubits per logical qubit.

quant-ph

Label-free Prediction of Vascular Connectivity in Perfused Microvascular Networks in vitro

Continuous monitoring and in-situ assessment of microvascular connectivity have significant implications for culturing vascularized organoids and optimizing the therapeutic strategies. However, commonly used methods for vascular connectivity assessment heavily rely on fluorescent labels that may either raise biocompatibility concerns or interrupt the normal cell growth process. To address this issue, a Vessel Connectivity Network (VC-Net) was developed for label-free assessment of vascular connectivity. To validate the VC-Net, microvascular networks (MVNs) were cultured in vitro and their microscopic images were acquired at different culturing conditions as a training dataset. The VC-Net employs a Vessel Queue Contrastive Learning (VQCL) method and a class imbalance algorithm to address the issues of limited sample size, indistinctive class features and imbalanced class distribution in the dataset. The VC-Net successfully evaluated the vascular connectivity with no significant deviation from that by fluorescence imaging. In addition, the proposed VC-Net successfully differentiated the connectivity characteristics between normal and tumor-related MVNs. In comparison with those cultured in the regular microenvironment, the averaged connectivity of MVNs cultured in the tumor-related microenvironment decreased by 30.8%, whereas the non-connected area increased by 37.3%. This study provides a new avenue for label-free and continuous assessment of organoid or tumor vascularization in vitro.

eess.IV

AquilaMoE: Efficient Training for MoE Models with Scale-Up and Scale-Out Strategies

In recent years, with the rapid application of large language models across various fields, the scale of these models has gradually increased, and the resources required for their pre-training have grown exponentially. Training an LLM from scratch will cost a lot of computation resources while scaling up from a smaller model is a more efficient approach and has thus attracted significant attention. In this paper, we present AquilaMoE, a cutting-edge bilingual 8*16B Mixture of Experts (MoE) language model that has 8 experts with 16 billion parameters each and is developed using an innovative training methodology called EfficientScale. This approach optimizes performance while minimizing data requirements through a two-stage process. The first stage, termed Scale-Up, initializes the larger model with weights from a pre-trained smaller model, enabling substantial knowledge transfer and continuous pretraining with significantly less data. The second stage, Scale-Out, uses a pre-trained dense model to initialize the MoE experts, further enhancing knowledge transfer and performance. Extensive validation experiments on 1.8B and 7B models compared various initialization schemes, achieving models that maintain and reduce loss during continuous pretraining. Utilizing the optimal scheme, we successfully trained a 16B model and subsequently the 8*16B AquilaMoE model, demonstrating significant improvements in performance and training efficiency.

cs.CL

MSR Codes with Linear Field Size and Smallest Sub-packetization for Any Number of Helper Nodes

The sub-packetization $\ell$ and the field size $q$ are of paramount importance in the MSR array code constructions. For optimal-access MSR codes, Balaji et al. proved that $\ell\geq s^{\left\lceil n/s \right\rceil}$, where $s = d-k+1$. Rawat et al. showed that this lower bound is attainable for all admissible values of $d$ when the field size is exponential in $n$. After that, tremendous efforts have been devoted to reducing the field size. However, till now, reduction to linear field size is only available for $d\in\{k+1,k+2,k+3\}$ and $d=n-1$. In this paper, we construct the first class of explicit optimal-access MSR codes with the smallest sub-packetization $\ell = s^{\left\lceil n/s \right\rceil}$ for all $d$ between $k+1$ and $n-1$, resolving an open problem in the survey (Ramkumar et al., Foundations and Trends in Communications and Information Theory: Vol. 19: No. 4). We further propose another class of explicit MSR code constructions (not optimal-access) with even smaller sub-packetization $s^{\left\lceil n/(s+1)\right\rceil }$ for all admissible values of $d$, making significant progress on another open problem in the survey. Previously, MSR codes with $\ell=s^{\left\lceil n/(s+1)\right\rceil }$ and $q=O(n)$ were only known for $d=k+1$ and $d=n-1$. The key insight that enables a linear field size in our construction is to reduce $\binom{n}{r}$ global constraints of non-vanishing determinants to $O_s(n)$ local ones, which is achieved by carefully designing the parity check matrices.

cs.IT

ABS+ Polar Codes: Exploiting More Linear Transforms on Adjacent Bits

ABS polar codes were recently proposed to speed up polarization by swapping certain pairs of adjacent bits after each layer of polar transform. In this paper, we observe that applying the Arikan transform $(U_i, U_{i+1}) \mapsto (U_{i}+U_{i+1}, U_{i+1})$ on certain pairs of adjacent bits after each polar transform layer leads to even faster polarization. In light of this, we propose ABS+ polar codes which incorporate the Arikan transform in addition to the swapping transform in ABS polar codes. In order to efficiently construct and decode ABS+ polar codes, we derive a new recursive relation between the joint distributions of adjacent bits through different layers of polar transforms. Simulation results over a wide range of parameters show that the CRC-aided SCL decoder of ABS+ polar codes improves upon that of ABS polar codes by 0.1dB--0.25dB while maintaining the same decoding time. Moreover, ABS+ polar codes improve upon standard polar codes by 0.2dB--0.45dB when they both use the CRC-aided SCL decoder with list size $32$. The implementations of all the algorithms in this paper are available at https://github.com/PlumJelly/ABS-Polar

cs.IT

Constructing MSR codes with subpacketization $2^{n/3}$ for $k+1$ helper nodes

Wang et al. (IEEE Transactions on Information Theory, vol. 62, no. 8, 2016) proposed an explicit construction of an $(n=k+2,k)$ Minimum Storage Regenerating (MSR) code with $2$ parity nodes and subpacketization $2^{k/3}$. The number of helper nodes for this code is $d=k+1=n-1$, and this code has the smallest subpacketization among all the existing explicit constructions of MSR codes with the same $n,k$ and $d$. In this paper, we present a new construction of MSR codes for a wider range of parameters. More precisely, we still fix $d=k+1$, but we allow the code length $n$ to be any integer satisfying $n\ge k+2$. The field size of our code is linear in $n$, and the subpacketization of our code is $2^{n/3}$. This value is slightly larger than the subpacketization of the construction by Wang et al. because their code construction only guarantees optimal repair for all the systematic nodes while our code construction guarantees optimal repair for all nodes.

cs.IT

All the codeword symbols in polar codes have the same SER under the SC decoder

We consider polar codes constructed from the $2\times 2$ kernel $\begin{bmatrix} 1 & 0 \\ \alpha & 1 \end{bmatrix}$ over a finite field $\mathbb{F}_{q}$, where $q=p^s$ is a power of a prime number $p$, and $\alpha$ satisfies that $\mathbb{F}_{p}(\alpha) = \mathbb{F}_{q}$. We prove that for any $\mathbb{F}_{q}$-symmetric memoryless channel, any code length, and any code dimension, all the codeword symbols in such polar codes have the same symbol error rate (SER) under the successive cancellation (SC) decoder.

cs.IT

Adjacent-Bits-Swapped Polar codes: A new code construction to speed up polarization

The construction of polar codes with code length $n=2^m$ involves $m$ layers of polar transforms. In this paper, we observe that after each layer of polar transforms, one can swap certain pairs of adjacent bits to accelerate the polarization process. More precisely, if the previous bit is more reliable than its next bit under the successive decoder, then switching the decoding order of these two adjacent bits will make the reliable bit even more reliable and the noisy bit even noisier. Based on this observation, we propose a new family of codes called the Adjacent-Bits-Swapped (ABS) polar codes. We add a permutation layer after each polar transform layer in the construction of the ABS polar codes. In order to choose which pairs of adjacent bits to swap in the permutation layers, we rely on a new polar transform that combines two independent channels with $4$-ary inputs. This new polar transform allows us to track the evolution of every pair of adjacent bits through different layers of polar transforms, and it also plays an essential role in the Successive Cancellation List (SCL) decoder for the ABS polar codes. Extensive simulation results show that ABS polar codes consistently outperform standard polar codes by 0.15dB--0.3dB when we use CRC-aided SCL decoder with list size $32$ for both codes. The implementations of all the algorithms in this paper are available at https://github.com/PlumJelly/ABS-Polar

cs.IT

A Dynamic Programming Method to Construct Polar Codes with Improved Performance

In the standard polar code construction, the message vector $(U_0,U_1,\dots,U_{n-1})$ is divided into information bits and frozen bits according to the reliability of each $U_i$ given $(U_0,U_1,\dots,U_{i-1})$ and all the channel outputs. While this reliability function is the most suitable measure to choose information bits under the Successive Cancellation (SC) decoder, there is a mismatch between this reliability function and the Successive Cancellation List (SCL) decoder because the SCL decoder also makes use of the information from the future frozen bits. We propose a Dynamic Programming (DP) construction of polar codes to resolve this mismatch. Our DP construction chooses different sets of information bits for different list sizes in order to optimize the performance of the constructed code under the SCL decoder. Simulation results show that our DP-polar codes consistently demonstrate $0.3$--$1$dB improvement over the standard polar codes under the SCL decoder with list size $32$ for various choices of code lengths and code rates.

cs.IT

Improving the List Decoding Version of the Cyclically Equivariant Neural Decoder

The cyclically equivariant neural decoder was recently proposed in [Chen-Ye, International Conference on Machine Learning, 2021] to decode cyclic codes. In the same paper, a list decoding procedure was also introduced for two widely used classes of cyclic codes -- BCH codes and punctured Reed-Muller (RM) codes. While the list decoding procedure significantly improves the Frame Error Rate (FER) of the cyclically equivariant neural decoder, the Bit Error Rate (BER) of the list decoding procedure is even worse than the unique decoding algorithm when the list size is small. In this paper, we propose an improved version of the list decoding algorithm for BCH codes and punctured RM codes. Our new proposal significantly reduces the BER while maintaining the same (in some cases even smaller) FER. More specifically, our new decoder provides up to $2$dB gain over the previous list decoder when measured by BER, and the running time of our new decoder is $15\%$ smaller. Code available at https://github.com/improvedlistdecoder/code

cs.IT