SearcharxivSearch

arXiv subjects

Lin Gan

Publications and source records attributed to Lin Gan.

At least 37 records · Page 2Linked to original sources

Gaussian Boson Sampling with Pseudo-Photon-Number Resolving Detectors and Quantum Computational Advantage

We report new Gaussian boson sampling experiments with pseudo-photon-number-resolving detection, which register up to 255 photon-click events. We consider partial photon distinguishability and develop a more complete model for the characterization of the noisy Gaussian boson sampling. In the quantum computational advantage regime, we use Bayesian tests and correlation function analysis to validate the samples against all current classical mockups. Estimating with the best classical algorithms to date, generating a single ideal sample from the same distribution on the supercomputer Frontier would take ~ 600 years using exact methods, whereas our quantum computer, Jiuzhang 3.0, takes only 1.27 us to produce a sample. Generating the hardest sample from the experiment using an exact algorithm would take Frontier ~ 3.1*10^10 years.

quant-ph

Lifetime-based Optimization for Simulating Quantum Circuits on a New Sunway Supercomputer

High-performance classical simulator for quantum circuits, in particular the tensor network contraction algorithm, has become an important tool for the validation of noisy quantum computing. In order to address the memory limitations, the slicing technique is used to reduce the tensor dimensions, but it could also lead to additional computation overhead that greatly slows down the overall performance. This paper proposes novel lifetime-based methods to reduce the slicing overhead and improve the computing efficiency, including an interpretation method to deal with slicing overhead, an in-place slicing strategy to find the smallest slicing set and an adaptive tensor network contraction path refiner customized for Sunway architecture. Experiments show that in most cases the slicing overhead with our in-place slicing strategy would be less than the cotengra, which is the most used graph path optimization software at present. Finally, the resulting simulation time is reduced to 96.1s for the Sycamore quantum processor RQC, with a sustainable single-precision performance of 308.6Pflops using over 41M cores to generate 1M correlated samples, which is more than 5 times performance improvement compared to 60.4 Pflops in 2021 Gordon Bell Prize work.

cs.DC

swTVM: Towards Optimized Tensor Code Generation for Deep Learning on Sunway Many-Core Processor

The flourish of deep learning frameworks and hardware platforms has been demanding an efficient compiler that can shield the diversity in both software and hardware in order to provide application portability. Among the existing deep learning compilers, TVM is well known for its efficiency in code generation and optimization across diverse hardware devices. In the meanwhile, the Sunway many-core processor renders itself as a competitive candidate for its attractive computational power in both scientific computing and deep learning workloads. This paper combines the trends in these two directions. Specifically, we propose swTVM that extends the original TVM to support ahead-of-time compilation for architecture requiring cross-compilation such as Sunway. In addition, we leverage the architecture features during the compilation such as core group for massive parallelism, DMA for high bandwidth memory transfer and local device memory for data locality, in order to generate efficient codes for deep learning workloads on Sunway. The experiment results show that the codes generated by swTVM achieves 1.79x on average compared to the state-of-the-art deep learning framework on Sunway, across six representative benchmarks. This work is the first attempt from the compiler perspective to bridge the gap of deep learning and Sunway processor particularly with productivity and efficiency in mind. We believe this work will encourage more people to embrace the power of deep learning and Sunway many-core processor.

cs.LG

Phase-Programmable Gaussian Boson Sampling Using Stimulated Squeezed Light

The tantalizing promise of quantum computational speedup in solving certain problems has been strongly supported by recent experimental evidence from a high-fidelity 53-qubit superconducting processor1 and Gaussian boson sampling (GBS) with up to 76 detected photons. Analogous to the increasingly sophisticated Bell tests that continued to refute local hidden variable theories, quantum computational advantage tests are expected to provide increasingly compelling experimental evidence against the Extended Church-Turing thesis. In this direction, continued competition between upgraded quantum hardware and improved classical simulations is required. Here, we report a new GBS experiment that produces up to 113 detection events out of a 144-mode photonic circuit. We develop a new high-brightness and scalable quantum light source, exploring the idea of stimulated squeezed photons, which has simultaneously near-unity purity and efficiency. This GBS is programmable by tuning the phase of the input squeezed states. We demonstrate a new method to efficiently validate the samples by inferring from computationally friendly subsystems, which rules out hypotheses including distinguishable photons and thermal states. We show that our noisy GBS experiment passes the nonclassicality test using an inequality, and we reveal non-trivial genuine high-order correlation in the GBS samples, which are evidence of robustness against possible classical simulation schemes. The photonic quantum computer, Jiuzhang 2.0, yields a Hilbert space dimension up to $10^{43}$, and a sampling rate $10^{24}$ faster than using brute-force simulation on supercomputers.

quant-ph

Quantum computational advantage using photons

Gaussian boson sampling exploits squeezed states to provide a highly efficient way to demonstrate quantum computational advantage. We perform experiments with 50 input single-mode squeezed states with high indistinguishability and squeezing parameters, which are fed into a 100-mode ultralow-loss interferometer with full connectivity and random transformation, and sampled using 100 high-efficiency single-photon detectors. The whole optical set-up is phase-locked to maintain a high coherence between the superposition of all photon number states. We observe up to 76 output photon-clicks, which yield an output state space dimension of $10^{30}$ and a sampling rate that is $10^{14}$ faster than using the state-of-the-art simulation strategy and supercomputers. The obtained samples are validated against various hypotheses including using thermal states, distinguishable photons, and uniform distribution.

quant-ph

Benchmarking 50-Photon Gaussian Boson Sampling on the Sunway TaihuLight

Boson sampling is expected to be one of an important milestones that will demonstrate quantum supremacy. The present work establishes the benchmarking of Gaussian boson sampling (GBS) with threshold detection based on the Sunway TaihuLight supercomputer. To achieve the best performance and provide a competitive scenario for future quantum computing studies, the selected simulation algorithm is fully optimized based on a set of innovative approaches, including a parallel scheme and instruction-level optimizing method. Furthermore, data precision and instruction scheduling are handled in a sophisticated manner by an adaptive precision optimization scheme and a DAG-based heuristic search algorithm, respectively. Based on these methods, a highly efficient and parallel quantum sampling algorithm is designed. The largest run enables us to obtain one Torontonian function of a 100 x 100 submatrix from 50-photon GBS within 20 hours in 128-bit precision and 2 days in 256-bit precision.

cs.DC

The Deep Learning Compiler: A Comprehensive Survey

The difficulty of deploying various deep learning (DL) models on diverse DL hardware has boosted the research and development of DL compilers in the community. Several DL compilers have been proposed from both industry and academia such as Tensorflow XLA and TVM. Similarly, the DL compilers take the DL models described in different DL frameworks as input, and then generate optimized codes for diverse DL hardware as output. However, none of the existing survey has analyzed the unique design architecture of the DL compilers comprehensively. In this paper, we perform a comprehensive survey of existing DL compilers by dissecting the commonly adopted design in details, with emphasis on the DL oriented multi-level IRs, and frontend/backend optimizations. Specifically, we provide a comprehensive comparison among existing DL compilers from various aspects. In addition, we present detailed analysis on the design of multi-level IRs and illustrate the commonly adopted optimization techniques. Finally, several insights are highlighted as the potential research directions of DL compiler. This is the first survey paper focusing on the design architecture of DL compilers, which we hope can pave the road for future research towards DL compiler.

cs.DC

Quantum Teleportation-Inspired Algorithm for Sampling Large Random Quantum Circuits

We show that low-depth random quantum circuits can be efficiently simulated by a quantum teleportation-inspired algorithm. By using logical qubits to redirect and teleport the quantum information in quantum circuits, the original circuits can be renormalized to new circuits with a smaller number of logical qubits. We demonstrate the algorithm to simulate several random quantum circuits, including 1D-chain 1000-qubit 42-depth, 2D-grid 125*8-qubit 42-depth and 2D-Bristlecone 72-qubit 32-depth circuits. Our results present a memory-efficient method with a clear physical picture to simulate low-depth random quantum circuits.

quant-ph

Efficient Channel Model for Homogeneous Weakly Coupled Multicore Fiber

To analyze a homogeneous weakly coupled multicore fiber (WC-MCF) based transmission system via simulation, we propose an efficient (fast and accurate) WC-MCF's channel model, which can describe the propagation effects including attenuation, walk-off, chromatic dispersion, self-phase modulation (SPM), and especially the frequency-dependent inter-core crosstalk (XT). We speed up the simulation with two orders of magnitude by simplifying the XT's calculation. On one hand, the calculation step size can be greatly increased by utilizing a new XT's coupling matrix. On the other hand, the calculation of XT can be further accelerated by down-sampling XT's coupling matrix in frequency domain. The XT power and average occurrence distance should be set manually based on the existing XT model to describe the frequency-dependent XT the same as a real WC-MCF. We numerically and experimentally observed that XT's de-correlation bandwidth decreases with relative time delay (RTD) by fractional linear function. The range of validity of the proposed channel model is also discussed with different walk-off and coupling strength. We believe the proposed efficient channel model can provide great help for analysis and optimization of homogeneous WC-MCF based optical communication systems.

physics.optics

High-speed PAM4-based Optical SDM Interconnects with Directly Modulated Long-wavelength VCSEL

This paper reports the demonstration of high-speed PAM-4 transmission using a 1.5-μm single-mode vertical cavity surface emitting laser (SM-VCSEL) over multicore fiber with 7 cores over different distances. We have successfully generated up to 70 Gbaud 4-level pulse amplitude modulation (PAM-4) signals with a VCSEL in optical back-to-back, and transmitted 50 Gbaud PAM-4 signals over both 1-km dispersion-uncompensated and 10-km dispersion-compensated in each core, enabling a total data throughput of 700 Gbps over the 7-core fiber. Moreover, 56 Gbaud PAM-4 over 1-km has also been shown, whereby unfortunately not all cores provide the required 3.8 $\times$ 10 $^{-3}$ bit error rate (BER) for the 7% overhead-hard decision forward error correction (7% OH HDFEC). The limited bandwidth of the VCSEL and the adverse chromatic dispersion of the fiber are suppressed with pre-equalization based on accurate end-to-end channel characterizations. With a digital post-equalization, BER performance below the 7% OH-HDFEC limit is achieved over all cores. The demonstrated results show a great potential to realize high-capacity and compact short-reach optical interconnects for data centers.

eess.SP

Layered Optical Flow Estimation Using a Deep Neural Network with a Soft Mask

Using a layered representation for motion estimation has the advantage of being able to cope with discontinuities and occlusions. In this paper, we learn to estimate optical flow by combining a layered motion representation with deep learning. Instead of pre-segmenting the image to layers, the proposed approach automatically generates a layered representation of optical flow using the proposed soft-mask module. The essential components of the soft-mask module are maxout and fuse operations, which enable a disjoint layered representation of optical flow and more accurate flow estimation. We show that by using masks the motion estimate results in a quadratic function of input features in the output layer. The proposed soft-mask module can be added to any existing optical flow estimation networks by replacing their flow output layer. In this work, we use FlowNet as the base network to which we add the soft-mask module. The resulting network is tested on three well-known benchmarks with both supervised and unsupervised flow estimation tasks. Evaluation results show that the proposed network achieve better results compared with the original FlowNet.

cs.CV

Investigation of channel model for weakly coupled multicore fiber

We investigate the evolution of decorrelation bandwidth of inter core crosstalk (IC-XT) in homogeneous weakly coupled multicore fibers (WC-MCFs). The modified mode-coupled equations (MCEs) are numerically solved by combining the fourth order Runge-Kutta method and compound Simpson integral method. It can be theoretically and numerically observed that the decorrelation bandwidth of IC-XT decreases with transmission distance by fractional linear function. The evolution rule of IC-XT's decorrelation bandwidth is further confirmed by experiments, which can be used as an evaluation criterion for channel model. Finally, we propose a new channel model with the coupling matrix of IC-XT generated automatically by phase transfer function (PTF), which is in good agreement with the above evaluation criterion. We believe the proposed channel model can provide a good simulation platform for homogeneous WC-MCF based communication systems.

physics.optics