Searcharxiv⌕ Search

arXiv subjects

Cong Guo

Publications and source records attributed to Cong Guo.

70 records · Page 4Linked to original sources

SQuant: On-the-Fly Data-Free Quantization via Diagonal Hessian Approximation

Quantization of deep neural networks (DNN) has been proven effective for compressing and accelerating DNN models. Data-free quantization (DFQ) is a promising approach without the original datasets under privacy-sensitive and confidential scenarios. However, current DFQ solutions degrade accuracy, need synthetic data to calibrate networks, and are time-consuming and costly. This paper proposes an on-the-fly DFQ framework with sub-second quantization time, called SQuant, which can quantize networks on inference-only devices with low computation and memory requirements. With the theoretical analysis of the second-order information of DNN task loss, we decompose and approximate the Hessian-based optimization objective into three diagonal sub-items, which have different areas corresponding to three dimensions of weight tensor: element-wise, kernel-wise, and output channel-wise. Then, we progressively compose sub-items and propose a novel data-free optimization objective in the discrete domain, minimizing Constrained Absolute Sum of Error (or CASE in short), which surprisingly does not need any dataset and is even not aware of network architecture. We also design an efficient algorithm without back-propagation to further reduce the computation complexity of the objective solver. Finally, without fine-tuning and synthetic datasets, SQuant accelerates the data-free quantization process to a sub-second level with >30% accuracy improvement over the existing data-free post-training quantization works, with the evaluated models under 4-bit quantization. We have open-sourced the SQuant framework at https://github.com/clevercool/SQuant.

cs.LG↗

Characterizing and Demystifying the Implicit Convolution Algorithm on Commercial Matrix-Multiplication Accelerators

Many of today's deep neural network accelerators, e.g., Google's TPU and NVIDIA's tensor core, are built around accelerating the general matrix multiplication (i.e., GEMM). However, supporting convolution on GEMM-based accelerators is not trivial. The naive method explicitly lowers the convolution to GEMM, commonly known as im2col, which introduces significant performance and memory overhead. Existing implicit im2col algorithms require unscalable hardware and are inefficient in supporting important convolution variants such as strided convolution. In this paper, we propose a memory-efficient and hardware-friendly implicit im2col algorithm used by Google's TPU, which dynamically converts a convolution into a GEMM with practically zero performance and memory overhead, fully unleashing the power of GEMM engines. Through comprehensive experimental results, we quantitatively argue that this algorithm has been adopted in commercial closed-source platforms, and we are the first to describe its high-level idea and implementation details. Finally, we show that our algorithm can also be generally applied to Nvidia's Tensor Cores (TC), matching and out-performing the measured performance on TCs.

cs.DC↗

Dual-side Sparse Tensor Core

Leveraging sparsity in deep neural network (DNN) models is promising for accelerating model inference. Yet existing GPUs can only leverage the sparsity from weights but not activations, which are dynamic, unpredictable, and hence challenging to exploit. In this work, we propose a novel architecture to efficiently harness the dual-side sparsity (i.e., weight and activation sparsity). We take a systematic approach to understand the (dis)advantages of previous sparsity-related architectures and propose a novel, unexplored paradigm that combines outer-product computation primitive and bitmap-based encoding format. We demonstrate the feasibility of our design with minimal changes to the existing production-scale inner-product-based Tensor Core. We propose a set of novel ISA extensions and co-design the matrix-matrix multiplication and convolution algorithms, which are the two dominant computation patterns in today's DNN models, to exploit our new dual-side sparse Tensor Core. Our evaluation shows that our design can fully unleash the dual-side DNN sparsity and improve the performance by up to one order of magnitude with \hl{small} hardware overhead.

cs.AR↗

Prospects of detecting the reactor $\bar{ν_e}$-Ar coherent elastic scattering with a low threshold dual-phase argon time projection chamber at Taishan

We propose to measure the coherent elastic neutrino nucleus scattering (CE$ν$NS) using a dual-phase liquid argon time projection chamber (TPC) with 200kg fiducial mass. The detector is expected to be adjacent to the JUNO-TAO experiment and to be about 35m from a reactor core with 4.6GW thermal power at Taishan. The antineutrino flux is approximately 6$\times10^{12}$cm$^{-1}$s$^{-1}$ at this location, leading to more than 11,000 coherent scattering events per day in the fiducial mass. However, the nuclear recoil energies concentrate in the sub-keV region, corresponding to less than ten ionisation electrons in the liquid argon. The detection of several ionisation electrons can be achieved in the dual-phase TPC due to the large amplification in the gas region. With a feasible detection threshold of four ionisation electrons, the signal rate is 955 per day. The detector is designed to be shielded well from cosmogenic backgrounds and ambient radioactivities to reach a 16% background-to-signal ratio in the energy region of interest. With the large CE$ν$NS sample, the expected sensitivity of measuring the weak mixing angle $\sin^{2}θ_{W}$, and of limiting the neutrino magnetic moment are discussed. In addition, a synergy between the reactor antineutrino CE$ν$NS experiment and the dark matter experiment is foreseen.

hep-ex↗

Accelerating Sparse DNN Models without Hardware-Support via Tile-Wise Sparsity

Network pruning can reduce the high computation cost of deep neural network (DNN) models. However, to maintain their accuracies, sparse models often carry randomly-distributed weights, leading to irregular computations. Consequently, sparse models cannot achieve meaningful speedup on commodity hardware (e.g., GPU) built for dense matrix computations. As such, prior works usually modify or design completely new sparsity-optimized architectures for exploiting sparsity. We propose an algorithm-software co-designed pruning method that achieves latency speedups on existing dense architectures. Our work builds upon the insight that the matrix multiplication generally breaks the large matrix into multiple smaller tiles for parallel execution. We propose a tiling-friendly "tile-wise" sparsity pattern, which maintains a regular pattern at the tile level for efficient execution but allows for irregular, arbitrary pruning at the global scale to maintain the high accuracy. We implement and evaluate the sparsity pattern on GPU tensor core, achieving a 1.95x speedup over the dense model.

cs.DC↗

Balancing Efficiency and Flexibility for DNN Acceleration via Temporal GPU-Systolic Array Integration

The research interest in specialized hardware accelerators for deep neural networks (DNN) spikes recently owing to their superior performance and efficiency. However, today's DNN accelerators primarily focus on accelerating specific "kernels" such as convolution and matrix multiplication, which are vital but only part of an end-to-end DNN-enabled application. Meaningful speedups over the entire application often require supporting computations that are, while massively parallel, ill-suited to DNN accelerators. Integrating a general-purpose processor such as a CPU or a GPU incurs significant data movement overhead and leads to resource under-utilization on the DNN accelerators. We propose Simultaneous Multi-mode Architecture (SMA), a novel architecture design and execution model that offers general-purpose programmability on DNN accelerators in order to accelerate end-to-end applications. The key to SMA is the temporal integration of the systolic execution model with the GPU-like SIMD execution model. The SMA exploits the common components shared between the systolic-array accelerator and the GPU, and provides lightweight reconfiguration capability to switch between the two modes in-situ. The SMA achieves up to 63% performance improvement while consuming 23% less energy than the baseline Volta architecture with TensorCore.

cs.DC↗

The liquid argon detector and measurement of SiPM array at liquid argon temperature

Particle detectors based on liquid argon (LAr) have recently become recognized as an extremely attractive technology for the direct detection of dark matter as well as the measurement of coherent elastic neutrino-nucleus scattering (CE$ν$NS). The Chinese argon group at Institute of High Energy Physics has been studying the LAr detector technology and a LAr detector has been operating steadily. A program of using a dual phase LAr detector to measure the CE$ν$NS at Taishang Nuclear Power Plant has been proposed and the R\&D work is ongoing. Considering the requirements of ultra-low radio-purity and high photon collection efficiency, SiPMs will be a good choice and will be used in the detector. In this proceeding, an introduction of the LAr detector and the measurement results of SiPM array at LAr temperature will be presented.

physics.ins-det↗

Developing the radium measurement system for the water Cherenkov detector of the Jiangmen Underground Neutrino Observatory

The Jiangmen Underground Neutrino Observatory is proposed to determine neutrino mass hierarchy using a 20~ktonne liquid scintillator detector. Strict radio-purity requirements have been put forward for all the components of the detector. According to the MC simulation results, the radon dissolved in the water Cherenkov detector should be below 200~mBq/m$^3$. Radium, the progenitor of radon, should also be taken seriously into account. In order to measure the radium concentration in water, a radium measurement system, which consists of a radium extraction system, a radon emanation chamber and a radon concentration measurement system, has been developed. In this paper, the updated radon concentration in gas measurement system with a one-day-measurement sensitivity of $\sim$5~mBq/m$^3$, the detail of the development of the radium concentration in water measurement system with a sensitivity of $\sim$23~mBq/m$^3$ as well as the measurement results of Daya Bay water samples will be presented.

astro-ph.IM↗

Status of the Jiangmen Underground Neutrino Observatory

The Jiangmen Underground Neutrino Observatory is a multipurpose neutrino experiment designed to determine neutrino mass hierarchy and precisely measure oscillation parameters by detecting reactor neutrinos from the Yangjiang and Taishan Nuclear Power Plants, observe supernova neutrinos, study the atmospheric, solar neutrinos and geo-neutrinos, and perform exotic searches, with a 20-thousand-ton liquid scintillator detector of unprecedented 3\% energy resolution (at 1 MeV) at 700-meter deep underground. In this proceeding, the subsystems of the experiment, including the cental detector, the online scintillator internal radioactivity investigation system, the PMT, the veto detector, the calibration system and the taishan antineutrino observatory, will be described. The construction is expected to be completed in 2021.

physics.ins-det↗

Calibration of liquid argon detector with $^{83m}Kr$ and $^{22}Na$ in different drift field

$^{83m}Kr$ and $^{22}Na$ have been used in calibrating a liquid argon (LAr) detector.$^{83m}Kr$ atoms are produced through the decay of $^{83}Rb$ and introduced into the LAr detector through the circulating purification system. The light yield reaches 7.26$\pm$0.02 photonelectrons/keV for 41.5keV from $^{83m}Kr$ and 7.66$\pm$0.01 photonelectrons/keV for the 511keV from $^{22}Na$, as a comparison. The light yield varies with the drift electric field from 50 to 200V/cm have been also reported. After stopping fill, the decay rate of $^{83m}Kr$ with a fitted half-life of 1.83$\pm$0.11 h, which is consistent with the reported value of 1.83$\pm$0.02 h.

physics.ins-det↗

Adversarial Defense Through Network Profiling Based Path Extraction

Recently, researchers have started decomposing deep neural network models according to their semantics or functions. Recent work has shown the effectiveness of decomposed functional blocks for defending adversarial attacks, which add small input perturbation to the input image to fool the DNN models. This work proposes a profiling-based method to decompose the DNN models to different functional blocks, which lead to the effective path as a new approach to exploring DNNs' internal organization. Specifically, the per-image effective path can be aggregated to the class-level effective path, through which we observe that adversarial images activate effective path different from normal images. We propose an effective path similarity-based method to detect adversarial images with an interpretable model, which achieve better accuracy and broader applicability than the state-of-the-art technique.

cs.LG↗

An Anomalous Circular Photogalvanic Effect in the Weyl Semimetal TaAs

Weyl semimetal (WSM) is expected to be an ideal spintronic material owing to its spin currents carried by the bulk and surface states with spin-momentum locking. A photocurrent generation in noncentrosymmetric WSM was also predicted owing to its broken inversion symmetry and linear energy dispersion which are unique to Weyl systems. In our recent measurements, the circular photogalvanic effect (CPGE) has been demonstrated in WSM of TaAs. The CPGE voltage is proportional to the helicity of the incident light and reverses its direction on changing the radiation helicity from left handed to right handed, a periodical oscillation therefore appears following with the alteration of optical polarization. We herein attribute the CPGE to the asymmetric optical excitation of the Weyl cone, which could result in an asymmetric distribution of photoexcited carriers in momentum space according to an optical spin selection rule.

cond-mat.str-el↗

Neutron Beam Tests of Barium Fluoride Crystal for Dark Matter Direct Detection

In order to test the capabilities of Barium Fluoride (BaF2) Crystal for dark matter direct detection, nuclear recoils are studied with mono-energetic neutron beam. The energy spectra of nuclear recoils, quenching factors for elastic scattering neutrons and discrimination capability between neutron inelastic scattering events and γ events are obtained for various recoil energies of the F content in BaF2.

physics.ins-det↗

Seasonal and Lunar month periods observed in natural neutron flux at high altitude

Air radon concentration measurement is useful for research on geophysical effects, but it is strongly sensitive to site geology and many geophysical and microclimatic processes such as wind, ventilation, air humidity and so on that induce very big fluctuations on the concentration of radon in air. On the contrary, monitoring the radon concentration in soil by measuring the thermal neutron flux reduces environmental effects. In this paper we report some experimental results on the natural thermal neutron flux as well as the concentration of air radon and its variations at 4300 m a.s.l. These results were obtained with unshielded thermal neutron scintillation detectors (en-detectors) and radon monitors located inside the ARGO-YBJ experimental hall. The correlation of these variations with the lunar month and 1-year period is undoubtedly confirmed. A method for earthquakes prediction provided by a global net of the en-detectors is currently under study.

physics.geo-ph↗

Neutron Beam Tests of $CsI(Na)$ and $CaF_{2}(Eu)$ Crystals for Dark Matter Direct Search

In recent decades, inorganic crystals have been widely used in dark matter direct search experiments. To contribute to the understanding of the capabilities of $CsI(Na)$ and $CaF_{2}(Eu)$ crystals, a mono-energetic neutron beam is utilized to study the properties of nuclear recoils, which are expected to be similar to signals of dark matter direct detection. The quenching factor of nuclear recoils in $CsI(Na)$ and $CaF_{2}(Eu)$, as well as an improved discrimination factor between nuclear recoils and $γ$ backgrounds in $CsI(Na)$, are reported.

physics.ins-det↗

Preliminary test results of LAr prototype detector

WIMPs are a well-motivated galactic dark matter candidate. Liquid argon (LAr) is an attractive target for the direct detection of WIMPs. The LAr prototype detector is designed to study the technology and property of LAr detector. The prototype detector have an active volume containing 0.65 kg of liquid argon. The liquid nitrogen(LN) cooling system allows the temperature of liquid argon to be maintained at the boiling point (87.8 K) with fluctuations less than 0.1 K. The prototype was calibrated with a Na$^{22}$ source, with the light yield 1.591$\pm$0.019 p.e./keV for the 511 keV gamma rays using the domestic-made argon purification system.

physics.ins-det↗