Searcharxiv⌕ Search

arXiv subjects

Ang Li

Publications and source records attributed to Ang Li.

At least 523 records · Page 29Linked to original sources

Comprehensive analysis of the tidal effect in gravitational waves and implication for cosmology

Detection of gravitational waves (GWs) produced by coalescence of compact binaries provides a novel way to measure the luminosity distance of GW events. Combining their redshift, they can act as standard sirens to constrain cosmological parameters. For various GW detector networks in 2nd-generation (2G), 2.5G and 3G, we comprehensively analyze the method to constrain the equation-of-state (EOS) of binary neutron-stars (BNSs) and extract their redshifts through the imprints of tidal effects in GW waveforms. We find for these events, the observations of electromagnetic counterparts in low-redshift range $z < 0.1$ are important for constraining the tidal effects. Considering 17 different EOSs of NSs or quark-stars, we find GW observations have strong capability to determine the EOS. Applying the events as standard sirens, and considering the constraints of NS's EOS derived from low-redshift observations as prior, we can constrain the dark-energy EOS parameters $w_0$ and $w_a$. In 3G era, the potential constraints are $Δw_0\in (0.0006,0.004)$ and $Δw_a\in(0.004,0.02)$, which are 1-3 orders smaller than those from traditional methods, including Type Ia supernovas and baryon acoustic oscillations. The constraints are also 1 order smaller than the method of GW standard siren by fixing the redshifts through short-hard $γ$-ray bursts, due to more available GW events in this method. Therefore, GW standard sirens, based on the tidal effect measurement, provide a realizable and much more powerful tool in cosmology.

gr-qc↗

FPDeep: Scalable Acceleration of CNN Training on Deeply-Pipelined FPGA Clusters

Deep Neural Networks (DNNs) have revolutionized numerous applications, but the demand for ever more performance remains unabated. Scaling DNN computations to larger clusters is generally done by distributing tasks in batch mode using methods such as distributed synchronous SGD. Among the issues with this approach is that to make the distributed cluster work with high utilization, the workload distributed to each node must be large, which implies nontrivial growth in the SGD mini-batch size. In this paper, we propose a framework called FPDeep, which uses a hybrid of model and layer parallelism to configure distributed reconfigurable clusters to train DNNs. This approach has numerous benefits. First, the design does not suffer from batch size growth. Second, novel workload and weight partitioning leads to balanced loads of both among nodes. And third, the entire system is a fine-grained pipeline. This leads to high parallelism and utilization and also minimizes the time features need to be cached while waiting for back-propagation. As a result, storage demand is reduced to the point where only on-chip memory is used for the convolution layers. We evaluate FPDeep with the Alexnet, VGG-16, and VGG-19 benchmarks. Experimental results show that FPDeep has good scalability to a large number of FPGAs, with the limiting factor being the FPGA-to-FPGA bandwidth. With 6 transceivers per FPGA, FPDeep shows linearity up to 83 FPGAs. Energy efficiency is evaluated with respect to GOPs/J. FPDeep provides, on average, 6.36x higher energy efficiency than comparable GPU servers.

cs.LG↗

MVStylizer: An Efficient Edge-Assisted Video Photorealistic Style Transfer System for Mobile Phones

Recent research has made great progress in realizing neural style transfer of images, which denotes transforming an image to a desired style. Many users start to use their mobile phones to record their daily life, and then edit and share the captured images and videos with other users. However, directly applying existing style transfer approaches on videos, i.e., transferring the style of a video frame by frame, requires an extremely large amount of computation resources. It is still technically unaffordable to perform style transfer of videos on mobile phones. To address this challenge, we propose MVStylizer, an efficient edge-assisted photorealistic video style transfer system for mobile phones. Instead of performing stylization frame by frame, only key frames in the original video are processed by a pre-trained deep neural network (DNN) on edge servers, while the rest of stylized intermediate frames are generated by our designed optical-flow-based frame interpolation algorithm on mobile phones. A meta-smoothing module is also proposed to simultaneously upscale a stylized frame to arbitrary resolution and remove style transfer related distortions in these upscaled frames. In addition, for the sake of continuously enhancing the performance of the DNN model on the edge server, we adopt a federated learning scheme to keep retraining each DNN model on the edge server with collected data from mobile clients and syncing with a global DNN model on the cloud server. Such a scheme effectively leverages the diversity of collected data from various mobile clients and efficiently improves the system performance. Our experiments demonstrate that MVStylizer can generate stylized videos with an even better visual quality compared to the state-of-the-art method while achieving 75.5$\times$ speedup for 1920$\times$1080 videos.

eess.IV↗

Visual Localization Using Semantic Segmentation and Depth Prediction

In this paper, we propose a monocular visual localization pipeline leveraging semantic and depth cues. We apply semantic consistency evaluation to rank the image retrieval results and a practical clustering technique to reject estimation outliers. In addition, we demonstrate a substantial performance boost achieved with a combination of multiple feature extractors. Furthermore, by using depth prediction with a deep neural network, we show that a significant amount of falsely matched keypoints are identified and eliminated. The proposed pipeline outperforms most of the existing approaches at the Long-Term Visual Localization benchmark 2020.

cs.CV↗

PoliteCamera: Respecting Strangers' Privacy in Mobile Photographing

Camera is a standard on-board sensor of modern mobile phones. It makes photo taking popular due to its convenience and high resolution. However, when users take a photo of a scenery, a building or a target person, a stranger may also be unintentionally captured in the photo. Such photos expose the location and activity of strangers, and hence may breach their privacy. In this paper, we propose a cooperative mobile photographing scheme called PoliteCamera to protect strangers' privacy. Through the cooperation between a photographer and a stranger, the stranger's face in a photo can be automatically blurred upon his request when the photo is taken. Since multiple strangers nearby the photographer might send out blurring requests but not all of them are in the photo, an adapted balanced convolutional neural network (ABCNN) is proposed to determine whether the requesting stranger is in the photo based on facial attributes. Evaluations demonstrate that the ABCNN can accurately predict facial attributes and PoliteCamera can provide accurate privacy protection for strangers.

cs.CR↗

The AVA-Kinetics Localized Human Actions Video Dataset

This paper describes the AVA-Kinetics localized human actions video dataset. The dataset is collected by annotating videos from the Kinetics-700 dataset using the AVA annotation protocol, and extending the original AVA dataset with these new AVA annotated Kinetics clips. The dataset contains over 230k clips annotated with the 80 AVA action classes for each of the humans in key-frames. We describe the annotation process and provide statistics about the new dataset. We also include a baseline evaluation using the Video Action Transformer Network on the AVA-Kinetics dataset, demonstrating improved performance for action classification on the AVA test set. The dataset can be downloaded from https://research.google.com/ava/

cs.CV↗

CSB-RNN: A Faster-than-Realtime RNN Acceleration Framework with Compressed Structured Blocks

Recurrent neural networks (RNNs) have been widely adopted in temporal sequence analysis, where realtime performance is often in demand. However, RNNs suffer from heavy computational workload as the model often comes with large weight matrices. Pruning schemes have been proposed for RNNs to eliminate the redundant (close-to-zero) weight values. On one hand, the non-structured pruning methods achieve a high pruning rate but introducing computation irregularity (random sparsity), which is unfriendly to parallel hardware. On the other hand, hardware-oriented structured pruning suffers from low pruning rate due to restricted constraints on allowable pruning structure. This paper presents CSB-RNN, an optimized full-stack RNN framework with a novel compressed structured block (CSB) pruning technique. The CSB pruned RNN model comes with both fine pruning granularity that facilitates a high pruning rate and regular structure that benefits the hardware parallelism. To address the challenges in parallelizing the CSB pruned model inference with fine-grained structural sparsity, we propose a novel hardware architecture with a dedicated compiler. Gaining from the architecture-compilation co-design, the hardware not only supports various RNN cell types, but is also able to address the challenging workload imbalance issue and therefore significantly improves the hardware efficiency.

cs.DC↗

Reconfigurable Intelligent Surface (RIS)-Enhanced Two-Way OFDM Communications

In this paper, we focus on the reconfigurable intelligent surface (RIS)-enhanced two-way device-to-device (D2D) multi-pair orthogonal-frequency-division-multiplexing (OFDM) communication systems. Specifically, we maximize the minimum bidirectional weighted sum-rate by jointly optimizing the sub-band allocation, the power allocation and the discrete phase shift (PS) design at the RIS. To tackle the main difficulty of the non-convex PS design at the RIS, we firstly formulate a semi-definite relaxation problem and further devise a low-complexity solution for the PS design by leveraging the projected sub-gradient method. We demonstrate the desirable performance gain for the proposed designs through numerical results.

eess.SP↗

Learning Low-rank Deep Neural Networks via Singular Vector Orthogonality Regularization and Singular Value Sparsification

Modern deep neural networks (DNNs) often require high memory consumption and large computational loads. In order to deploy DNN algorithms efficiently on edge or mobile devices, a series of DNN compression algorithms have been explored, including factorization methods. Factorization methods approximate the weight matrix of a DNN layer with the multiplication of two or multiple low-rank matrices. However, it is hard to measure the ranks of DNN layers during the training process. Previous works mainly induce low-rank through implicit approximations or via costly singular value decomposition (SVD) process on every training step. The former approach usually induces a high accuracy loss while the latter has a low efficiency. In this work, we propose SVD training, the first method to explicitly achieve low-rank DNNs during training without applying SVD on every step. SVD training first decomposes each layer into the form of its full-rank SVD, then performs training directly on the decomposed weights. We add orthogonality regularization to the singular vectors, which ensure the valid form of SVD and avoid gradient vanishing/exploding. Low-rank is encouraged by applying sparsity-inducing regularizers on the singular values of each layer. Singular value pruning is applied at the end to explicitly reach a low-rank model. We empirically show that SVD training can significantly reduce the rank of DNN layers and achieve higher reduction on computation load under the same accuracy, comparing to not only previous factorization methods but also state-of-the-art filter pruning methods.

cs.LG↗

RF-Rhythm: Secure and Usable Two-Factor RFID Authentication

Passive RFID technology is widely used in user authentication and access control. We propose RF-Rhythm, a secure and usable two-factor RFID authentication system with strong resilience to lost/stolen/cloned RFID cards. In RF-Rhythm, each legitimate user performs a sequence of taps on his/her RFID card according to a self-chosen secret melody. Such rhythmic taps can induce phase changes in the backscattered signals, which the RFID reader can detect to recover the user's tapping rhythm. In addition to verifying the RFID card's identification information as usual, the backend server compares the extracted tapping rhythm with what it acquires in the user enrollment phase. The user passes authentication checks if and only if both verifications succeed. We also propose a novel phase-hopping protocol in which the RFID reader emits Continuous Wave (CW) with random phases for extracting the user's secret tapping rhythm. Our protocol can prevent a capable adversary from extracting and then replaying a legitimate tapping rhythm from sniffed RFID signals. Comprehensive user experiments confirm the high security and usability of RF-Rhythm with false-positive and false-negative rates close to zero.

eess.SP↗

Near-Optimal Interference Exploitation 1-Bit Massive MIMO Precoding via Partial Branch-and-Bound

In this paper, we focus on 1-bit precoding for large-scale antenna systems in the downlink based on the concept of constructive interference (CI). By formulating the optimization problem that aims to maximize the CI effect subject to the 1-bit constraint on the transmit signals, we mathematically prove that, when relaxing the 1-bit constraint, the majority of the obtained transmit signals already satisfy the 1-bit constraint. Based on this important observation, we propose a 1-bit precoding method via a partial branch-and-bound (P-BB) approach, where the BB procedure is only performed for the entries that do not comply with the 1-bit constraint. The proposed P-BB enables the use of the BB framework in large-scale antenna scenarios, which was not applicable due to its prohibitive complexity. Numerical results demonstrate a near-optimal error rate performance for the proposed 1-bit precoding algorithm.

eess.SP↗

Data Efficient Training for Reinforcement Learning with Adaptive Behavior Policy Sharing

Deep Reinforcement Learning (RL) is proven powerful for decision making in simulated environments. However, training deep RL model is challenging in real world applications such as production-scale health-care or recommender systems because of the expensiveness of interaction and limitation of budget at deployment. One aspect of the data inefficiency comes from the expensive hyper-parameter tuning when optimizing deep neural networks. We propose Adaptive Behavior Policy Sharing (ABPS), a data-efficient training algorithm that allows sharing of experience collected by behavior policy that is adaptively selected from a pool of agents trained with an ensemble of hyper-parameters. We further extend ABPS to evolve hyper-parameters during training by hybridizing ABPS with an adapted version of Population Based Training (ABPS-PBT). We conduct experiments with multiple Atari games with up to 16 hyper-parameter/architecture setups. ABPS achieves superior overall performance, reduced variance on top 25% agents, and equivalent performance on the best agent compared to conventional hyper-parameter tuning with independent training, even though ABPS only requires the same number of environmental interactions as training a single agent. We also show that ABPS-PBT further improves the convergence speed and reduces the variance.

cs.LG↗

Prediction, Consistency, Curvature: Representation Learning for Locally-Linear Control

Many real-world sequential decision-making problems can be formulated as optimal control with high-dimensional observations and unknown dynamics. A promising approach is to embed the high-dimensional observations into a lower-dimensional latent representation space, estimate the latent dynamics model, then utilize this model for control in the latent space. An important open question is how to learn a representation that is amenable to existing control algorithms? In this paper, we focus on learning representations for locally-linear control algorithms, such as iterative LQR (iLQR). By formulating and analyzing the representation learning problem from an optimal control perspective, we establish three underlying principles that the learned representation should comprise: 1) accurate prediction in the observation space, 2) consistency between latent and observation space dynamics, and 3) low curvature in the latent space transitions. These principles naturally correspond to a loss function that consists of three terms: prediction, consistency, and curvature (PCC). Crucially, to make PCC tractable, we derive an amortized variational bound for the PCC loss function. Extensive experiments on benchmark domains demonstrate that the new variational-PCC learning algorithm benefits from significantly more stable and reproducible training, and leads to superior control performance. Further ablation studies give support to the importance of all three PCC components for learning a good latent space for control.

cs.LG↗

Kinetic Control of Morphology and Composition in Ge/GeSn Core/Shell Nanowires

The growth of Sn-rich group-IV semiconductors at the nanoscale provides new paths for understanding the fundamental properties of metastable GeSn alloys. Here, we demonstrate the effect of the growth conditions on the morphology and composition of Ge/GeSn core/shell nanowires by correlating the experimental observations with a theoretical interpretation based on a multi-scale approach. We show that the cross-sectional morphology of Ge/GeSn core/shell nanowires changes from hexagonal to dodecagonal upon increasing the supply of the Sn precursor. This transformation strongly influences the Sn distribution as a higher Sn content is measured under the {112} growth front. Ab-initio DFT calculations provide an atomic-scale explanation by showing that Sn incorporation is favored at the {112} surfaces, where the Ge bonds are tensile-strained. A phase-field continuum model was developed to reproduce the morphological transformation and the Sn distribution within the wire, shedding light on the complex growth mechanism and unveiling the relation between segregation and faceting. The tunability of the photoluminescence emission with the change in composition and morphology of the GeSn shell highlights the potential of the core/shell nanowire system for opto-electronic devices operating at mid-infrared wavelengths.

cond-mat.mtrl-sci↗

A Parallel Sparse Tensor Benchmark Suite on CPUs and GPUs

Tensor computations present significant performance challenges that impact a wide spectrum of applications ranging from machine learning, healthcare analytics, social network analysis, data mining to quantum chemistry and signal processing. Efforts to improve the performance of tensor computations include exploring data layout, execution scheduling, and parallelism in common tensor kernels. This work presents a benchmark suite for arbitrary-order sparse tensor kernels using state-of-the-art tensor formats: coordinate (COO) and hierarchical coordinate (HiCOO) on CPUs and GPUs. It presents a set of reference tensor kernel implementations that are compatible with real-world tensors and power law tensors extended from synthetic graph generation techniques. We also propose Roofline performance models for these kernels to provide insights of computer platforms from sparse tensor view.

cs.DC↗

Hard superconducting gap and diffusion-induced superconductors in Ge-Si nanowires

We show a hard induced superconducting gap in a Ge-Si nanowire Josephson transistor up to in-plane magnetic fields of $250$ mT, an important step towards creating and detecting Majorana zero modes in this system. A hard induced gap requires a highly homogeneous tunneling heterointerface between the superconducting contacts and the semiconducting nanowire. This is realized by annealing devices at $180$ $^\circ$C during which aluminium inter-diffuses and replaces the germanium in a section of the nanowire. Next to Al, we find a superconductor with lower critical temperature ($T_\mathrm{C}=0.9$ K) and a higher critical field ($B_\mathrm{C}=0.9-1.2$ T). We can therefore selectively switch either superconductor to the normal state by tuning the temperature and the magnetic field and observe that the additional superconductor induces a proximity supercurrent in the semiconducting part of the nanowire even when the Al is in the normal state. In another device where the diffusion of Al rendered the nanowire completely metallic, a superconductor with a much higher critical temperature ($T_\mathrm{C}=2.9$ K) and critical field ($B_\mathrm{C}=3.4$ T) is found. The small size of diffusion-induced superconductors inside nanowires may be of special interest for applications requiring high magnetic fields in arbitrary direction.

cond-mat.mes-hall↗

Improved Knowledge Distillation via Teacher Assistant

Despite the fact that deep neural networks are powerful models and achieve appealing results on many tasks, they are too large to be deployed on edge devices like smartphones or embedded sensor nodes. There have been efforts to compress these networks, and a popular method is knowledge distillation, where a large (teacher) pre-trained network is used to train a smaller (student) network. However, in this paper, we show that the student network performance degrades when the gap between student and teacher is large. Given a fixed student network, one cannot employ an arbitrarily large teacher, or in other words, a teacher can effectively transfer its knowledge to students up to a certain size, not smaller. To alleviate this shortcoming, we introduce multi-step knowledge distillation, which employs an intermediate-sized network (teacher assistant) to bridge the gap between the student and the teacher. Moreover, we study the effect of teacher assistant size and extend the framework to multi-step distillation. Theoretical analysis and extensive experiments on CIFAR-10,100 and ImageNet datasets and on CNN and ResNet architectures substantiate the effectiveness of our proposed approach.

cs.LG↗

Hybrid Precoding Design for Reconfigurable Intelligent Surface aided mmWave Communication Systems

In this letter, we focus on the hybrid precoding (HP) design for the reconfigurable intelligent surface (RIS) aided multi-user (MU) millimeter wave (mmWave) communication systems. Specifically, we aim to minimize the mean-squared-error (MSE) between the received symbols and the transmitted symbols by jointly optimizing the analog-digital HP at the base-station (BS) and the phase shifts (PSs) at the RIS, where the non-convex element-wise constant-modulus constraints for the analog precoding and the PSs are tackled by resorting to the gradient-projection (GP) method. We analytically prove the convergence of the proposed algorithm and demonstrate the desirable performance gain for the proposed design through numerical results.

cs.NI↗