Searcharxiv⌕ Search

arXiv subjects

Ang Li

Publications and source records attributed to Ang Li.

At least 559 records · Page 31Linked to original sources

Robust Hybrid Precoding for Beam Misalignment in Millimeter-Wave Communications

In this paper, we focus on the phenomenon of beam misalignment in Millimeter-wave (mmWave) multi-receiver communication systems, and propose robust hybrid precoding designs that alleviate the performance loss caused by this effect. We consider two distinct design methodologies: I) the synthesis of a `flat mainlobe' beam model which maximizes the minimum effective array gain over the beam misalignment range, and II) the inclusion of the `error statistics' into the design, where the array response incorporating the distribution of the misalignment error is derived. For both design methodologies, we propose a hybrid precoding design that approximates the robust fully-digital precoder, which is obtained via alternating optimization based on the gradient projection (GP) method. We also propose a low-complexity alternative to the GP algorithm based on the least square projection (LSP), and we further deploy a second-stage digital precoder to mitigate any residual inter-receiver interference after the hybrid analog-digital precoding. Numerical results show that the robust hybrid precoding designs can effectively alleviate the performance degradation incurred by beam misalignment.

cs.NI↗

Hybrid Precoder and Combiner for Imperfect Beam Alignment in mmWave MIMO Systems

In this letter, we aim to design a robust hybrid precoder and combiner against beam misalignment in millimeter-wave (mmWave) communication systems. We consider the inclusion of the `error statistics' into the precoder and combiner design, where the array response that incorporates the distribution of the misalignment error is first derived. An iterative algorithm is then proposed to design the robust hybrid precoder and combiner to maximize the array gain in the presence of beam misalignment. To further enhance the spectral efficiency, a second-stage digital precoder and combiner are included to mitigate the inter-stream interference. Numerical results show that the proposed robust hybrid precoder and combiner design can effectively alleviate the performance degradation incurred by beam misalignment.

eess.SP↗

Evaluating Modern GPU Interconnect: PCIe, NVLink, NV-SLI, NVSwitch and GPUDirect

High performance multi-GPU computing becomes an inevitable trend due to the ever-increasing demand on computation capability in emerging domains such as deep learning, big data and planet-scale simulations. However, the lack of deep understanding on how modern GPUs can be connected and the real impact of state-of-the-art interconnect technology on multi-GPU application performance become a hurdle. In this paper, we fill the gap by conducting a thorough evaluation on five latest types of modern GPU interconnects: PCIe, NVLink-V1, NVLink-V2, NVLink-SLI and NVSwitch, from six high-end servers and HPC platforms: NVIDIA P100-DGX-1, V100-DGX-1, DGX-2, OLCF's SummitDev and Summit supercomputers, as well as an SLI-linked system with two NVIDIA Turing RTX-2080 GPUs. Based on the empirical evaluation, we have observed four new types of GPU communication network NUMA effects: three are triggered by NVLink's topology, connectivity and routing, while one is caused by PCIe chipset design issue. These observations indicate that, for an application running in a multi-GPU node, choosing the right GPU combination can impose considerable impact on GPU communication efficiency, as well as the application's overall performance. Our evaluation can be leveraged in building practical multi-GPU performance models, which are vital for GPU task allocation, scheduling and migration in a shared environment (e.g., AI cloud and HPC centers), as well as communication-oriented performance tuning.

cs.AR↗

PASTA: A Parallel Sparse Tensor Algorithm Benchmark Suite

Tensor methods have gained increasingly attention from various applications, including machine learning, quantum chemistry, healthcare analytics, social network analysis, data mining, and signal processing, to name a few. Sparse tensors and their algorithms become critical to further improve the performance of these methods and enhance the interpretability of their output. This work presents a sparse tensor algorithm benchmark suite (PASTA) for single- and multi-core CPUs. To the best of our knowledge, this is the first benchmark suite for sparse tensor world. PASTA targets on: 1) helping application users to evaluate different computer systems using its representative computational workloads; 2) providing insights to better utilize existed computer architecture and systems and inspiration for the future design. This benchmark suite is publicly released https://gitlab.com/tensorworld/pasta.

cs.DC↗

A Generalized Framework for Population Based Training

Population Based Training (PBT) is a recent approach that jointly optimizes neural network weights and hyperparameters which periodically copies weights of the best performers and mutates hyperparameters during training. Previous PBT implementations have been synchronized glass-box systems. We propose a general, black-box PBT framework that distributes many asynchronous "trials" (a small number of training steps with warm-starting) across a cluster, coordinated by the PBT controller. The black-box design does not make assumptions on model architectures, loss functions or training procedures. Our system supports dynamic hyperparameter schedules to optimize both differentiable and non-differentiable metrics. We apply our system to train a state-of-the-art WaveNet generative model for human voice synthesis. We show that our PBT system achieves better accuracy, less sensitivity and faster convergence compared to existing methods, given the same computational resource.

cs.AI↗

Quark mean-field model for nuclear matter with or without bag

We propose the new quark mean-field bag (QMFB) model by incorporating the bag confinement mechanism in the original quark mean-field model. Nuclear matter and neutron star properties are studied with the QMFB model. For the study of the bag effect, we newly fit 12 parameter sets by reproducing the empirical saturation properties of nuclear matter. Quark confinement is found to be mainly demonstrated by the bag after it is included in the model, instead of the confining potential. For nuclear matter, the bag decreases the binding energy and increases the symmetry energy. For neutron star, the bag affects significantly the radius $R$ of a $1.4M_\odot$ star, with the maximum mass only slightly modified. The bag also has a large suppression effect on the well-accepted $R$ vs $L$ dependence, with $L$ the symmetry energy slope at the saturation density.

nucl-th↗

Multiplexing More Streams in the MU-MISO Downlink by Interference Exploitation Precoding

In this paper, we study the interference exploitation precoding for the scenario where the number of streams simultaneously transmitted by the base station (BS) is larger than that of transmit antennas at the BS, and derive the optimal precoding structure by employing the pseudo inverse. We show that the optimal pre-scaling vector is equal to a linear combination of the right singular vectors that correspond to zero singular values of the coefficient matrix. By formulating the dual problem, the optimal precoding matrix can be expressed as a function of the dual variables in a closed form, and an equivalent quadratic programming (QP) formulation is further derived for computational complexity reduction. Numerical results validate our analysis and demonstrate significant performance improvements for interference exploitation precoding for the considered scenario.

eess.SP↗

Dense matter with eXTP

In this White Paper we present the potential of the Enhanced X-ray Timing and Polarimetry (eXTP) mission for determining the nature of dense matter; neutron star cores host an extreme density regime which cannot be replicated in a terrestrial laboratory. The tightest statistical constraints on the dense matter equation of state will come from pulse profile modelling of accretion-powered pulsars, burst oscillation sources, and rotation-powered pulsars. Additional constraints will derive from spin measurements, burst spectra, and properties of the accretion flows in the vicinity of the neutron star. Under development by an international Consortium led by the Institute of High Energy Physics of the Chinese Academy of Science, the eXTP mission is expected to be launched in the mid 2020s.

astro-ph.HE↗

Interference Exploitation Precoding for Multi-Level Modulations: Closed-Form Solutions

In this paper, we study closed-form interference-exploitation precoding for multi-level modulations in the downlink of multi-user multiple-input single-output (MU-MISO) systems. We consider two distinct cases: first, for the case where the number of served users is not larger than the number of transmit antennas at the base station (BS), we mathematically derive the optimal precoding structure based on the Karush-Kuhn-Tucker (KKT) conditions. By formulating the dual problem, the precoding problem for multi-level modulations is transformed into a pre-scaling operation using quadratic programming (QP) optimization. We further consider the case where the number of served users is larger than the number of transmit antennas at the BS. By employing the pseudo inverse, we show that the optimal solution of the pre-scaling vector is equivalent to a linear combination of the right singular vectors corresponding to zero singular values, and derive the equivalent QP formulation. We also present the condition under which multiplexing more streams than the number of transmit antennas is achievable. For both considered scenarios, we propose a modified iterative algorithm to obtain the optimal precoding matrix, as well as a sub-optimal closed-form precoder. Numerical results validate our derivations on the optimal precoding structures for multi-level modulations, and demonstrate the superiority of interference-exploitation precoding for both scenarios.

cs.IT↗

1-Bit Massive MIMO Downlink Based on Constructive Interference

In this paper, we focus on the multiuser massive multiple-input single-output (MISO) downlink with low-cost 1-bit digital-to-analog converters (DACs) for PSK modulation, and propose a low-complexity refinement process that is applicable to any existing 1-bit precoding approaches based on the constructive interference (CI) formulation. With the decomposition of the signals along the detection thresholds, we first formulate a simple symbol-scaling method as the performance metric. The low-complexity refinement approach is subsequently introduced, where we aim to improve the introduced symbol-scaling performance metric by modifying the transmit signal on one antenna at a time. Numerical results validate the effectiveness of the proposed refinement method on existing approaches for massive MIMO with 1-bit DACs, and the performance improvements are most significant for the low-complexity quantized zero-forcing (ZF) method.

eess.SP↗

Hybrid Analog-Digital Precoding for Interference Exploitation

We study the multi-user massive multiple-input-single-output (MISO) and focus on the downlink systems where the base station (BS) employs hybrid analog-digital precoding with low-cost 1-bit digital-to-analog converters (DACs). In this paper, we propose a hybrid downlink transmission scheme where the analog precoder is formed based on the SVD decomposition. In the digital domain, instead of designing a linear transmit precoding matrix, we directly design the transmit signals by exploiting the concept of constructive interference. The optimization problem is then formulated based on the geometry of the modulation constellations and is shown to be non-convex. We relax the above optimization and show that the relaxed optimization can be transformed into a linear programming that can be efficiently solved. Numerical results validate the superiority of the proposed scheme for the hybrid massive MIMO downlink systems.

eess.SP↗

Consistency-aware Shading Orders Selective Fusion for Intrinsic Image Decomposition

We address the problem of decomposing a single image into reflectance and shading. The difficulty comes from the fact that the components of image---the surface albedo, the direct illumination, and the ambient illumination---are coupled heavily in observed image. We propose to infer the shading by ordering pixels by their relative brightness, without knowing the absolute values of the image components beforehand. The pairwise shading orders are estimated in two ways: brightness order and low-order fittings of local shading field. The brightness order is a non-local measure, which can be applied to any pair of pixels including those whose reflectance and shading are both different. The low-order fittings are used for pixel pairs within local regions of smooth shading. Together, they can capture both global order structure and local variations of the shading. We propose a Consistency-aware Selective Fusion (CSF) to integrate the pairwise orders into a globally consistent order. The iterative selection process solves the conflicts between the pairwise orders obtained by different estimation methods. Inconsistent or unreliable pairwise orders will be automatically excluded from the fusion to avoid polluting the global order. Experiments on the MIT Intrinsic Image dataset show that the proposed model is effective at recovering the shading including deep shadows. Our model also works well on natural images from the IIW dataset, the UIUC Shadow dataset and the NYU-Depth dataset, where the colors of direct lights and ambient lights are quite different.

cs.CV↗

Privacy-Preserving Multiparty Learning For Logistic Regression

In recent years, machine learning techniques are widely used in numerous applications, such as weather forecast, financial data analysis, spam filtering, and medical prediction. In the meantime, massive data generated from multiple sources further improve the performance of machine learning tools. However, data sharing from multiple sources brings privacy issues for those sources since sensitive information may be leaked in this process. In this paper, we propose a framework enabling multiple parties to collaboratively and accurately train a learning model over distributed datasets while guaranteeing the privacy of data sources. Specifically, we consider logistic regression model for data training and propose two approaches for perturbing the objective function to preserve ε-differential privacy. The proposed solutions are tested on real datasets, including Bank Marketing and Credit Card Default prediction. Experimental results demonstrate that the proposed multiparty learning framework is highly efficient and accurate.

cs.CR↗

Privacy-Preserving Outsourcing of Large-Scale Nonlinear Programming to the Cloud

The increasing massive data generated by various sources has given birth to big data analytics. Solving large-scale nonlinear programming problems (NLPs) is one important big data analytics task that has applications in many domains such as transport and logistics. However, NLPs are usually too computationally expensive for resource-constrained users. Fortunately, cloud computing provides an alternative and economical service for resource-constrained users to outsource their computation tasks to the cloud. However, one major concern with outsourcing NLPs is the leakage of user's private information contained in NLP formulations and results. Although much work has been done on privacy-preserving outsourcing of computation tasks, little attention has been paid to NLPs. In this paper, we for the first time investigate secure outsourcing of general large-scale NLPs with nonlinear constraints. A secure and efficient transformation scheme at the user side is proposed to protect user's private information; at the cloud side, generalized reduced gradient method is applied to effectively solve the transformed large-scale NLPs. The proposed protocol is implemented on a cloud computing testbed. Experimental evaluations demonstrate that significant time can be saved for users and the proposed mechanism has the potential for practical use.

cs.CR↗

PhotoSafer: Content-Based and Context-Aware Private Photo Protection for Smartphones

Nowadays many people store photos in smartphones. Many of the photos contain sensitive, private information, such as a photocopy of driver's license and credit card. An arising privacy concern is with the unauthorized accesses to such private photos by installed apps. Coarse-grained access control systems such as the Android permission system offer all-or-nothing access to photos stored on smartphones, and users are unaware of the exact behavior of installed apps. Our analysis finds that 82% of the top 200 free apps in a popular Android app store have complete access to stored photos and network on a user's smartphone, which indicates possible private photo leakage. In addition, our user survey reveals that 87.5% of the 112 respondents are not aware that certain apps can access their photos without informing users, and all the respondents believe that the stored photos on their smartphones contain different types of private information. Hence, we propose PhotoSafer, a content-based, context-aware private photo protection system for Android phones. PhotoSafer can detect private photos based on photo content with a well-trained deep convolutional neural network, and control access to photos based on system status (e.g., screen locked or not) and app-running status (e.g., app in the background). Evaluations demonstrate that PhotoSafer can accurately identify private photos in real time. The efficacy and efficiency of the implemented prototype system show the potential for practical use.

cs.CR↗

SymmNet: A Symmetric Convolutional Neural Network for Occlusion Detection

Detecting the occlusion from stereo images or video frames is important to many computer vision applications. Previous efforts focus on bundling it with the computation of disparity or optical flow, leading to a chicken-and-egg problem. In this paper, we leverage convolutional neural network to liberate the occlusion detection task from the interleaved, traditional calculation framework. We propose a Symmetric Network (SymmNet) to directly exploit information from an image pair, without estimating disparity or motion in advance. The proposed network is structurally left-right symmetric to learn the binocular occlusion simultaneously, aimed at jointly improving both results. The comprehensive experiments show that our model achieves state-of-the-art results on detecting the stereo and motion occlusion.

cs.CV↗

Josephson effect in a few-hole quantum dot

We use a Ge-Si core-shell nanowire to realise a Josephson field-effect transistor with highly transparent contacts to superconducting leads. By changing the electric field we gain access to two distinct regimes not combined before in a single device: In the accumulation mode the device is highly transparent and the supercurrent is carried by multiple subbands, while near depletion supercurrent is carried by single-particle levels of a strongly coupled quantum dot operating in the few-hole regime. These results establish Ge-Si nanowires as an important platform for hybrid superconductor-semiconductor physics and Majorana fermions.

cond-mat.mes-hall↗

C-WSL: Count-guided Weakly Supervised Localization

We introduce count-guided weakly supervised localization (C-WSL), an approach that uses per-class object count as a new form of supervision to improve weakly supervised localization (WSL). C-WSL uses a simple count-based region selection algorithm to select high-quality regions, each of which covers a single object instance during training, and improves existing WSL methods by training with the selected regions. To demonstrate the effectiveness of C-WSL, we integrate it into two WSL architectures and conduct extensive experiments on VOC2007 and VOC2012. Experimental results show that C-WSL leads to large improvements in WSL and that the proposed approach significantly outperforms the state-of-the-art methods. The results of annotation experiments on VOC2007 suggest that a modest extra time is needed to obtain per-class object counts compared to labeling only object categories in an image. Furthermore, we reduce the annotation time by more than $2\times$ and $38\times$ compared to center-click and bounding-box annotations.

cs.CV↗