Searcharxiv⌕ Search

arXiv subjects

Ang Li

Publications and source records attributed to Ang Li.

At least 487 records · Page 27Linked to original sources

BCNN: Binary Complex Neural Network

Binarized neural networks, or BNNs, show great promise in edge-side applications with resource limited hardware, but raise the concerns of reduced accuracy. Motivated by the complex neural networks, in this paper we introduce complex representation into the BNNs and propose Binary complex neural network -- a novel network design that processes binary complex inputs and weights through complex convolution, but still can harvest the extraordinary computation efficiency of BNNs. To ensure fast convergence rate, we propose novel BCNN based batch normalization function and weight initialization function. Experimental results on Cifar10 and ImageNet using state-of-the-art network models (e.g., ResNet, ResNetE and NIN) show that BCNN can achieve better accuracy compared to the original BNN models. BCNN improves BNN by strengthening its learning capability through complex representation and extending its applicability to complex-valued input data. The source code of BCNN will be released on GitHub.

cs.NE↗

Constraints on the maximum mass of neutron stars with a quark core from GW170817 and NICER PSR J0030+0451 data

We perform a Bayesian analysis of the maximum mass $M_{\rm TOV}$ of neutron stars with a quark core, incorporating the observational data from tidal deformability of the GW170817 binary neutron star merger as detected by LIGO/Virgo and the mass and radius of PSR J0030+0451 as detected by \nicer. The analysis is performed under the assumption that the hadron-quark phase transition is of first order, where the low-density hadronic matter described in a unified manner by the soft QMF or the stiff DD2 equation of state (EOS) transforms into a high-density phase of quark matter modeled by the generic "Constant-sound-speed" (CSS) parameterization. The mass distribution measured for the $2.14 \,{\rm M}_{\odot}$ pulsar, MSP J0740+6620, is used as the lower limit on $M_{\rm TOV}$. We find the most probable values of the hybrid star maximum mass are $M_{\rm TOV}=2.36^{+0.49}_{-0.26}\,{\rm M}_{\odot}$ ($2.39^{+0.47}_{-0.28}\,{\rm M}_{\odot}$) for QMF (DD2), with an absolute upper bound around $2.85\,{\rm M}_{\odot}$, to the $90\%$ posterior credible level. Such results appear robust with respect to the uncertainties in the hadronic EOS. We also discuss astrophysical implications of this result, especially on the post-merger product of GW170817, short gamma-ray bursts, and other likely binary neutron star mergers.

astro-ph.HE↗

Symbol-Level Precoding Made Practical for Multi-Level Modulations via Block-Level Rescaling

In this paper, we propose an interference exploitation symbol-level precoding (SLP) method for multi-level modulations via an in-block power allocation scheme to greatly reduce the signaling overhead. Existing SLP approaches require the symbol-level broadcast of the rescaling factor to the users for correct demodulation, which hinders the practical implementation of SLP. The proposed approach allows a block-level broadcast of the rescaling factor as done in traditional block-level precoding, greatly reducing the signaling overhead for SLP without sacrificing the performance. Our derivations further show that the proposed in-block power allocation enjoys an exact closed-form solution and thus does not increase the complexity at the base station (BS). In addition to the significant alleviation of the signaling overhead validated by the effective throughput result, numerical results demonstrate that the proposed power allocation approach also improves the error-rate performance of the existing SLP. Accordingly, the proposed approach enables the practical use of SLP in multi-level modulations.

cs.IT↗

Fluid forces and vortex patterns of an oscillating cylinder pair in still water with both side-by-side and tandem configurations

Models of cylinders in the oscillatory flow can be found virtually everywhere in the marine industry, such as pump towers experiencing sloshing load in a LNG ship liquid tank. However, compared to the problem of a cylinder in the uniform flow, a cylinder in the oscillatory flow is less studied, let alone multiple cylinders. Therefore, we experimentally and numerically studied two identical circular cylinders oscillating in the still water with either a side-by-side or a tandem configuration for a wide range of Keulegan-Carpenter number and Stokes number $β$. The experiment result shows that the hydrodynamic performance of an oscillating cylinder pair in the still water is greatly altered due to the interference between the multiple structures with different configurations. In specific, compared to the single-cylinder case, the drag coefficient is greatly enhanced when two cylinders are placed side-by-side at a small gap ratio, while dual cylinders in a tandem configuration obtain a smaller drag coefficient and oscillating lift coefficient. In order to reveal the detailed flow physics that result in significant fluid forces alternations, the detailed flow visualization is provided by the numerical simulation: the small gap between two cylinders in a side-by-side configuration will result in a strong gap jet that enhances the energy dissipation and increase the drag, while due to the flow blocking effect for two cylinders in a tandem configuration, the drag coefficient decreases.

physics.flu-dyn↗

Sound velocity in dense stellar matter with strangeness and compact stars

The phase state of dense matter in the intermediate density range ($\sim$1-10 times the nuclear saturation density) is both intriguing and unclear and could have important observable effects in the present gravitational wave era of neutron stars. As the matter density increases in compact stars, the sound velocity is expected to approach the conformal limit ($c_s/c=1/\sqrt{3}$) at high densities and should also fulfill the causality limit ($c_s/c<1$). However, its detailed behavior remains a hot topic of debate. It was suggested that the sound velocity of dense matter could be an important indicator for a deconfinement phase transition, where a particular shape might be expected for its density dependence. In this work, we explore the general properties of the sound velocity and the adiabatic index of dense matter in hybrid stars, as well as in neutron stars and quark stars. Various conditions are employed for hadron-quark phase transition with varying interface tension. We find that the expected behavior of the sound velocity can also be achieved by the nonperturbative properties of the quark phase, in addition to a deconfinement phase transition. And it leads to a more compact star with a similar mass. We then propose a new class of quark star equation of states, which could be tested by future high-precision radius measurements of pulsar-like objects.

nucl-th↗

High-resolution ARPES endstation for in-situ electronic structure investigations at SSRF

Angle-resolved photoemission spectroscopy (ARPES) is one of the most powerful experimental techniques in condensed matter physics. Synchrotron ARPES, which uses photons with high flux and continuously tunable energy, has become particularly important. However, an excellent synchrotron ARPES system must have features such as a small beam spot, super-high energy resolution, and a user-friendly operation interface. A synchrotron beamline and an endstation (BL03U) were designed and constructed at the Shanghai Synchrotron Radiation Facility. The beam spot size at the sample position is 7.5 (V) $μ$m $\times$ 67 (H) $μ$m, and the fundamental photon range is 7-165 eV; the ARPES system enables photoemission with an energy resolution of 2.67 meV@21.2 eV. In addition, the ARPES system of this endstation is equipped with a six-axis cryogenic sample manipulator (the lowest temperature is 7 K) and is integrated with an oxide molecular beam epitaxy system and a scanning tunneling microscope, which can provide an advanced platform for in-situ characterization of the fine electronic structure of condensed matter.

physics.ins-det↗

$R$-mode Stability of GW190814's Secondary Component as a Supermassive and Superfast Pulsar

The nature of GW190814's secondary component $m_2$ of mass $(2.50-2.67)\,\text{M}_{\odot}$ in the mass gap between the currently known maximum mass of neutron stars and the minimum mass of black holes is currently under hot debate. Among the many possibilities proposed in the literature, the $m_2$ was suggested as a superfast pulsar while its r-mode stability against the run-away gravitational radiation through the Chandrasekhar-Friedman-Schutz mechanism is still unknown. Using those fulfilling all currently known astrophysical and nuclear physics constraints among a sample of 33 unified equation of states (EOSs) constructed previously by Fortin {\it et al.} (2016) using the same nuclear interactions from the crust to the core consistently, we compare the minimum frequency required for the $m_2$ to rotationally sustain a mass higher than $2.50\,\text{M}_{\odot}$ with the critical frequency above which the r-mode instability occurs. We use two extreme damping models assuming the crust is either perfectly rigid or elastic. Using the stability of 19 observed low-mass x-ray binaries as an indication that the rigid crust damping of the r-mode dominates within the models studied, we find that the $m_2$ is r-mode stable while rotating with a frequency higher than 870.2 Hz (0.744 times its Kepler frequency of 1169.6 Hz) as long as its temperate is lower than about $3.9\times 10^7 K$, further supporting the proposal that GW190814's secondary component is a supermassive and superfast pulsar.

astro-ph.HE↗

Optimization and Generalization of Regularization-Based Continual Learning: a Loss Approximation Viewpoint

Neural networks have achieved remarkable success in many cognitive tasks. However, when they are trained sequentially on multiple tasks without access to old data, their performance on early tasks tend to drop significantly. This problem is often referred to as catastrophic forgetting, a key challenge in continual learning of neural networks. The regularization-based approach is one of the primary classes of methods to alleviate catastrophic forgetting. In this paper, we provide a novel viewpoint of regularization-based continual learning by formulating it as a second-order Taylor approximation of the loss function of each task. This viewpoint leads to a unified framework that can be instantiated to derive many existing algorithms such as Elastic Weight Consolidation and Kronecker factored Laplace approximation. Based on this viewpoint, we study the optimization aspects (i.e., convergence) as well as generalization properties (i.e., finite-sample guarantees) of regularization-based continual learning. Our theoretical results indicate the importance of accurate approximation of the Hessian matrix. The experimental results on several benchmarks provide empirical validation of our theoretical findings.

cs.LG↗

PredCoin: Defense against Query-based Hard-label Attack

Many adversarial attacks and defenses have recently been proposed for Deep Neural Networks (DNNs). While most of them are in the white-box setting, which is impractical, a new class of query-based hard-label (QBHL) black-box attacks pose a significant threat to real-world applications (e.g., Google Cloud, Tencent API). Till now, there has been no generalizable and practical approach proposed to defend against such attacks. This paper proposes and evaluates PredCoin, a practical and generalizable method for providing robustness against QBHL attacks. PredCoin poisons the gradient estimation step, an essential component of most QBHL attacks. PredCoin successfully identifies gradient estimation queries crafted by an attacker and introduces uncertainty to the output. Extensive experiments show that PredCoin successfully defends against four state-of-the-art QBHL attacks across various settings and tasks while preserving the target model's overall accuracy. PredCoin is also shown to be robust and effective against several defense-aware attacks, which may have full knowledge regarding the internal mechanisms of PredCoin.

cs.CR↗

Growth and Strain Relaxation Mechanisms of InAs/InP/GaAsSb Core-Dual-Shell Nanowires

The combination of core/shell geometry and band gap engineering in nanowire heterostructures can be employed to realize systems with novel transport and optical properties. Here, we report on the growth of InAs/InP/GaAsSb core-dual-shell nanowires by catalyst-free chemical beam epitaxy on Si(111) substrates. Detailed morphological, structural, and compositional analyses of the nanowires as a function of growth parameters were carried out by scanning and transmission electron microscopy and by energy-dispersive X-ray spectroscopy. Furthermore, by combining the scanning transmission electron microscopy-Moire technique with geometric phase analysis, we studied the residual strain and the relaxation mechanisms in this system. We found that InP shell facets are well-developed along all the crystallographic directions only when the nominal thickness is above 1 nm, suggesting an island-growth mode. Moreover, the crystallographic analysis indicates that both InP and GaAsSb shells grow almost coherently to the InAs core along the 112 direction and elastically compressed along the 110 direction. For InP shell thickness above 8 nm, some dislocations and roughening occur at the interfaces. This study provides useful general guidelines for the fabrication of high-quality devices based on these core-dual-shell nanowires.

cond-mat.mtrl-sci↗

Cramér-Rao Bound Optimization for Joint Radar-Communication Design

In this paper, we propose multi-input multi-output (MIMO) beamforming designs towards joint radar sensing and multi-user communications. We employ the Cramér-Rao bound (CRB) as a performance metric of target estimation, under both point and extended target scenarios. We then propose minimizing the CRB of radar sensing while guaranteeing a pre-defined level of signal-to-interference-plus-noise ratio (SINR) for each communication user. For the single-user scenario, we derive a closed form for the optimal solution for both cases of point and extended targets. For the multi-user scenario, we show that both problems can be relaxed into semidefinite programming by using the semidefinite relaxation approach, and prove that the global optimum can always be obtained. Finally, we demonstrate numerically that the globally optimal solutions are reachable via the proposed methods, which provide significant gains in target estimation performance over state-of-the-art benchmarks.

eess.SP↗

1-Bit Massive MIMO Transmission: Embracing Interference with Symbol-Level Precoding

The deployment of large-scale antenna arrays for cellular base stations (BSs), termed as `Massive MIMO', has been a key enabler for meeting the ever-increasing capacity requirement for 5G communication systems and beyond. Despite their promising performance, fully-digital massive MIMO systems require a vast amount of hardware components including radio frequency chains, power amplifiers, digital-to-analog converters (DACs), etc., resulting in a huge increase in terms of the total power consumption and hardware costs for cellular BSs. Towards both spectrally-efficient and energy-efficient massive MIMO deployment, a number of hardware limited architectures have been proposed, including hybrid analog-digital structures, constant-envelope transmission, and use of low-resolution DACs. In this paper, we overview the recent interest in improving the error-rate performance of massive MIMO systems deployed with 1-bit DACs through precoding at the symbol level. This line of research goes beyond traditional interference suppression or cancellation techniques by managing interference on a symbol-by-symbol basis. This provides unique opportunities for interference-aware precoding tailored for practical massive MIMO systems. Firstly, we characterize constructive interference (CI) and elaborate on how CI can benefit the 1-bit signal design by exploiting the traditionally undesired multi-user interference as well as the interference from imperfect hardware components. Subsequently, we overview several solutions for 1-bit signal design to illustrate the gains achievable by exploiting CI. Finally, we identify some challenges and future research directions for 1-bit massive MIMO systems that are yet to be explored.

cs.IT↗

On Provable Backdoor Defense in Collaborative Learning

As collaborative learning allows joint training of a model using multiple sources of data, the security problem has been a central concern. Malicious users can upload poisoned data to prevent the model's convergence or inject hidden backdoors. The so-called backdoor attacks are especially difficult to detect since the model behaves normally on standard test data but gives wrong outputs when triggered by certain backdoor keys. Although Byzantine-tolerant training algorithms provide convergence guarantee, provable defense against backdoor attacks remains largely unsolved. Methods based on randomized smoothing can only correct a small number of corrupted pixels or labels; methods based on subset aggregation cause a severe drop in classification accuracy due to low data utilization. We propose a novel framework that generalizes existing subset aggregation methods. The framework shows that the subset selection process, a deciding factor for subset aggregation methods, can be viewed as a code design problem. We derive the theoretical bound of data utilization ratio and provide optimal code construction. Experiments on non-IID versions of MNIST and CIFAR-10 show that our method with optimal codes significantly outperforms baselines using non-overlapping partition and random selection. Additionally, integration with existing coding theory results shows that special codes can track the location of the attackers. Such capability provides new countermeasures to backdoor attacks.

cs.CR↗

Hermes: Decentralized Dynamic Spectrum Access System for Massive Devices Deployment in 5G

With the incoming 5G network, the ubiquitous Internet of Things (IoT) devices can benefit our daily life, such as smart cameras, drones, etc. With the introduction of the millimeter-wave band and the thriving number of IoT devices, it is critical to design new dynamic spectrum access (DSA) system to coordinate the spectrum allocation across massive devices in 5G. In this paper, we present Hermes, the first decentralized DSA system for massive devices deployment. Specifically, we propose an efficient multi-agent reinforcement learning algorithm and introduce a novel shuffle mechanism, addressing the drawbacks of collision and fairness in existing decentralized systems. We implement Hermes in 5G network via simulations. Extensive evaluations show that Hermes significantly reduces collisions and improves fairness compared to the state-of-the-art decentralized methods. Furthermore, Hermes is able to adapt the environmental changes within 0.5 seconds, showing its deployment practicability in dynamic environment of 5G.

cs.NI↗

Accelerating Binarized Neural Networks via Bit-Tensor-Cores in Turing GPUs

Despite foreseeing tremendous speedups over conventional deep neural networks, the performance advantage of binarized neural networks (BNNs) has merely been showcased on general-purpose processors such as CPUs and GPUs. In fact, due to being unable to leverage bit-level-parallelism with a word-based architecture, GPUs have been criticized for extremely low utilization (1%) when executing BNNs. Consequently, the latest tensorcores in NVIDIA Turing GPUs start to experimentally support bit computation. In this work, we look into this brand new bit computation capability and characterize its unique features. We show that the stride of memory access can significantly affect performance delivery and a data-format co-design is highly desired to support the tensorcores for achieving superior performance than existing software solutions without tensorcores. We realize the tensorcore-accelerated BNN design, particularly the major functions for fully-connect and convolution layers -- bit matrix multiplication and bit convolution. Evaluations on two NVIDIA Turing GPUs show that, with ResNet-18, our BTC-BNN design can process ImageNet at a rate of 5.6K images per second, 77% faster than state-of-the-art. Our BNN approach is released on https://github.com/pnnl/TCBNN.

cs.DC↗

Fast and Scalable Sparse Triangular Solver for Multi-GPU Based HPC Architectures

Designing efficient and scalable sparse linear algebra kernels on modern multi-GPU based HPC systems is a daunting task due to significant irregular memory references and workload imbalance across the GPUs. This is particularly the case for Sparse Triangular Solver (SpTRSV) which introduces additional two-dimensional computation dependencies among subsequent computation steps. Dependency information is exchanged and shared among GPUs, thus warrant for efficient memory allocation, data partitioning, and workload distribution as well as fine-grained communication and synchronization support. In this work, we demonstrate that directly adopting unified memory can adversely affect the performance of SpTRSV on multi-GPU architectures, despite linking via fast interconnect like NVLinks and NVSwitches. Alternatively, we employ the latest NVSHMEM technology based on Partitioned Global Address Space programming model to enable efficient fine-grained communication and drastic synchronization overhead reduction. Furthermore, to handle workload imbalance, we propose a malleable task-pool execution model which can further enhance the utilization of GPUs. By applying these techniques, our experiments on the NVIDIA multi-GPU supernode V100-DGX-1 and DGX-2 systems demonstrate that our design can achieve on average 3.53x (up to 9.86x) speedup on a DGX-1 system and 3.66x (up to 9.64x) speedup on a DGX-2 system with 4-GPUs over the Unified-Memory design. The comprehensive sensitivity and scalability studies also show that the proposed zero-copy SpTRSV is able to fully utilize the computing and communication resources of the multi-GPU system.

cs.DC↗

GraphFL: A Federated Learning Framework for Semi-Supervised Node Classification on Graphs

Graph-based semi-supervised node classification (GraphSSC) has wide applications, ranging from networking and security to data mining and machine learning, etc. However, existing centralized GraphSSC methods are impractical to solve many real-world graph-based problems, as collecting the entire graph and labeling a reasonable number of labels is time-consuming and costly, and data privacy may be also violated. Federated learning (FL) is an emerging learning paradigm that enables collaborative learning among multiple clients, which can mitigate the issue of label scarcity and protect data privacy as well. Therefore, performing GraphSSC under the FL setting is a promising solution to solve real-world graph-based problems. However, existing FL methods 1) perform poorly when data across clients are non-IID, 2) cannot handle data with new label domains, and 3) cannot leverage unlabeled data, while all these issues naturally happen in real-world graph-based problems. To address the above issues, we propose the first FL framework, namely GraphFL, for semi-supervised node classification on graphs. Our framework is motivated by meta-learning methods. Specifically, we propose two GraphFL methods to respectively address the non-IID issue in graph data and handle the tasks with new label domains. Furthermore, we design a self-training method to leverage unlabeled graph data. We adopt representative graph neural networks as GraphSSC methods and evaluate GraphFL on multiple graph datasets. Experimental results demonstrate that GraphFL significantly outperforms the compared FL baseline and GraphFL with self-training can obtain better performance.

cs.LG↗

Provable Defense against Privacy Leakage in Federated Learning from Representation Perspective

Federated learning (FL) is a popular distributed learning framework that can reduce privacy risks by not explicitly sharing private data. However, recent works demonstrated that sharing model updates makes FL vulnerable to inference attacks. In this work, we show our key observation that the data representation leakage from gradients is the essential cause of privacy leakage in FL. We also provide an analysis of this observation to explain how the data presentation is leaked. Based on this observation, we propose a defense against model inversion attack in FL. The key idea of our defense is learning to perturb data representation such that the quality of the reconstructed data is severely degraded, while FL performance is maintained. In addition, we derive certified robustness guarantee to FL and convergence guarantee to FedAvg, after applying our defense. To evaluate our defense, we conduct experiments on MNIST and CIFAR10 for defending against the DLG attack and GS attack. Without sacrificing accuracy, the results demonstrate that our proposed defense can increase the mean squared error between the reconstructed data and the raw data by as much as more than 160X for both DLG attack and GS attack, compared with baseline defense methods. The privacy of the FL system is significantly improved.

cs.LG↗