Searcharxiv⌕ Search

arXiv subjects

Ang Li

Publications and source records attributed to Ang Li.

At least 505 records · Page 28Linked to original sources

Accelerated Deep Reinforcement Learning Based Load Shedding for Emergency Voltage Control

Load shedding has been one of the most widely used and effective emergency control approaches against voltage instability. With increased uncertainties and rapidly changing operational conditions in power systems, existing methods have outstanding issues in terms of either speed, adaptiveness, or scalability. Deep reinforcement learning (DRL) was regarded and adopted as a promising approach for fast and adaptive grid stability control in recent years. However, existing DRL algorithms show two outstanding issues when being applied to power system control problems: 1) computational inefficiency that requires extensive training and tuning time; and 2) poor scalability making it difficult to scale to high dimensional control problems. To overcome these issues, an accelerated DRL algorithm named PARS was developed and tailored for power system voltage stability control via load shedding. PARS features high scalability and is easy to tune with only five main hyperparameters. The method was tested on both the IEEE 39-bus and IEEE 300-bus systems, and the latter is by far the largest scale for such a study. Test results show that, compared to other methods including model-predictive control (MPC) and proximal policy optimization(PPO) methods, PARS shows better computational efficiency (faster convergence), more robustness in learning, excellent scalability and generalization capability.

eess.SY↗

Towards Latency-aware DNN Optimization with GPU Runtime Analysis and Tail Effect Elimination

Despite the superb performance of State-Of-The-Art (SOTA) DNNs, the increasing computational cost makes them very challenging to meet real-time latency and accuracy requirements. Although DNN runtime latency is dictated by model property (e.g., architecture, operations), hardware property (e.g., utilization, throughput), and more importantly, the effective mapping between these two, many existing approaches focus only on optimizing model property such as FLOPS reduction and overlook the mismatch between DNN model and hardware properties. In this work, we show that the mismatch between the varied DNN computation workloads and GPU capacity can cause the idle GPU tail effect, leading to GPU under-utilization and low throughput. As a result, the FLOPs reduction cannot bring effective latency reduction, which causes sub-optimal accuracy versus latency trade-offs. Motivated by this, we propose a GPU runtime-aware DNN optimization methodology to eliminate such GPU tail effect adaptively on GPU platforms. Our methodology can be applied on top of existing SOTA DNN optimization approaches to achieve better latency and accuracy trade-offs. Experiments show 11%-27% latency reduction and 2.5%-4.0% accuracy improvement over several SOTA DNN pruning and NAS methods, respectively

cs.AR↗

On-chip spectrometer using stratified waveguides filters

We present an ultra-compact single-shot spectrometer on silicon platform with broad operation bandwidth and high resolution. It consists of 32 stratified waveguide filters (SWFs) with diverse transmission spectra for sampling the unknown spectrum of the input signal and a specially designed ultra-compact structure for splitting the incident signal into 32 filters with low imbalance. Each SWF has a footprint less than 1um x 30um, while the 1x32 splitter and 32 filters in total occupy an area of about 35um x 260um, which to the best of our knowledge, is the smallest footprint spectrometer realized on silicon photonic platform. Experimental characteristics of the fabricated spectrometer demonstrate a broad operating bandwidth of 180nm centered at 1550nm and narrowband peaks with 0.45nm Full-Width-Half-Maximum (FWHM) can be clearly resolved. This concept can also be implemented using other material platforms for operation in optical spectral bands of interest for various applications.

physics.optics↗

An $\mathbb{R}$-motivic $v_{1}-$self-map of periodicity $1$

We consider a nontrivial action of $\mathrm{C}_2$ on the type $1$ spectrum $\mathcal{Y} := \mathrm{M}_2(1) \wedge \mathrm{C}(η)$, which is well-known for admitting a $1$-periodic $v_1-$self-map. The resultant finite $\mathrm{C}_2$-equivariant spectrum $\mathcal{Y}^{\mathrm{C}_2}$ can also be viewed as the complex points of a finite $\mathbb{R}$-motivic spectrum $\mathcal{Y}^\mathbb{R}$. In this paper, we show that one of the $1$-periodic $v_1-$self-maps of $\mathcal{Y}$ can be lifted to a self-map of $\mathcal{Y}^{\mathrm{C}_2}$ as well as $\mathcal{Y}^{\mathbb{R}}$. Further, the cofiber of the self-map of $\mathcal{Y}^{\mathbb{R}}$ is a realization of the subalgebra $\mathcal{A}^\mathbb{R}(1)$ of the $\mathbb{R}$-motivic Steenrod algebra. We also show that the $\mathrm{C}_2$-equivariant self-map is nilpotent on the geometric fixed-points of $\mathcal{Y}^{\mathrm{C}_2}$.

math.AT↗

Tidal deformability and gravitational-wave phase evolution of magnetised compact-star binaries

The evolution of the gravitational-wave phase in the signal produced by inspiralling binaries of compact stars is modified by the nonzero deformability of the two stars. Hence, the measurement of these corrections has the potential of providing important information on the equation of state of nuclear matter. Extensive work has been carried out over the last decade to quantify these corrections, but it has so far been restricted to stars with zero intrinsic magnetic fields. While the corrections introduced by the magnetic tension and magnetic pressure are expected to be subdominant, it is nevertheless useful to determine the precise conditions under which these corrections become important. To address this question, we have carried out a second-order perturbative analysis of the tidal deformability of magnetised compact stars under a variety of magnetic-field strengths and equations of state describing either neutron stars or quark stars. Overall, we find that magnetically induced corrections to the tidal deformability will produce changes in the gravitational-wave phase evolution that are unlikely to be detected for realistic magnetic field i.e., $B\sim 10^{10} - 10^{12}\,{\rm G}$. At the same time, if the magnetic field is unrealistically large, i.e., $B\sim 10^{16}\,{\rm G}$, these corrections would produce a sizeable contribution to the phase evolution, especially for quark stars. In the latter case, the induced phase differences would represent a unique tool to measure the properties of the magnetic fields, providing information that is otherwise hard to quantify.

astro-ph.HE↗

Learning to Incentivize Other Learning Agents

The challenge of developing powerful and general Reinforcement Learning (RL) agents has received increasing attention in recent years. Much of this effort has focused on the single-agent setting, in which an agent maximizes a predefined extrinsic reward function. However, a long-term question inevitably arises: how will such independent agents cooperate when they are continually learning and acting in a shared multi-agent environment? Observing that humans often provide incentives to influence others' behavior, we propose to equip each RL agent in a multi-agent environment with the ability to give rewards directly to other agents, using a learned incentive function. Each agent learns its own incentive function by explicitly accounting for its impact on the learning of recipients and, through them, the impact on its own extrinsic objective. We demonstrate in experiments that such agents significantly outperform standard RL and opponent-shaping agents in challenging general-sum Markov games, often by finding a near-optimal division of labor. Our work points toward more opportunities and challenges along the path to ensure the common good in a multi-agent future.

cs.LG↗

The Effectiveness of Memory Replay in Large Scale Continual Learning

We study continual learning in the large scale setting where tasks in the input sequence are not limited to classification, and the outputs can be of high dimension. Among multiple state-of-the-art methods, we found vanilla experience replay (ER) still very competitive in terms of both performance and scalability, despite its simplicity. However, a degraded performance is observed for ER with small memory. A further visualization of the feature space reveals that the intermediate representation undergoes a distributional drift. While existing methods usually replay only the input-output pairs, we hypothesize that their regularization effect is inadequate for complex deep models and diverse tasks with small replay buffer size. Following this observation, we propose to replay the activation of the intermediate layers in addition to the input-output pairs. Considering that saving raw activation maps can dramatically increase memory and compute cost, we propose the Compressed Activation Replay technique, where compressed representations of layer activation are saved to the replay buffer. We show that this approach can achieve superior regularization effect while adding negligible memory overhead to replay method. Experiments on both the large-scale Taskonomy benchmark with a diverse set of tasks and standard common datasets (Split-CIFAR and Split-miniImageNet) demonstrate the effectiveness of the proposed method.

cs.LG↗

Constraining hadron-quark phase transition parameters within the quark-mean-field model using multimessenger observations of neutron stars

We extend the quark mean-field (QMF) model for nuclear matter and study the possible presence of quark matter inside the cores of neutron stars. A sharp first-order hadron-quark phase transition is implemented combining the QMF for the hadronic phase with "constant-speed-of-sound" parametrization for the high-density quark phase. The interplay of the nuclear symmetry energy slope parameter, $L$, and the dimensionless phase transition parameters (the transition density $n_{\rm trans}/n_0$, the transition strength $Δ\varepsilon/\varepsilon_{\rm trans}$, and the sound speed squared in quark matter $c^2_{\rm QM}$) are then systematically explored for the hybrid star proprieties, especially the maximum mass $M_{\rm max}$ and the radius and the tidal deformability of a typical $1.4 \,M_{\odot}$ star. We show the strong correlation between the symmetry energy slope $L$ and the typical stellar radius $R_{1.4}$, similar to that previously found for neutron stars without a phase transition. With the inclusion of phase transition, we obtain robust limits on the maximum mass ($M_{\rm max}< 3.6 \,M_{\odot}$) and the radius of $1.4 \,M_{\odot}$ stars ($R_{1.4}\gtrsim 9.6~\rm km$), and we find that a too-weak ($Δ\varepsilon/\varepsilon_{\rm trans}\lesssim 0.2$) phase transition taking place at low densities $\lesssim 1.3-1.5 \, n_0$ is strongly disfavored. We also demonstrate that future measurements of the radius and tidal deformability of $\sim 1.4 \,M_{\odot}$ stars, as well as the mass measurement of very massive pulsars, can help reveal the presence and amount of quark matter in compact objects.

nucl-th↗

Benchmarking Machine Learning Techniques with Di-Higgs Production at the LHC

Many domains of high energy physics analysis are starting to explore machine learning techniques. Powerful methods can be used to identify and measure rare processes from previously insurmountable backgrounds. One of the most profound Standard Model signatures still to be discovered at the LHC is the pair production of Higgs bosons through the Higgs self-coupling. The small cross section of this process makes detection very difficult even for the decay channel with the largest branching fraction ($hh\rightarrow b\bar{b}b\bar{b}$). This paper benchmarks a variety of approaches (boosted decision trees, various neural network architectures, semi-supervised algorithms) against one another to catalog a few of the various techniques available to high energy physicists as the era of the HL-LHC approaches.

hep-ph↗

Generative Image Inpainting with Submanifold Alignment

Image inpainting aims at restoring missing regions of corrupted images, which has many applications such as image restoration and object removal. However, current GAN-based generative inpainting models do not explicitly exploit the structural or textural consistency between restored contents and their surrounding contexts.To address this limitation, we propose to enforce the alignment (or closeness) between the local data submanifolds (or subspaces) around restored images and those around the original (uncorrupted) images during the learning process of GAN-based inpainting models. We exploit Local Intrinsic Dimensionality (LID) to measure, in deep feature space, the alignment between data submanifolds learned by a GAN model and those of the original data, from a perspective of both images (denoted as iLID) and local patches (denoted as pLID) of images. We then apply iLID and pLID as regularizations for GAN-based inpainting models to encourage two levels of submanifold alignment: 1) an image-level alignment for improving structural consistency, and 2) a patch-level alignment for improving textural details. Experimental results on four benchmark datasets show that our proposed model can generate more accurate results than state-of-the-art models.

cs.CV↗

Reinforcement Learning-based Black-Box Evasion Attacks to Link Prediction in Dynamic Graphs

Link prediction in dynamic graphs (LPDG) is an important research problem that has diverse applications such as online recommendations, studies on disease contagion, organizational studies, etc. Various LPDG methods based on graph embedding and graph neural networks have been recently proposed and achieved state-of-the-art performance. In this paper, we study the vulnerability of LPDG methods and propose the first practical black-box evasion attack. Specifically, given a trained LPDG model, our attack aims to perturb the graph structure, without knowing to model parameters, model architecture, etc., such that the LPDG model makes as many wrong predicted links as possible. We design our attack based on a stochastic policy-based RL algorithm. Moreover, we evaluate our attack on three real-world graph datasets from different application domains. Experimental results show that our attack is both effective and efficient.

cs.CR↗

Short-Term and Long-Term Context Aggregation Network for Video Inpainting

Video inpainting aims to restore missing regions of a video and has many applications such as video editing and object removal. However, existing methods either suffer from inaccurate short-term context aggregation or rarely explore long-term frame information. In this work, we present a novel context aggregation network to effectively exploit both short-term and long-term frame information for video inpainting. In the encoding stage, we propose boundary-aware short-term context aggregation, which aligns and aggregates, from neighbor frames, local regions that are closely related to the boundary context of missing regions into the target frame. Furthermore, we propose dynamic long-term context aggregation to globally refine the feature map generated in the encoding stage using long-term frame features, which are dynamically updated throughout the inpainting process. Experiments show that it outperforms state-of-the-art methods with better inpainting results and fast inpainting speed.

cs.CV↗

AWB-GCN: A Graph Convolutional Network Accelerator with Runtime Workload Rebalancing

Deep learning systems have been successfully applied to Euclidean data such as images, video, and audio. In many applications, however, information and their relationships are better expressed with graphs. Graph Convolutional Networks (GCNs) appear to be a promising approach to efficiently learn from graph data structures, having shown advantages in many critical applications. As with other deep learning modalities, hardware acceleration is critical. The challenge is that real-world graphs are often extremely large and unbalanced; this poses significant performance demands and design challenges. In this paper, we propose Autotuning-Workload-Balancing GCN (AWB-GCN) to accelerate GCN inference. To address the issue of workload imbalance in processing real-world graphs, three hardware-based autotuning techniques are proposed: dynamic distribution smoothing, remote switching, and row remapping. In particular, AWB-GCN continuously monitors the sparse graph pattern, dynamically adjusts the workload distribution among a large number of processing elements (up to 4K PEs), and, after converging, reuses the ideal configuration. Evaluation is performed using an Intel D5005 FPGA with five commonly-used datasets. Results show that 4K-PE AWB-GCN can significantly elevate PE utilization by 7.7x on average and demonstrate considerable performance speedups over CPUs (3255x), GPUs (80.3x), and a prior GCN accelerator (5.1x).

cs.DC↗

TIPRDC: Task-Independent Privacy-Respecting Data Crowdsourcing Framework for Deep Learning with Anonymized Intermediate Representations

The success of deep learning partially benefits from the availability of various large-scale datasets. These datasets are often crowdsourced from individual users and contain private information like gender, age, etc. The emerging privacy concerns from users on data sharing hinder the generation or use of crowdsourcing datasets and lead to hunger of training data for new deep learning applications. One na\"ıve solution is to pre-process the raw data to extract features at the user-side, and then only the extracted features will be sent to the data collector. Unfortunately, attackers can still exploit these extracted features to train an adversary classifier to infer private attributes. Some prior arts leveraged game theory to protect private attributes. However, these defenses are designed for known primary learning tasks, the extracted features work poorly for unknown learning tasks. To tackle the case where the learning task may be unknown or changing, we present TIPRDC, a task-independent privacy-respecting data crowdsourcing framework with anonymized intermediate representation. The goal of this framework is to learn a feature extractor that can hide the privacy information from the intermediate representations; while maximally retaining the original information embedded in the raw data for the data collector to accomplish unknown learning tasks. We design a hybrid training method to learn the anonymized intermediate representation: (1) an adversarial training process for hiding private information from features; (2) maximally retain original information using a neural-network-based mutual information estimator.

cs.LG↗

Orbital selectivity of layer resolved tunneling on iron superconductor Ba0.6K0.4Fe2As2

We use scanning tunneling microscopy/spectroscopy (STM/S) to elucidate the Cooper pairing of the iron pnictide superconductor Ba0.6K0.4Fe2As2. By a cold-cleaving technique, we obtain atomically resolved termination surfaces with different layer identities. Remarkably, we observe that the low-energy tunneling spectrum related to superconductivity has an unprecedented dependence on the layer-identity. By cross-referencing with the angle-revolved photoemission results and the tunneling data of LiFeAs, we find that tunneling on each termination surface probes superconductivity through selecting distinct Fe-3d orbitals. These findings imply the real-space orbital features of the Cooper pairing in the iron pnictide superconductors, and propose a new and general concept that, for complex multi-orbital material, tunneling on different terminating layers can feature orbital selectivity.

cond-mat.supr-con↗

The $v_1$-Periodic Region of the Complex-Motivic Ext

We establish a $v_1$-periodicity theorem in Ext over the complex-motivic Steenrod algebra. The element $h_1$ of Ext, which detects the homotopy class $η$ in the motivic Adams spectral sequence, is non-nilpotent and therefore generates $h_1$-towers. Our result is that, apart from these $h_1$-towers, $v_1$-periodicity behaves as it does classically.

math.AT↗

LotteryFL: Personalized and Communication-Efficient Federated Learning with Lottery Ticket Hypothesis on Non-IID Datasets

Federated learning is a popular distributed machine learning paradigm with enhanced privacy. Its primary goal is learning a global model that offers good performance for the participants as many as possible. The technology is rapidly advancing with many unsolved challenges, among which statistical heterogeneity (i.e., non-IID) and communication efficiency are two critical ones that hinder the development of federated learning. In this work, we propose LotteryFL -- a personalized and communication-efficient federated learning framework via exploiting the Lottery Ticket hypothesis. In LotteryFL, each client learns a lottery ticket network (i.e., a subnetwork of the base model) by applying the Lottery Ticket hypothesis, and only these lottery networks will be communicated between the server and clients. Rather than learning a shared global model in classic federated learning, each client learns a personalized model via LotteryFL; the communication cost can be significantly reduced due to the compact size of lottery networks. To support the training and evaluation of our framework, we construct non-IID datasets based on MNIST, CIFAR-10 and EMNIST by taking feature distribution skew, label distribution skew and quantity skew into consideration. Experiments on these non-IID datasets demonstrate that LotteryFL significantly outperforms existing solutions in terms of personalization and communication cost.

cs.LG↗

Hybrid Models for Open Set Recognition

Open set recognition requires a classifier to detect samples not belonging to any of the classes in its training set. Existing methods fit a probability distribution to the training samples on their embedding space and detect outliers according to this distribution. The embedding space is often obtained from a discriminative classifier. However, such discriminative representation focuses only on known classes, which may not be critical for distinguishing the unknown classes. We argue that the representation space should be jointly learned from the inlier classifier and the density estimator (served as an outlier detector). We propose the OpenHybrid framework, which is composed of an encoder to encode the input data into a joint embedding space, a classifier to classify samples to inlier classes, and a flow-based density estimator to detect whether a sample belongs to the unknown category. A typical problem of existing flow-based models is that they may assign a higher likelihood to outliers. However, we empirically observe that such an issue does not occur in our experiments when learning a joint representation for discriminative and generative components. Experiments on standard open set benchmarks also reveal that an end-to-end trained OpenHybrid model significantly outperforms state-of-the-art methods and flow-based baselines.

cs.CV↗