Searcharxiv⌕ Search

arXiv subjects

Mei Yang

Publications and source records attributed to Mei Yang.

32 records · Page 2Linked to original sources

Physics-Guided Graph Neural Networks for Real-time AC/DC Power Flow Analysis

The increasing scale of alternating current and direct current (AC/DC) hybrid systems necessitates a faster power flow analysis tool than ever. This letter thus proposes a specific physics-guided graph neural network (PG-GNN). The tailored graph modelling of AC and DC grids is firstly advanced to enhance the topology adaptability of the PG-GNN. To eschew unreliable experience emulation from data, AC/DC physics are embedded in the PG-GNN using duality. Augmented Lagrangian method-based learning scheme is then presented to help the PG-GNN better learn nonconvex patterns in an unsupervised label-free manner. Multi-PG-GNN is finally conducted to master varied DC control modes. Case study shows that, relative to the other 7 data-driven rivals, only the proposed method matches the performance of the model-based benchmark, also beats it in computational efficiency beyond 10 times.

eess.SY↗

Interpreting Vulnerabilities of Multi-Instance Learning to Adversarial Perturbations

Multi-Instance Learning (MIL) is a recent machine learning paradigm which is immensely useful in various real-life applications, like image analysis, video anomaly detection, text classification, etc. It is well known that most of the existing machine learning classifiers are highly vulnerable to adversarial perturbations. Since MIL is a weakly supervised learning, where information is available for a set of instances, called bag and not for every instances, adversarial perturbations can be fatal. In this paper, we have proposed two adversarial perturbation methods to analyze the effect of adversarial perturbations to interpret the vulnerability of MIL methods. Out of the two algorithms, one can be customized for every bag, and the other is a universal one, which can affect all bags in a given data set and thus has some generalizability. Through simulations, we have also shown the effectiveness of the proposed algorithms to fool the state-of-the-art (SOTA) MIL methods. Finally, we have discussed through experiments, about taking care of these kind of adversarial perturbations through a simple strategy. Source codes are available at https://github.com/InkiInki/MI-UAP.

cs.CV↗

Fault-Aware Neural Code Rankers

Large language models (LLMs) have demonstrated an impressive ability to generate code for various programming tasks. In many instances, LLMs can generate a correct program for a task when given numerous trials. Consequently, a recent trend is to do large scale sampling of programs using a model and then filtering/ranking the programs based on the program execution on a small number of known unit tests to select one candidate solution. However, these approaches assume that the unit tests are given and assume the ability to safely execute the generated programs (which can do arbitrary dangerous operations such as file manipulations). Both of the above assumptions are impractical in real-world software development. In this paper, we propose CodeRanker, a neural ranker that can predict the correctness of a sampled program without executing it. Our CodeRanker is fault-aware i.e., it is trained to predict different kinds of execution information such as predicting the exact compile/runtime error type (e.g., an IndexError or a TypeError). We show that CodeRanker can significantly increase the pass@1 accuracy of various code generation models (including Codex, GPT-Neo, GPT-J) on APPS, HumanEval and MBPP datasets.

cs.PL↗

In-Network Accumulation: Extending the Role of NoC for DNN Acceleration

Network-on-Chip (NoC) plays a significant role in the performance of a DNN accelerator. The scalability and modular design property of the NoC help in improving the performance of a DNN execution by providing flexibility in running different kinds of workloads. Data movement in a DNN workload is still a challenging task for DNN accelerators and hence a novel approach is required. In this paper, we propose the In-Network Accumulation (INA) method to further accelerate a DNN workload execution on a many-core spatial DNN accelerator for the Weight Stationary (WS) dataflow model. The INA method expands the router's function to support partial sum accumulation. This method avoids the overhead of injecting and ejecting an incoming partial sum to the local processing element. The simulation results on AlexNet, ResNet-50, and VGG-16 workloads show that the proposed INA method achieves 1.22x improvement in latency and 2.16x improvement in power consumption on the WS dataflow model.

cs.AR↗

Real-Time Monocular Human Depth Estimation and Segmentation on Embedded Systems

Estimating a scene's depth to achieve collision avoidance against moving pedestrians is a crucial and fundamental problem in the robotic field. This paper proposes a novel, low complexity network architecture for fast and accurate human depth estimation and segmentation in indoor environments, aiming to applications for resource-constrained platforms (including battery-powered aerial, micro-aerial, and ground vehicles) with a monocular camera being the primary perception module. Following the encoder-decoder structure, the proposed framework consists of two branches, one for depth prediction and another for semantic segmentation. Moreover, network structure optimization is employed to improve its forward inference speed. Exhaustive experiments on three self-generated datasets prove our pipeline's capability to execute in real-time, achieving higher frame rates than contemporary state-of-the-art frameworks (114.6 frames per second on an NVIDIA Jetson Nano GPU with TensorRT) while maintaining comparable accuracy.

cs.CV↗

Efficient On-Chip Multicast Routing based on Dynamic Partition Merging

Networks-on-chips (NoCs) have become the mainstream communication infrastructure for chip multiprocessors (CMPs) and many-core systems. The commonly used parallel applications and emerging machine learning-based applications involve a significant amount of collective communication patterns. In CMP applications, multicast is widely used in multithreaded programs and protocols for barrier/clock synchronization and cache coherence. Multicast routing plays an important role on the system performance of a CMP. Existing partition-based multicast routing algorithms all use static destination set partition strategy which lacks the global view of path optimization. In this paper, we propose an efficient Dynamic Partition Merging (DPM)-based multicast routing algorithm. The proposed algorithm divides the multicast destination set into partitions dynamically by comparing the routing cost of different partition merging options and selecting the merged partitions with lower cost. The simulation results of synthetic traffic and PARSEC benchmark applications confirm that the proposed algorithm outperforms the existing path-based routing algorithms. The proposed algorithm is able to improve up to 23\% in average packet latency and 14\% in power consumption against the existing multipath routing algorithm when tested in PARSEC benchmark workloads.

cs.NI↗

Improving the Performance of a NoC-based CNN Accelerator with Gather Support

The increasing application of deep learning technology drives the need for an efficient parallel computing architecture for Convolutional Neural Networks (CNNs). A significant challenge faced when designing a many-core CNN accelerator is to handle the data movement between the processing elements. The CNN workload introduces many-to-one traffic in addition to one-to-one and one-to-many traffic. As the de-facto standard for on-chip communication, Network-on-Chip (NoC) can support various unicast and multicast traffic. For many-to-one traffic, repetitive unicast is employed which is not an efficient way. In this paper, we propose to use the gather packet on mesh-based NoCs employing output stationary systolic array in support of many-to-one traffic. The gather packet will collect the data from the intermediate nodes eventually leading to the destination efficiently. This method is evaluated using the traffic traces generated from the convolution layer of AlexNet and VGG-16 with improvement in the latency and power than the repetitive unicast method.

cs.LG↗

Data Streaming and Traffic Gathering in Mesh-based NoC for Deep Neural Network Acceleration

The increasing popularity of deep neural network (DNN) applications demands high computing power and efficient hardware accelerator architecture. DNN accelerators use a large number of processing elements (PEs) and on-chip memory for storing weights and other parameters. As the communication backbone of a DNN accelerator, networks-on-chip (NoC) play an important role in supporting various dataflow patterns and enabling processing with communication parallelism in a DNN accelerator. However, the widely used mesh-based NoC architectures inherently cannot support the efficient one-to-many and many-to-one traffic largely existing in DNN workloads. In this paper, we propose a modified mesh architecture with a one-way/two-way streaming bus to speedup one-to-many (multicast) traffic, and the use of gather packets to support many-to-one (gather) traffic. The analysis of the runtime latency of a convolutional layer shows that the two-way streaming architecture achieves better improvement than the one-way streaming architecture for an Output Stationary (OS) dataflow architecture. The simulation results demonstrate that the gather packets can help to reduce the runtime latency up to 1.8 times and network power consumption up to 1.7 times, compared with the repetitive unicast method on modified mesh architectures supporting two-way streaming.

cs.LG↗

First Solar energetic particles measured on the Lunar far-side

On 2019 May 6, the Lunar Lander Neutron & Dosimetry (LND) Experiment on board the Chang'E-4 on the far-side of the Moon detected its first small solar energetic particle (SEP) event with proton energies up to 21MeV. Combined proton energy spectra are studied based on the LND, SOHO/EPHIN and ACE/EPAM measurements which show that LND could provide a complementary dataset from a special location on the Moon, contributing to our existing observations and understanding of space environment. Velocity dispersion analysis (VDA) has been applied to the impulsive electron event and weak proton enhancement and the results demonstrate that electrons are released only 22 minutes after the flare onset and $\sim$15 minutes after type II radio burst, while protons are released more than one hour after the electron release. The impulsive enhancement of the in-situ electrons and the derived early release time indicate a good magnetic connection between the source and Earth. However, stereoscopic remote-sensing observations from Earth and STA suggest that the SEPs are associated with an active region nearly 100$^\circ$ away from the magnetic footpoint of Earth. This suggests that the propagation of these SEPs could not follow a nominal Parker spiral under the ballistic mapping model and the release and propagation mechanism of electrons and protons are likely to differ significantly during this event.

astro-ph.SR↗

Multi-Object Portion Tracking in 4D Fluorescence Microscopy Imagery with Deep Feature Maps

3D fluorescence microscopy of living organisms has increasingly become an essential and powerful tool in biomedical research and diagnosis. An exploding amount of imaging data has been collected, whereas efficient and effective computational tools to extract information from them are still lagging behind. This is largely due to the challenges in analyzing biological data. Interesting biological structures are not only small, but are often morphologically irregular and highly dynamic. Although tracking cells in live organisms has been studied for years, existing tracking methods for cells are not effective in tracking subcellular structures, such as protein complexes, which feature in continuous morphological changes including split and merge, in addition to fast migration and complex motion. In this paper, we first define the problem of multi-object portion tracking to model the protein object tracking process. A multi-object tracking method with portion matching is proposed based on 3D segmentation results. The proposed method distills deep feature maps from deep networks, then recognizes and matches object portions using an extended search. Experimental results confirm that the proposed method achieves 2.96% higher on consistent tracking accuracy and 35.48% higher on event identification accuracy than the state-of-art methods.

eess.IV↗

Mid-infrared polarized emission from black phosphorus light-emitting diodes

The mid-infrared (MIR) spectral range is of immense use for civilian and military applications. The large number of vibrational absorption bands in this range can be used for gas sensing, process control and spectroscopy. In addition, there exists transparency windows in the atmosphere such as that between 3.6-3.8 $μ$m, which are ideal for free-space optical communication, range finding and thermal imaging. A number of different semiconductor platforms have been used for MIR light-emission. This includes InAsSb/InAs quantum wells, InSb/AlInSb, GaInAsSbP pentanary alloys, and intersubband transitions in group III-V compounds. These approaches, however, are costly and lack the potential for integration on silicon and silicon-on-insulator platforms. In this respect, two-dimensional (2D) materials are particularly attractive due to the ease with which they can be heterointegrated. Weak interactions between neighbouring atomic layers in these materials allows for deposition on arbitrary substrates and van der Waals heterostructures enable the design of devices with targeted optoelectronic properties. In this Letter, we demonstrate a light-emitting diode (LED) based on the 2D semiconductor black phosphorus (BP). The device, which is composed of a BP/molybdenum disulfide (MoS$_2$) heterojunction emits polarized light at $λ$ = 3.68 $μ$m with room-temperature internal and external quantum efficiencies (IQE and EQE) of ~1$\%$ and $\sim3\times10^{-2}\%$, respectively. The ability to tune the bandgap, and consequently emission wavelength of BP, with layer number, strain and electric field make it a particularly attractive platform for MIR emission.

physics.app-ph↗

An improved design method for conventional straight dipole magnets

The standard design method for conventional straight dipole magnets is improved in this paper. The good field region is not symmetric with respect to the magnet mechanical center, and its width is not enlarged to include the beam sagitta. The integrated field quality is obtained by integrating the field along nominal beam paths. 2D and 3D design procedures of the improved design method are introduced, and two application examples of straight dipole magnets are presented. It is shown that the differences in integrated field quality between different field integration paths cannot be neglected. Compared with the traditional design method of straight dipole magnets, the advantage of the improved method is that the integrated field quality is accurate; the pole width, magnet dimension and weight of a straight dipole magnet can be reduced.

physics.acc-ph↗

Co-Emulation of Scan-Chain Based Designs Utilizing SCE-MI Infrastructure

As the complexity of the scan algorithm is dependent on the number of design registers, large SoC scan designs can no longer be verified in RTL simulation unless partitioned into smaller sub-blocks. This paper proposes a methodology to decrease scan-chain verification time utilizing SCE-MI, a widely used communication protocol for emulation, and an FPGA-based emulation platform. A high-level (SystemC) testbench and FPGA synthesizable hardware transactor models are developed for the scan-chain ISCAS89 S400 benchmark circuit for high-speed communication between the host CPU workstation and the FPGA emulator. The emulation results are compared to other verification methodologies (RTL Simulation, Simulation Acceleration, and Transaction-based emulation), and found to be 82% faster than regular RTL simulation. In addition, the emulation runs in the MHz speed range, allowing the incorporation of software applications, drivers, and operating systems, as opposed to the Hz range in RTL simulation or sub-megahertz range as accomplished in transaction-based emulation. In addition, the integration of scan testing and acceleration/emulation platforms allows more complex DFT methods to be developed and tested on a large scale system, decreasing the time to market for products.

cs.OH↗

Service Composition in Service-Oriented Wireless Sensor Networks with Persistent Queries

Service-oriented wireless sensor network(WSN) has been recently proposed as an architecture to rapidly develop applications in WSNs. In WSNs, a query task may require a set of services and may be carried out repetitively with a given frequency during its lifetime. A service composition solution shall be provided for each execution of such a persistent query task. Due to the energy saving strategy, some sensors may be scheduled to be in sleep mode periodically. Thus, a service composition solution may not always be valid during the lifetime of a persistent query. When a query task needs to be conducted over a new service composition solution, a routing update procedure is involved which consumes energy. In this paper, we study service composition design which minimizes the number of service composition solutions during the lifetime of a persistent query. We also aim to minimize the total service composition cost when the minimum number of required service composition solutions is derived. A greedy algorithm and a dynamic programming algorithm are proposed to complete these two objectives respectively. The optimality of both algorithms provides the service composition solutions for a persistent query with minimum energy consumption.

cs.NI↗