SearcharxivSearch

arXiv subjects

Peng Lin

Publications and source records attributed to Peng Lin.

At least 37 records · Page 2Linked to original sources

Roadmap for Unconventional Computing with Nanotechnology

In the "Beyond Moore's Law" era, with increasing edge intelligence, domain-specific computing embracing unconventional approaches will become increasingly prevalent. At the same time, adopting a variety of nanotechnologies will offer benefits in energy cost, computational speed, reduced footprint, cyber resilience, and processing power. The time is ripe for a roadmap for unconventional computing with nanotechnologies to guide future research, and this collection aims to fill that need. The authors provide a comprehensive roadmap for neuromorphic computing using electron spins, memristive devices, two-dimensional nanomaterials, nanomagnets, and various dynamical systems. They also address other paradigms such as Ising machines, Bayesian inference engines, probabilistic computing with p-bits, processing in memory, quantum memories and algorithms, computing with skyrmions and spin waves, and brain-inspired computing for incremental learning and problem-solving in severely resource-constrained environments. These approaches have advantages over traditional Boolean computing based on von Neumann architecture. As the computational requirements for artificial intelligence grow 50 times faster than Moore's Law for electronics, more unconventional approaches to computing and signal processing will appear on the horizon, and this roadmap will help identify future needs and challenges. In a very fertile field, experts in the field aim to present some of the dominant and most promising technologies for unconventional computing that will be around for some time to come. Within a holistic approach, the goal is to provide pathways for solidifying the field and guiding future impactful discoveries.

cs.ET

Stacking Factorizing Partitioned Expressions in Hybrid Bayesian Network Models

Hybrid Bayesian networks (HBN) contain complex conditional probabilistic distributions (CPD) specified as partitioned expressions over discrete and continuous variables. The size of these CPDs grows exponentially with the number of parent nodes when using discrete inference, resulting in significant inefficiency. Normally, an effective way to reduce the CPD size is to use a binary factorization (BF) algorithm to decompose the statistical or arithmetic functions in the CPD by factorizing the number of connected parent nodes to sets of size two. However, the BF algorithm was not designed to handle partitioned expressions. Hence, we propose a new algorithm called stacking factorization (SF) to decompose the partitioned expressions. The SF algorithm creates intermediate nodes to incrementally reconstruct the densities in the original partitioned expression, allowing no more than two continuous parent nodes to be connected to each child node in the resulting HBN. SF can be either used independently or combined with the BF algorithm. We show that the SF+BF algorithm significantly reduces the CPD size and contributes to lowering the tree-width of a model, thus improving efficiency.

cs.AI

Darwin3: A large-scale neuromorphic chip with a Novel ISA and On-Chip Learning

Spiking Neural Networks (SNNs) are gaining increasing attention for their biological plausibility and potential for improved computational efficiency. To match the high spatial-temporal dynamics in SNNs, neuromorphic chips are highly desired to execute SNNs in hardware-based neuron and synapse circuits directly. This paper presents a large-scale neuromorphic chip named Darwin3 with a novel instruction set architecture(ISA), which comprises 10 primary instructions and a few extended instructions. It supports flexible neuron model programming and local learning rule designs. The Darwin3 chip architecture is designed in a mesh of computing nodes with an innovative routing algorithm. We used a compression mechanism to represent synaptic connections, significantly reducing memory usage. The Darwin3 chip supports up to 2.35 million neurons, making it the largest of its kind in neuron scale. The experimental results showed that code density was improved up to 28.3x in Darwin3, and neuron core fan-in and fan-out were improved up to 4096x and 3072x by connection compression compared to the physical memory depth. Our Darwin3 chip also provided memory saving between 6.8X and 200.8X when mapping convolutional spiking neural networks (CSNN) onto the chip, demonstrating state-of-the-art performance in accuracy and latency compared to other neuromorphic chips.

cs.NE

Random resistive memory-based deep extreme point learning machine for unified visual processing

Visual sensors, including 3D LiDAR, neuromorphic DVS sensors, and conventional frame cameras, are increasingly integrated into edge-side intelligent machines. Realizing intensive multi-sensory data analysis directly on edge intelligent machines is crucial for numerous emerging edge applications, such as augmented and virtual reality and unmanned aerial vehicles, which necessitates unified data representation, unprecedented hardware energy efficiency and rapid model training. However, multi-sensory data are intrinsically heterogeneous, causing significant complexity in the system development for edge-side intelligent machines. In addition, the performance of conventional digital hardware is limited by the physically separated processing and memory units, known as the von Neumann bottleneck, and the physical limit of transistor scaling, which contributes to the slowdown of Moore's law. These limitations are further intensified by the tedious training of models with ever-increasing sizes. We propose a novel hardware-software co-design, random resistive memory-based deep extreme point learning machine (DEPLM), that offers efficient unified point set analysis. We show the system's versatility across various data modalities and two different learning tasks. Compared to a conventional digital hardware-based system, our co-design system achieves huge energy efficiency improvements and training cost reduction when compared to conventional systems. Our random resistive memory-based deep extreme point learning machine may pave the way for energy-efficient and training-friendly edge AI across various data modalities and tasks.

cs.LG

Pruning random resistive memory for optimizing analogue AI

The rapid advancement of artificial intelligence (AI) has been marked by the large language models exhibiting human-like intelligence. However, these models also present unprecedented challenges to energy consumption and environmental sustainability. One promising solution is to revisit analogue computing, a technique that predates digital computing and exploits emerging analogue electronic devices, such as resistive memory, which features in-memory computing, high scalability, and nonvolatility. However, analogue computing still faces the same challenges as before: programming nonidealities and expensive programming due to the underlying devices physics. Here, we report a universal solution, software-hardware co-design using structural plasticity-inspired edge pruning to optimize the topology of a randomly weighted analogue resistive memory neural network. Software-wise, the topology of a randomly weighted neural network is optimized by pruning connections rather than precisely tuning resistive memory weights. Hardware-wise, we reveal the physical origin of the programming stochasticity using transmission electron microscopy, which is leveraged for large-scale and low-cost implementation of an overparameterized random neural network containing high-performance sub-networks. We implemented the co-design on a 40nm 256K resistive memory macro, observing 17.3% and 19.9% accuracy improvements in image and audio classification on FashionMNIST and Spoken digits datasets, as well as 9.8% (2%) improvement in PR (ROC) in image segmentation on DRIVE datasets, respectively. This is accompanied by 82.1%, 51.2%, and 99.8% improvement in energy efficiency thanks to analogue in-memory computing. By embracing the intrinsic stochasticity and in-memory computing, this work may solve the biggest obstacle of analogue computing systems and thus unleash their immense potential for next-generation AI hardware.

cs.ET

FreeEagle: Detecting Complex Neural Trojans in Data-Free Cases

Trojan attack on deep neural networks, also known as backdoor attack, is a typical threat to artificial intelligence. A trojaned neural network behaves normally with clean inputs. However, if the input contains a particular trigger, the trojaned model will have attacker-chosen abnormal behavior. Although many backdoor detection methods exist, most of them assume that the defender has access to a set of clean validation samples or samples with the trigger, which may not hold in some crucial real-world cases, e.g., the case where the defender is the maintainer of model-sharing platforms. Thus, in this paper, we propose FreeEagle, the first data-free backdoor detection method that can effectively detect complex backdoor attacks on deep neural networks, without relying on the access to any clean samples or samples with the trigger. The evaluation results on diverse datasets and model architectures show that FreeEagle is effective against various complex backdoor attacks, even outperforming some state-of-the-art non-data-free backdoor detection methods.

cs.CR

Automatic Scene-based Topic Channel Construction System for E-Commerce

Scene marketing that well demonstrates user interests within a certain scenario has proved effective for offline shopping. To conduct scene marketing for e-commerce platforms, this work presents a novel product form, scene-based topic channel which typically consists of a list of diverse products belonging to the same usage scenario and a topic title that describes the scenario with marketing words. As manual construction of channels is time-consuming due to billions of products as well as dynamic and diverse customers' interests, it is necessary to leverage AI techniques to automatically construct channels for certain usage scenarios and even discover novel topics. To be specific, we first frame the channel construction task as a two-step problem, i.e., scene-based topic generation and product clustering, and propose an E-commerce Scene-based Topic Channel construction system (i.e., ESTC) to achieve automated production, consisting of scene-based topic generation model for the e-commerce domain, product clustering on the basis of topic similarity, as well as quality control based on automatic model filtering and human screening. Extensive offline experiments and online A/B test validates the effectiveness of such a novel product form as well as the proposed system. In addition, we also introduce the experience of deploying the proposed system on a real-world e-commerce recommendation platform.

cs.CL

Echo state graph neural networks with analogue random resistor arrays

Recent years have witnessed an unprecedented surge of interest, from social networks to drug discovery, in learning representations of graph-structured data. However, graph neural networks, the machine learning models for handling graph-structured data, face significant challenges when running on conventional digital hardware, including von Neumann bottleneck incurred by physically separated memory and processing units, slowdown of Moore's law due to transistor scaling limit, and expensive training cost. Here we present a novel hardware-software co-design, the random resistor array-based echo state graph neural network, which addresses these challenges. The random resistor arrays not only harness low-cost, nanoscale and stackable resistors for highly efficient in-memory computing using simple physical laws, but also leverage the intrinsic stochasticity of dielectric breakdown to implement random projections in hardware for an echo state network that effectively minimizes the training cost thanks to its fixed and random weights. The system demonstrates state-of-the-art performance on both graph classification using the MUTAG and COLLAB datasets and node classification using the CORA dataset, achieving 34.2x, 93.2x, and 570.4x improvement of energy efficiency and 98.27%, 99.46%, and 95.12% reduction of training cost compared to conventional graph learning on digital hardware, respectively, which may pave the way for the next generation AI system for graph learning.

cs.ET

Multi-window SRS imaging using a rapid widely tunable fiber laser

Spectroscopic stimulated Raman scattering (SRS) imaging has become a useful tool finding a broad range of applications. Yet, wider adoption is hindered by the bulky and environmentally-sensitive solid-state optical parametric oscillator (OPO) in current SRS microscope. Moreover, chemically-informative multi-window SRS imaging across C-H, C-D and fingerprint Raman regions is challenging due to the slow wavelength tuning speed of the solid-state OPO. In this work, we present a multi-window SRS imaging system based on a compact and robust fiber laser with rapid and widely tuning capability. To address the relative intensity noise intrinsic to fiber laser, we implemented auto-balanced detection which enhances the signal-to-noise ratio of stimulated Raman loss imaging by 23 times. We demonstrate high-quality SRS metabolic imaging of fungi, cancer cells, and Caenorhabditis elegans across the C-H, C-D and fingerprint Raman windows. Our re-sults showcase the potential of the compact multi-window SRS system for a broad range of applications.

physics.optics

On the computational solution of vector-density based continuum dislocation dynamics models: a comparison of two plastic distortion and stress update algorithms

Continuum dislocation dynamics models of mesoscale plasticity consist of dislocation transport-reaction equations coupled with crystal mechanics equations. The coupling between these two sets of equations is such that dislocation transport gives rise to the evolution of plastic distortion (strain), while the evolution of the latter fixes the stress from which the dislocation velocity field is found via a mobility law. Earlier solutions of these equations employed a staggered solution scheme for the two sets of equations in which the plastic distortion was updated via time integration of its rate, as found from Orowan's law. In this work, we show that such a direct time integration scheme can suffer from accumulation of numerical errors. We introduce an alternative scheme based on field dislocation mechanics that ensures consistency between the plastic distortion and the dislocation content in the crystal. The new scheme is based on calculating the compatible and incompatible parts of the plastic distortion separately, and the incompatible part is calculated from the current dislocation density field. Stress field and dislocation transport calculations were implemented within a finite element based discretization of the governing equations, with the crystal mechanics part solved by a conventional Galerkin method and the dislocation transport equations by the least squares method. A simple test is first performed to show the accuracy of the two schemes for updating the plastic distortion, which shows that the solution method based on field dislocation mechanics is more accurate. This method then was used to simulate an austenitic steel crystal under uniaxial loading and multiple slip conditions.

cond-mat.mtrl-sci

Incorporating point defect generation due to jog formation into the vector density-based continuum dislocation dynamics approach

During plastic deformation of crystalline materials, point defects such as vacancies and interstitials are generated by jogs on moving dislocations. A detailed model for jog formation and transport during plastic deformation was developed within the vector density-based continuum dislocation dynamics framework (Lin and El-Azab, 2020; Xia and El-Azab, 2015). As a part of this model, point defect generation associated with jog transport was formulated in terms of the volume change due to the non-conservative motion of jogs. Balance equations for the vacancies and interstitials including their rate of generation due to jog transport were also formulated. A two-way coupling between point defects and dislocation dynamics was then completed by including the stress contributed by the eigen-strain of point defects. A jog drag stress was further introduced into the mobility law of dislocations to account for the energy dissipation during point defects generation. A number of test problems and a fully coupled simulation of dislocation dynamics and point defect generation and diffusion were performed. The results show that there is an asymmetry of vacancy and interstitial generation due to the different formation energies of the two types of defects. The results also show that a higher hardening rate and a higher dislocation density are obtained when the point defect generation mechanism is coupled to dislocation dynamics.

cond-mat.mtrl-sci

Implementation of annihilation and junction reactions in vector density-based continuum dislocation dynamics

In a continuum dislocation dynamics formulation by Xia and El-Azab, dislocations are represented by a set of vector density fields, one per crystallographic slip systems. The space-time evolution of these densities is obtained by solving a set of dislocation transport equations coupled with crystal mechanics. Here, we present an approach for incorporating dislocation annihilation and junction reactions into the dislocation transport equations. These reactions consume dislocations and result in nothing as in the annihilation reactions, or produce new dislocations of different types as in the case of junction reactions. Collinear annihilation, glissile junctions, and sessile junctions are particularly emphasized here. A generalized energy-based criterion for junction reactions is established in terms of the dislocation density and Burgers vectors of the reacting species, and the reaction rate terms for junction reactions are formulated in terms of the dislocation densities. In order to illustrate how the dislocation network changes as a result of junction formation and annihilation in a continuum dislocation dynamics setting, we present some numerical examples focusing on the reactions processes themselves. The results show that our modeling approach is able to capture the respective dislocation network changes associated with dislocation reactions in FCC crystals: dislocations of opposite line directions encountering each other on collinear slip systems annihilate to connect the dislocations on the two slip systems, glissile junctions form on new slip system behave like Frank-Read sources, and sessile junctions form and expand along the intersection of the slip planes of the reacting dislocation species. A collective-dynamics test showing the frequency of occurrence of junctions of different types relative to each other is also presented.

physics.comp-ph

On the implementation of dislocation reactions in continuum dislocation dynamics modeling of mesoscale plasticity

The continuum dislocation dynamics framework for mesoscale plasticity is intended to capture the dislocation density evolution and the deformation of crystals when subjected to mechanical loading. It does so by solving a set of transport equations for dislocations concurrently with crystal mechanics equations, with the latter being cast in the form of an eigenstrain problem. Incorporating dislocation reactions in the dislocation transport equations is essential for making such continuum dislocation dynamics predictive. A formulation is proposed to incorporate dislocation reactions in the transport equations of the vector density-based continuum dislocation dynamics. This formulation aims to rigorously enforce dislocation line continuity using the concept of virtual dislocations that close all dislocation loops involved in cross slip, annihilation, and glissile and sessile junction reactions. The addition of virtual dislocations enables us to accurately enforce the divergence free condition upon the numerical solution of the dislocation transport equations for all slip systems individually. A set of tests were performed to illustrate the accuracy of the formulation and the solution of the transport equations within the vector density-based continuum dislocation dynamics. Comparing the results from these tests with an earlier approach in which the divergence free constraint was enforced on the total dislocation density tensor or the sum of two densities when only cross slip is considered shows that the new approach yields highly accurate results. Bulk simulations were performed for a face centered cubic crystal based on the new formulation and the results were compared with discrete dislocation dynamics predictions of the same. The microstructural features obtained from continuum dislocation dynamics were also analyzed with reference to relevant experimental observations.

cond-mat.mtrl-sci

Distributed velocity-constrained consensus of discrete-time multi-agent systems with nonconvex constraints, switching topologies, and delays

In this paper, a distributed velocity-constrained consensus problem is studied for discrete-time multi-agent systems, where each agent's velocity is constrained to lie in a nonconvex set. A distributed constrained control algorithm is proposed to enable all agents to converge to a common point using only local information. {The gains of the algorithm for all agents need not to be the same or predesigned and can be adjusted by each agent itself based on its own and neighbors' information.} It is shown that the algorithm is robust to arbitrarily bounded communication delays and arbitrarily switching communication graphs provided that the union of the graphs has directed spanning trees among each certain time interval. The analysis approach is based on multiple novel model transformations, proper control parameter selections, boundedness analysis of state-dependent stochastic matrices, exploitation of the convexity of stochastic matrices, and the joint connectivity of the communication graphs. Numerical examples are included to illustrate the theoretical results.

math.OC

Distributed Continuous-Time and Discrete-Time Optimization With Nonuniform Unbounded Convex Constraint Sets and Nonuniform Stepsizes

This paper is devoted to distributed continuous-time and discrete-time optimization problems with nonuniform convex constraint sets and nonuniform stepsizes for general differentiable convex objective functions. The communication graphs are not required to be strongly connected at any time, the gradients of the local objective functions are not required to be bounded when their independent variables tend to infinity, and the constraint sets are not required to be bounded. For continuous-time multi-agent systems, a distributed continuous algorithm is first introduced where the stepsizes and the convex constraint sets are both nonuniform. It is shown that all agents reach a consensus while minimizing the team objective function even when the constraint sets are unbounded. After that, the obtained results are extended to discrete-time multi-agent systems and then the case where each agent remains in a corresponding convex constraint set is studied. To ensure all agents to remain in a bounded region, a switching mechanism is introduced in the algorithms. It is shown that the distributed optimization problems can be solved, even though the discretization of the algorithms might deviate the convergence of the agents from the minimum of the objective functions. Finally, numerical examples are included to show the obtained theoretical results.

math.OC

Distributed optimization with nonconvex velocity constraints, nonuniform position constraints and nonuniform stepsizes

This note is devoted to the distributed optimization problem of multi-agent systems with nonconvex velocity constraints, nonuniform position constraints and nonuniform stepsizes. Two distributed constrained algorithms with nonconvex velocity constraints and nonuniform stepsizes are proposed in the absence and the presence of nonuniform position constraints by introducing a switching mechanism to guarantee all agents' position states to remain in a bounded region. The algorithm gains need not to be predesigned and can be selected by each agent using its own and neighbours' information. By a model transformation, the original nonlinear time-varying system is converted into a linear time-varying one with a nonlinear error term. Based on the properties of stochastic matrices, it is shown that the optimization problem can be solved as long as the communication topologies are jointly strongly connected and balanced. Numerical examples are given to show the obtained theoretical results.

math.OC

A Hardware Friendly Unsupervised Memristive Neural Network with Weight Sharing Mechanism

Memristive neural networks (MNNs), which use memristors as neurons or synapses, have become a hot research topic recently. However, most memristors are not compatible with mainstream integrated circuit technology and their stabilities in large-scale are not very well so far. In this paper, a hardware friendly MNN circuit is introduced, in which the memristive characteristics are implemented by digital integrated circuit. Through this method, spike timing dependent plasticity (STDP) and unsupervised learning are realized. A weight sharing mechanism is proposed to bridge the gap of network scale and hardware resource. Experiment results show the hardware resource is significantly saved with it, maintaining good recognition accuracy and high speed. Moreover, the tendency of resource increase is slower than the expansion of network scale, which infers our method's potential on large scale neuromorphic network's realization.

cs.ET

Long short-term memory networks in memristor crossbars

Recent breakthroughs in recurrent deep neural networks with long short-term memory (LSTM) units has led to major advances in artificial intelligence. State-of-the-art LSTM models with significantly increased complexity and a large number of parameters, however, have a bottleneck in computing power resulting from limited memory capacity and data communication bandwidth. Here we demonstrate experimentally that LSTM can be implemented with a memristor crossbar, which has a small circuit footprint to store a large number of parameters and in-memory computing capability that circumvents the 'von Neumann bottleneck'. We illustrate the capability of our system by solving real-world problems in regression and classification, which shows that memristor LSTM is a promising low-power and low-latency hardware platform for edge inference.

cs.ET