SearcharxivSearch

arXiv subjects

Huaqiang Wu

Publications and source records attributed to Huaqiang Wu.

13 recordsLinked to original sources

Training deep physical neural networks with local physical information bottleneck

Deep learning has revolutionized modern society but faces growing energy and latency constraints. Deep physical neural networks (PNNs) are interconnected computing systems that directly exploit analog dynamics for energy-efficient, ultrafast AI execution. Realizing this potential, however, requires universal training methods tailored to physical intricacies. Here, we present the Physical Information Bottleneck (PIB), a general and efficient framework that integrates information theory and local learning, enabling deep PNNs to learn under arbitrary physical dynamics. By allocating matrix-based information bottlenecks to each unit, we demonstrate supervised, unsupervised, and reinforcement learning across electronic memristive chips and optical computing platforms. PIB also adapts to severe hardware faults and allows for parallel training via geographically distributed resources. Bypassing auxiliary digital models and contrastive measurements, PIB recasts PNN training as an intrinsic, scalable information-theoretic process compatible with diverse physical substrates.

cs.LG

Neuromorphic spatiotemporal optical flow: Enabling ultrafast visual perception beyond human capabilities

Optical flow, inspired by the mechanisms of biological visual systems, calculates spatial motion vectors within visual scenes that are necessary for enabling robotics to excel in complex and dynamic working environments. However, current optical flow algorithms, despite human-competitive task performance on benchmark datasets, remain constrained by unacceptable time delays (~0.6 seconds per inference, 4X human processing speed) in practical deployment. Here, we introduce a neuromorphic optical flow approach that addresses delay bottlenecks by encoding temporal information directly in a synaptic transistor array to assist spatial motion analysis. Compared to conventional spatial-only optical flow methods, our spatiotemporal neuromorphic optical flow offers the spatial-temporal consistency of motion information, rapidly identifying regions of interest in as little as 1-2 ms using the temporal motion cues derived from the embedded temporal information in the two-dimensional floating gate synaptic transistors. Thus, the visual input can be selectively filtered to achieve faster velocity calculations and various task execution. At the hardware level, due to the atomically sharp interfaces between distinct functional layers in two-dimensional van der Waals heterostructures, the synaptic transistor offers high-frequency response (~100 μs), robust non-volatility (>10000 s), and excellent endurance (>8000 cycles), enabling robust visual processing. In software benchmarks, our system outperforms state-of-the-art algorithms with a 400% speedup, frequently surpassing human-level performance while maintaining or enhancing accuracy by utilizing the temporal priors provided by the embedded temporal information.

cs.CV

Distributed Representations Enable Robust Multi-Timescale Symbolic Computation in Neuromorphic Hardware

Programming recurrent spiking neural networks (RSNNs) to robustly perform multi-timescale computation remains a difficult challenge. To address this, we describe a single-shot weight learning scheme to embed robust multi-timescale dynamics into attractor-based RSNNs, by exploiting the properties of high-dimensional distributed representations. We embed finite state machines into the RSNN dynamics by superimposing a symmetric autoassociative weight matrix and asymmetric transition terms, which are each formed by the vector binding of an input and heteroassociative outer-products between states. Our approach is validated through simulations with highly nonideal weights; an experimental closed-loop memristive hardware setup; and on Loihi 2, where it scales seamlessly to large state machines. This work introduces a scalable approach to embed robust symbolic computation through recurrent dynamics into neuromorphic hardware, without requiring parameter fine-tuning or significant platform-specific optimisation. Moreover, it demonstrates that distributed symbolic representations serve as a highly capable representation-invariant language for cognitive algorithms in neuromorphic hardware.

cs.NE

Synergistic Development of Perovskite Memristors and Algorithms for Robust Analog Computing

Analog computing using non-volatile memristors has emerged as a promising solution for energy-efficient deep learning. New materials, like perovskites-based memristors are recently attractive due to their cost-effectiveness, energy efficiency and flexibility. Yet, challenges in material diversity and immature fabrications require extensive experimentation for device development. Moreover, significant non-idealities in these memristors often impede them for computing. Here, we propose a synergistic methodology to concurrently optimize perovskite memristor fabrication and develop robust analog DNNs that effectively address the inherent non-idealities of these memristors. Employing Bayesian optimization (BO) with a focus on usability, we efficiently identify optimal materials and fabrication conditions for perovskite memristors. Meanwhile, we developed "BayesMulti", a DNN training strategy utilizing BO-guided noise injection to improve the resistance of analog DNNs to memristor imperfections. Our approach theoretically ensures that within a certain range of parameter perturbations due to memristor non-idealities, the prediction outcomes remain consistent. Our integrated approach enables use of analog computing in much deeper and wider networks, which significantly outperforms existing methods in diverse tasks like image classification, autonomous driving, species identification, and large vision-language models, achieving up to 100-fold improvements. We further validate our methodology on a 10$\times$10 optimized perovskite memristor crossbar, demonstrating high accuracy in a classification task and low energy consumption. This study offers a versatile solution for efficient optimization of various analog computing systems, encompassing both devices and algorithms.

cs.LG

Scaling Limits of Memristor-Based Routers for Asynchronous Neuromorphic Systems

Multi-core neuromorphic systems typically use on-chip routers to transmit spikes among cores. These routers require significant memory resources and consume a large part of the overall system's energy budget. A promising alternative approach to using standard CMOS and SRAM-based routers is to exploit the features of memristive crossbar arrays and use them as programmable switch-matrices that route spikes. However, the scaling of these crossbar arrays presents physical challenges, such as "IR drop" on the metal lines due to the parasitic resistance, and leakage current accumulation on multiple active memristors in their "off" state. While reliability challenges of this type have been extensively studied in synchronous systems for compute-in-memory matrix-vector multiplication (MVM) accelerators and storage class memory, little effort has been devoted so far to characterizing the scaling limits of memristor-based crossbar routers. Here, we study the challenges of memristive crossbar arrays, when used as routing channels to transmit spikes in asynchronous Spiking Neural Network (SNN) hardware. We validate our analytical findings with experimental results obtained from a 4K-ReRAM chip which demonstrates its functionality as a routing crossbar. We determine the functionality bounds on the routing due to the IR drop and leak problem, based on theoretical modeling, circuit simulations for a 22nm FDSOI technology, and experimental measurements. This work highlights the limitations of this approach and provides useful guidelines for engineering the memristor device properties in memristive crossbar routers for multi-core asynchronous neuromorphic systems.

cs.ET

Multiferroic Magnon Spin-Torque Based Reconfigurable Logic-In-Memory

Magnons, bosonic quasiparticles carrying angular momentum, can flow through insulators for information transmission with minimal power dissipation. However, it remains challenging to develop a magnon-based logic due to the lack of efficient electrical manipulation of magnon transport. Here we present a magnon logic-in-memory device in a spin-source/multiferroic/ferromagnet structure, where multiferroic magnon modes can be electrically excited and controlled. In this device, magnon information is encoded to ferromagnetic bits by the magnon-mediated spin torque. We show that the ferroelectric polarization can electrically modulate the magnon spin-torque by controlling the non-collinear antiferromagnetic structure in multiferroic bismuth ferrite thin films with coupled antiferromagnetic and ferroelectric orders. By manipulating the two coupled non-volatile state variables (ferroelectric polarization and magnetization), we further demonstrate reconfigurable logic-in-memory operations in a single device. Our findings highlight the potential of multiferroics for controlling magnon information transport and offer a pathway towards room-temperature voltage-controlled, low-power, scalable magnonics for in-memory computing.

physics.app-ph

Acoustic-Driven Magnetic Skyrmion Motion

Magnetic skyrmions have great potential for developing novel spintronic devices. The electrical manipulation of skyrmions has mainly relied on current-induced spin-orbit torques. A recent theoretical model suggested that the skyrmions could be more efficiently manipulated by surface acoustic waves (SAW), an elastic wave that can couple with magnetic moment through magnetoelastic effect. However, the directional motion of skyrmions that is driven by SAW is still missing. Here, we experimentally demonstrate the motion of Néel-type skyrmions in Ta/CoFeB/MgO/Ta multilayers driven by propagating SAW pulses from on-chip piezoelectric transducers. Our results reveal that the elastic wave with longitudinal and shear vertical displacements (Rayleigh wave) traps skyrmions, while the shear horizontal wave effectively drives the motion of skyrmions. In particular, a longitudinal motion along the SAW propagation direction and a transverse motion due to topological charge, are observed and further confirmed by our micromagnetic simulations. This work demonstrates a promising approach based on acoustic waves for manipulating skyrmions, which could offer new opportunities for ultra-low power spintronics.

cond-mat.mtrl-sci

Large-Scale Integrated Flexible Tactile Sensor Array for Sensitive Smart Robotic Touch

In the long pursuit of smart robotics, it has been envisioned to empower robots with human-like senses, especially vision and touch. While tremendous progress has been made in image sensors and computer vision over the past decades, the tactile sense abilities are lagging behind due to the lack of large-scale flexible tactile sensor array with high sensitivity, high spatial resolution, and fast response. In this work, we have demonstrated a 64x64 flexible tactile sensor array with a record-high spatial resolution of 0.9 mm (equivalently 28.2 pixels per inch), by integrating a high-performance piezoresistive film (PRF) with a large-area active matrix of carbon nanotube thin-film transistors. PRF with self-formed microstructures exhibited high pressure-sensitivity of ~385 kPa-1 for MWCNTs concentration of 6%, while the 14% one exhibited fast response time of ~3 ms, good linearity, broad detection range beyond 1400 kPa, and excellent cyclability over 3000 cycles. Using this fully integrated tactile sensor array, the footprint maps of an artificial honeybee were clearly identified. Furthermore, we hardware-implemented a smart tactile system by integrating the PRF-based sensor array with a memristor-based computing-in-memory chip to record and recognize handwritten digits and Chinese calligraphy, achieving high classification accuracies of 98.8% and 97.3% in hardware, respectively. The integration of sensor networks with deep learning hardware may enable edge or near-sensor computing with significantly reduced power consumption and latency. Our work could pave the road to building large-scale intelligent sensor networks for next-generation smart robotics.

cond-mat.mtrl-sci

Edge AI without Compromise: Efficient, Versatile and Accurate Neurocomputing in Resistive Random-Access Memory

Realizing today's cloud-level artificial intelligence functionalities directly on devices distributed at the edge of the internet calls for edge hardware capable of processing multiple modalities of sensory data (e.g. video, audio) at unprecedented energy-efficiency. AI hardware architectures today cannot meet the demand due to a fundamental "memory wall": data movement between separate compute and memory units consumes large energy and incurs long latency. Resistive random-access memory (RRAM) based compute-in-memory (CIM) architectures promise to bring orders of magnitude energy-efficiency improvement by performing computation directly within memory. However, conventional approaches to CIM hardware design limit its functional flexibility necessary for processing diverse AI workloads, and must overcome hardware imperfections that degrade inference accuracy. Such trade-offs between efficiency, versatility and accuracy cannot be addressed by isolated improvements on any single level of the design. By co-optimizing across all hierarchies of the design from algorithms and architecture to circuits and devices, we present NeuRRAM - the first multimodal edge AI chip using RRAM CIM to simultaneously deliver a high degree of versatility for diverse model architectures, record energy-efficiency $5\times$ - $8\times$ better than prior art across various computational bit-precisions, and inference accuracy comparable to software models with 4-bit weights on all measured standard AI benchmarks including accuracy of 99.0% on MNIST and 85.7% on CIFAR-10 image classification, 84.7% accuracy on Google speech command recognition, and a 70% reduction in image reconstruction error on a Bayesian image recovery task. This work paves a way towards building highly efficient and reconfigurable edge AI hardware platforms for the more demanding and heterogeneous AI applications of the future.

cs.AR

Large-scale neuromorphic optoelectronic computing with a reconfigurable diffractive processing unit

Application-specific optical processors have been considered disruptive technologies for modern computing that can fundamentally accelerate the development of artificial intelligence (AI) by offering substantially improved computing performance. Recent advancements in optical neural network architectures for neural information processing have been applied to perform various machine learning tasks. However, the existing architectures have limited complexity and performance; and each of them requires its own dedicated design that cannot be reconfigured to switch between different neural network models for different applications after deployment. Here, we propose an optoelectronic reconfigurable computing paradigm by constructing a diffractive processing unit (DPU) that can efficiently support different neural networks and achieve a high model complexity with millions of neurons. It allocates almost all of its computational operations optically and achieves extremely high speed of data modulation and large-scale network parameter updating by dynamically programming optical modulators and photodetectors. We demonstrated the reconfiguration of the DPU to implement various diffractive feedforward and recurrent neural networks and developed a novel adaptive training approach to circumvent the system imperfections. We applied the trained networks for high-speed classifying of handwritten digit images and human action videos over benchmark datasets, and the experimental results revealed a comparable classification accuracy to the electronic computing approaches. Furthermore, our prototype system built with off-the-shelf optoelectronic components surpasses the performance of state-of-the-art graphics processing units (GPUs) by several times on computing speed and more than an order of magnitude on system energy efficiency.

eess.IV

Current-induced in-plane magnetization switching in biaxial ferrimagnetic insulator

Ferrimagnetic insulators (FiMI) have been intensively used in microwave and magneto-optical devices as well as spin caloritronics, where their magnetization direction plays a fundamental role on the device performance. The magnetization is generally switched by applying external magnetic fields. Here we investigate current-induced spin-orbit torque (SOT) switching of the magnetization in Y3Fe5O12 (YIG)/Pt bilayers with in-plane magnetic anisotropy, where the switching is detected by spin Hall magnetoresistance. Reversible switching is found at room temperature for a threshold current density of 10^7 A cm^-2. The YIG sublattices with antiparallel and unequal magnetic moments are aligned parallel or antiparallel to the direction of current pulses, which is consistent to the Neel order switching in antiferromagnetic system. It is proposed that such a switching behavior may be triggered by the antidamping-torque acting on the two antiparallel sublattices of FiMI. Our finding not only broadens the magnetization switching by electrical means and promotes the understanding of magnetization switching, but also paves the way for all-electrically modulated microwave devices and spin caloritronics with low power consumption.

cond-mat.mtrl-sci

Thermal generation, manipulation and detection of skyrmions

Recent years have witnessed significant progresses in realizing skyrmions in chiral magnets1-4 and asymmetric magnetic multilayers5-13, as well as their electrical manipulation2,7,8,10. Equally important, thermal generation, manipulation and detection of skyrmions can be exploited for prototypical new architecture with integrated computation14 and energy harvesting15. It has yet to verify if skyrmions can be purely generated by heating16,17, and if their resultant direction of motion driven by temperature gradients follows the diffusion or, oppositely, the magnonic spin torque17-21. Here, we address these important issues in microstructured devices made of multilayers: (Ta_CoFeB_MgO)15, (Pt_CoFeB_MgO_Ta)15 and (Pt_Co_Ta)15 integrated with on-chip heaters, by using a full-field soft X-ray microscopy. The thermal generation of densely packed skyrmions is attributed to the low energy barrier at the device edge, together with the thermally induced morphological transition from stripe domains to skyrmions. The unidirectional diffusion of skyrmions from the hot region towards the cold region is experimentally observed. It can be theoretically explained by the combined contribution from repulsive forces between skyrmions, and thermal spin-orbit torques in competing with magnonic spin torques17,18,20,21 and entropic forces22. These thermally generated skyrmions can be further electrically detected by measuring the accompanied anomalous Nernst voltages23. The on-chip thermoelectric generation, manipulation and detection of skyrmions could open another exciting avenue for enabling skyrmionics, and promote interdisciplinary studies among spin caloritronics15, magnonics24 and skyrmionics3,4,12.

cond-mat.mes-hall

Magnetoelectric coupling induced by interfacial orbital reconstruction

The magnetoelectric coupling effect with profound physics and enormous potential applications has provoked a great number of research activities in materials science. Here, we report that the reversible orbital reconstruction driven by ferroelectric polarization modulates the magnetic performance of ferroelectric ferromagnetic heterostructure. Mn in plane orbital occupancy and related interfacial exotic magnetic state are enhanced and weakened by the negative and positive electric field, respectively. Our findings thus not only present a broad opportunity to fill the missing member, orbital in the mechanism of magnetoelectric coupling, but also make the orbital degree of freedom straight forward to the application in microelectronic device.

cond-mat.str-el