SearcharxivSearch

arXiv subjects

Nathan Leroux

Publications and source records attributed to Nathan Leroux.

10 recordsLinked to original sources

Gain Cell-Based Analog Content Addressable Memory for Dynamic Associative tasks in AI

Analog Content Addressable Memories (aCAMs) have proven useful for associative in-memory computing applications like Decision Trees, Finite State Machines, and Hyper-dimensional Computing. While non-volatile implementations using FeFETs and ReRAM devices offer speed, power, and area advantages, they suffer from slow write speeds and limited write cycles, making them less suitable for computations involving fully dynamic data patterns. To address these limitations, in this work, we propose a capacitor gain cell-based aCAM designed for dynamic processing, where frequent memory updates are required. Our system compares analog input voltages to boundaries stored in capacitors, enabling efficient dynamic tasks. We demonstrate the application of aCAM within transformer attention mechanisms by replacing the softmax-scaled dot-product similarity with aCAM similarity, achieving competitive results. Circuit simulations on a TSMC 28 nm node show promising performance in terms of energy efficiency, precision, and latency, making it well-suited for fast, dynamic AI applications.

cs.ET

Analog In-Memory Computing Attention Mechanism for Fast and Energy-Efficient Large Language Models

Transformer networks, driven by self-attention, are central to Large Language Models. In generative Transformers, self-attention uses cache memory to store token projections, avoiding recomputation at each time step. However, GPU-stored projections must be loaded into SRAM for each new generation step, causing latency and energy bottlenecks. We present a custom self-attention in-memory computing architecture based on emerging charge-based memories called gain cells, which can be efficiently written to store new tokens during sequence generation and enable parallel analog dot-product computation required for self-attention. However, the analog gain cell circuits introduce non-idealities and constraints preventing the direct mapping of pre-trained models. To circumvent this problem, we design an initialization algorithm achieving text processing performance comparable to GPT-2 without training from scratch. Our architecture respectively reduces attention latency and energy consumption by up to two and five orders of magnitude compared to GPUs, marking a significant step toward ultra-fast, low-power generative Transformers.

cs.NE

Classification of multi-frequency RF signals by extreme learning, using magnetic tunnel junctions as neurons and synapses

Extracting information from radiofrequency (RF) signals using artificial neural networks at low energy cost is a critical need for a wide range of applications from radars to health. These RF inputs are composed of multiples frequencies. Here we show that magnetic tunnel junctions can process analogue RF inputs with multiple frequencies in parallel and perform synaptic operations. Using a backpropagation-free method called extreme learning, we classify noisy images encoded by RF signals, using experimental data from magnetic tunnel junctions functioning as both synapses and neurons. We achieve the same accuracy as an equivalent software neural network. These results are a key step for embedded radiofrequency artificial intelligence.

cond-mat.mes-hall

Online Transformers with Spiking Neurons for Fast Prosthetic Hand Control

Transformers are state-of-the-art networks for most sequence processing tasks. However, the self-attention mechanism often used in Transformers requires large time windows for each computation step and thus makes them less suitable for online signal processing compared to Recurrent Neural Networks (RNNs). In this paper, instead of the self-attention mechanism, we use a sliding window attention mechanism. We show that this mechanism is more efficient for continuous signals with finite-range dependencies between input and target, and that we can use it to process sequences element-by-element, this making it compatible with online processing. We test our model on a finger position regression dataset (NinaproDB8) with Surface Electromyographic (sEMG) signals measured on the forearm skin to estimate muscle activities. Our approach sets the new state-of-the-art in terms of accuracy on this dataset while requiring only very short time windows of 3.5 ms at each inference step. Moreover, we increase the sparsity of the network using Leaky-Integrate and Fire (LIF) units, a bio-inspired neuron model that activates sparsely in time solely when crossing a threshold. We thus reduce the number of synaptic operations up to a factor of $\times5.3$ without loss of accuracy. Our results hold great promises for accurate and fast online processing of sEMG signals for smooth prosthetic hand control and is a step towards Transformers and Spiking Neural Networks (SNNs) co-integration for energy efficient temporal signal processing.

cs.NE

Multilayer spintronic neural networks with radio-frequency connections

Spintronic nano-synapses and nano-neurons perform complex cognitive computations with high accuracy thanks to their rich, reproducible and controllable magnetization dynamics. These dynamical nanodevices could transform artificial intelligence hardware, provided that they implement state-of-the art deep neural networks. However, there is today no scalable way to connect them in multilayers. Here we show that the flagship nano-components of spintronics, magnetic tunnel junctions, can be connected into multilayer neural networks where they implement both synapses and neurons thanks to their magnetization dynamics, and communicate by processing, transmitting and receiving radio frequency (RF) signals. We build a hardware spintronic neural network composed of nine magnetic tunnel junctions connected in two layers, and show that it natively classifies nonlinearly-separable RF inputs with an accuracy of 97.7%. Using physical simulations, we demonstrate that a large network of nanoscale junctions can achieve state-of the-art identification of drones from their RF transmissions, without digitization, and consuming only a few milliwatts, which is a gain of more than four orders of magnitude in power consumption compared to currently used techniques. This study lays the foundation for deep, dynamical, spintronic neural networks.

cs.ET

Convolutional Neural Networks with Radio-Frequency Spintronic Nano-Devices

Convolutional neural networks are state-of-the-art and ubiquitous in modern signal processing and machine vision. Nowadays, hardware solutions based on emerging nanodevices are designed to reduce the power consumption of these networks. Spintronics devices are promising for information processing because of the various neural and synaptic functionalities they offer. However, due to their low OFF/ON ratio, performing all the multiplications required for convolutions in a single step with a crossbar array of spintronic memories would cause sneak-path currents. Here we present an architecture where synaptic communications have a frequency selectivity that prevents crosstalk caused by sneak-path currents. We first demonstrate how a chain of spintronic resonators can function as synapses and make convolutions by sequentially rectifying radio-frequency signals encoding consecutive sets of inputs. We show that a parallel implementation is possible with multiple chains of spintronic resonators to avoid storing intermediate computational steps in memory. We propose two different spatial arrangements for these chains. For each of them, we explain how to tune many artificial synapses simultaneously, exploiting the synaptic weight sharing specific to convolutions. We show how information can be transmitted between convolutional layers by using spintronic oscillators as artificial microwave neurons. Finally, we simulate a network of these radio-frequency resonators and spintronic oscillators to solve the MNIST handwritten digits dataset, and obtain results comparable to software convolutional neural networks. Since it can run convolutional neural networks fully in parallel in a single step with nano devices, the architecture proposed in this paper is promising for embedded applications requiring machine vision, such as autonomous driving.

cs.ET

Hardware realization of the multiply and accumulate operation on radio-frequency signals with magnetic tunnel junctions

Artificial neural networks are a valuable tool for radio-frequency (RF) signal classification in many applications, but digitization of analog signals and the use of general purpose hardware non-optimized for training make the process slow and energetically costly. Recent theoretical work has proposed to use nano-devices called magnetic tunnel junctions, which exhibit intrinsic RF dynamics, to implement in hardware the Multiply and Accumulate (MAC) operation, a key building block of neural networks, directly using analogue RF signals. In this article, we experimentally demonstrate that a magnetic tunnel junction can perform multiplication of RF powers, with tunable positive and negative synaptic weights. Using two magnetic tunnel junctions connected in series we demonstrate the MAC operation and use it for classification of RF signals. These results open the path to embedded systems capable of analyzing RF signals with neural networks directly after the antenna, at low power cost and high speed.

cond-mat.dis-nn

Tuning the performance of a micrometer-sized Stirling engine through reservoir engineering

Colloidal heat engines are paradigmatic models to understand the conversion of heat into work in a noisy environment - a domain where biological and synthetic nano/micro machines function. While the operation of these engines across thermal baths is well-understood, how they function across baths with noise statistics that is non-Gaussian and also lacks memory, the simplest departure from equilibrium, remains unclear. Here we quantified the performance of a colloidal Stirling engine operating between an engineered \textit{memoryless} non-Gaussian bath and a Gaussian one. In the quasistatic limit, the non-Gaussian engine functioned like an equilibrium one as predicted by theory. On increasing the operating speed, due to the nature of noise statistics, the onset of irreversibility for the non-Gaussian engine preceded its thermal counterpart and thus shifted the operating speed at which power is maximum. The performance of nano/micro machines can be tuned by altering only the nature of reservoir noise statistics.

cond-mat.stat-mech

Wireless communication between two magnetic tunnel junctions acting as oscillator and diode

Magnetic tunnel junctions are nanoscale spintronic devices with microwave generation and detection capabilities. Here we use the rectification effect called "spin-diode" in a magnetic tunnel junction to wirelessly detect the microwave emission of another junction in the auto-oscillatory regime. We show that the rectified spin-diode voltage measured at the receiving junction end can be reconstructed from the independently measured auto-oscillation and spin diode spectra in each junction. Finally we adapt the auto-oscillator model to the case of spin-torque oscillator and spin-torque diode and we show that accurately reproduces the experimentally observed features. These results will be useful to design circuits and chips based on spintronic nanodevices communicating through microwaves.

cond-mat.mes-hall

Reservoir computing with the frequency, phase and amplitude of spin-torque nano-oscillators

Spin-torque nano-oscillators can emulate neurons at the nanoscale. Recent works show that the non-linearity of their oscillation amplitude can be leveraged to achieve waveform classification for an input signal encoded in the amplitude of the input voltage. Here we show that the frequency and the phase of the oscillator can also be used to recognize waveforms. For this purpose, we phase-lock the oscillator to the input waveform, which carries information in its modulated frequency. In this way we considerably decrease amplitude, phase and frequency noise. We show that this method allows classifying sine and square waveforms with an accuracy above 99% when decoding the output from the oscillator amplitude, phase or frequency. We find that recognition rates are directly related to the noise and non-linearity of each variable. These results prove that spin-torque nano-oscillators offer an interesting platform to implement different computing schemes leveraging their rich dynamical features.

physics.app-ph