SearcharxivSearch

arXiv subjects

Farshad Moradi

Publications and source records attributed to Farshad Moradi.

14 recordsLinked to original sources

Event-Driven Language Models with Sparse Neural Activity for Neuromorphic Hardware

Inference with transformer-based large language models (LLMs) is often limited by the memory-bound KV cache and quadratic attention cost. State-space models (SSMs) mitigate this through linear attention and fixed-size recurrent states, but their large dense linear projections remain computationally expensive even after quantization. We introduce a method that induces sparse neural activity in heavily quantized linear-attention models with minimal performance loss. Activations below a per-projection trainable threshold ($\pm \Delta$) are nullified while preserving crucial outliers, achieving comparable performance to dense models with up to 4$\times$ fewer effective arithmetic operations. Targeting a multi-core, multi-chip neuromorphic platform, where event-driven execution converts unstructured sparsity into throughput at both the compute and communication levels, a capability GPU architectures fundamentally lack, we project up to 37$\times$ higher throughput and 16$\times$ lower power versus edge GPU inference of a comparable transformer-based model, and up to 5.4$\times$ improvements over the non-sparsified baseline. These results position sparse, quantized linear-attention models as a natural fit for deploying LLMs on event-driven multi-core platforms.

cs.NE

Non-uniform Memory Partitioning For Low-Power Spiking Neural Networks

Spiking Neural Networks (SNNs) naturally excel in processing temporally rich and sparse data. However, because of their time-stepped processing, memory access, specifically to synaptic weights stored in SRAM (static random-access memory), tends to dominate total power consumption. To address this issue, without incurring a large area overhead, we propose to leverage the greatly varying average firing rate of neurons in the network to efficiently allocate synaptic weights to an on-chip memory consisting of multiple non-uniformly sized memory banks. By assigning weights of frequently firing neurons to shallow, low-access cost memory and less actively accessed weights to deeper, high-density memories, the average power consumption of the synaptic weight memory is decreased without incurring a large area overhead. To benchmark our proposed architecture and find optimal configurations of memory arrangements, we perform an automatic exploration based on application requirements and hardware constraints. For memory designs synthesized in 28-nm CMOS technology, we show that our architecture can achieve a synaptic weight memory access power reduction of up to 61\% compared to a conventional design, with a 2.1$\times$ lower area overhead, as compared to a traditional uniformly partitioned memory bank that achieves a comparable reduction.

cs.AR

Reconfigurable Multistate MRAM Synapses with Vortex STNO based Neurons for Scalable In-Memory Convolutional Neural Networks

Magnetic tunnel junction (MTJ)-based magnetic random-access memory (MRAM) is a promising platform for neuromorphic and in-memory computing owing to its non-volatility, high endurance, fast switching dynamics and CMOS compatibility. However, conventional spin-transfer torque and spin-orbit torque MRAM implementations for neural networks often suffer from high critical switching currents, large latency, thermal instability and significant read-write overheads. Here, we demonstrate a unified multistate MRAM-spin-torque nano-oscillator (STNO) architecture that integrates synapses and neurons on a single chip for convolutional neural network (CNN) applications. The system employs 1x8 multistate MRAM arrays as programmable synapses coupled with a vortex-based STNO neuron, enabling both individual and collective programming through fieldline-driven write channels. Multiple configurable resistance states are achieved by tuning internal and external magnetic fields together with bias currents, allowing quantized positive and negative synaptic weights for configurable kernel and pooling operations. The proposed architecture is evaluated through simulation on MNIST, SVHN, CIFAR-10, Google Speech Commands (GSC) and RadioML datasets, achieving accuracy of 99.76%, 87.93%, 78.14%, 87.96% and 56.46% respectively. Based on fabricated device dimensions, the complete architecture occupies ~6171.2 {\mu}m2 with an average energy consumption of 200.08 pJ per training and inference cycle for MNIST, highlighting its potential for scalable low-power neuromorphic computing

physics.app-ph

Closed-loop control of seizure activity via real-time seizure forecasting by reservoir neuromorphic computing

Closed-loop brain stimulation holds potential as personalized treatment for drug-resistant epilepsy (DRE) but still suffers from limitations that result in highly variable efficacy. First, stimulation is typically delivered upon detection of the seizure to abort rather than prevent it; second, the stimulation parameters are established by trial and error, requiring lengthy rounds of fine-tuning, which delay steady-state therapeutic efficacy. Here, we address these limitations by leveraging the potential of neuromorphic computing. We present a neuromorphic reservoir computing hardware system capable of driving real-time personalized free-run stimulations based on seizure forecasting, wherein each forecast triggers an electrical pulse rather than an arbitrarily predefined fixed-frequency stimulus train. The system achieves 83.33% accuracy in forecasting seizure occurrences during the training phase. We validate the system using hippocampal spheroids coupled to 3D microelectrode array as a simplified testbed, achieving seizure reduction >97% during the real-time processing while primarily using instantaneous stimulation frequencies within 20 Hz, well below what typically used in clinical practice. Our work demonstrates the potential of neuromorphic systems as a next-generation neuromodulation strategy for personalized DRE treatment, leveraging their sparse and event-driven processing for real-time applications.

cs.AI

Score-based Generative Diffusion Models to Synthesize Full-dose FDG Brain PET from MRI in Epilepsy Patients

Fluorodeoxyglucose (FDG) PET to evaluate patients with epilepsy is one of the most common applications for simultaneous PET/MRI, given the need to image both brain structure and metabolism, but is suboptimal due to the radiation dose in this young population. Little work has been done synthesizing diagnostic quality PET images from MRI data or MRI data with ultralow-dose PET using advanced generative AI methods, such as diffusion models, with attention to clinical evaluations tailored for the epilepsy population. Here we compared the performance of diffusion- and non-diffusion-based deep learning models for the MRI-to-PET image translation task for epilepsy imaging using simultaneous PET/MRI in 52 subjects (40 train/2 validate/10 hold-out test). We tested three different models: 2 score-based generative diffusion models (SGM-Karras Diffusion [SGM-KD] and SGM-variance preserving [SGM-VP]) and a Transformer-Unet. We report results on standard image processing metrics as well as clinically relevant metrics, including congruency measures (Congruence Index and Congruency Mean Absolute Error) that assess hemispheric metabolic asymmetry, which is a key part of the clinical analysis of these images. The SGM-KD produced the best qualitative and quantitative results when synthesizing PET purely from T1w and T2 FLAIR images with the least mean absolute error in whole-brain specific uptake value ratio (SUVR) and highest intraclass correlation coefficient. When 1% low-dose PET images are included in the inputs, all models improve significantly and are interchangeable for quantitative performance and visual quality. In summary, SGMs hold great potential for pure MRI-to-PET translation, while all 3 model types can synthesize full-dose FDG-PET accurately using MRI and ultralow-dose PET.

eess.IV

Hybrid Opto-Electrical Excitation of Spin-Transfer Torque Nano-Oscillators for Advanced Computing

Neuromorphic computing, inspired by the brain's parallel and energy-efficient processing, offers a transformative approach to artificial intelligence. In this study, we fabricated optimized spin-transfer torque nano-oscillators (STNOs) and investigated their dynamic behaviors using a hybrid excitation scheme combining AC laser illumination and DC bias currents. Laser-induced thermal gradients generate pulsed thermoelectric voltages ($V_{\text{AC}}$) via the Tunnel Magneto-Seebeck (TMS) effect, while the addition of bias currents enhances this response, producing both $V_{\text{AC}}$ and a DC component ($V_{\text{DC}}$). Magnetic field sweeps reveal distinct switching between parallel (P) and antiparallel (AP) magnetization states in both voltage components, supporting multistate memory applications. Millivolt-range thermovoltage signals in open-circuit conditions demonstrate CMOS compatibility, enabling simplified, scalable neuromorphic systems. Under biased conditions, enhanced thermovoltage outputs exhibit intriguing phenomena, including spikes correlated with Barkhausen jumps and double-switching behavior, offering insights into magnetization dynamics and vortex transitions. These features resemble neural spiking behavior, suggesting applications in spiking neural networks, reservoir computing, multistate logic, analog computing, and high-resolution sensing. By bridging spintronic phenomena with practical applications, this work provides a versatile platform for next-generation AI technologies and adaptive computing architectures.

physics.optics

Spin Hall Nano-Oscillator Empirical Electrical Model for Optimal On-chip Detector Design

As nascent nonlinear oscillators, nano-constriction spin Hall nano-oscillators (SHNOs) represent a promising potential for integration into more complicated systems such as neural networks, magnetic field sensors, and radio frequency (RF) signal classification, their tunable high-frequency operating regime, easy synchronization, and CMOS compatibility can streamline the process. To implement SHNOs in any of these networks, the electrical features of a single device are needed before designing the signal detection CMOS circuitry. This study centers on presenting an empirical electrical model of the SHNO based on a comprehensive characterization of the output impedance of a single SHNO, and its available output power in the range of 2-10 GHz at various bias currents.

cond-mat.mes-hall

In the realm of hybrid Brain: Human Brain and AI

With the recent developments in neuroscience and engineering, it is now possible to record brain signals and decode them. Also, a growing number of stimulation methods have emerged to modulate and influence brain activity. Current brain-computer interface (BCI) technology is mainly on therapeutic outcomes, it already demonstrated its efficiency as assistive and rehabilitative technology for patients with severe motor impairments. Recently, artificial intelligence (AI) and machine learning (ML) technologies have been used to decode brain signals. Beyond this progress, combining AI with advanced BCIs in the form of implantable neurotechnologies grants new possibilities for the diagnosis, prediction, and treatment of neurological and psychiatric disorders. In this context, we envision the development of closed loop, intelligent, low-power, and miniaturized neural interfaces that will use brain inspired AI techniques with neuromorphic hardware to process the data from the brain. This will be referred to as Brain Inspired Brain Computer Interfaces (BI-BCIs). Such neural interfaces would offer access to deeper brain regions and better understanding for brain's functions and working mechanism, which improves BCIs operative stability and system's efficiency. On one hand, brain inspired AI algorithms represented by spiking neural networks (SNNs) would be used to interpret the multimodal neural signals in the BCI system. On the other hand, due to the ability of SNNs to capture rich dynamics of biological neurons and to represent and integrate different information dimensions such as time, frequency, and phase, it would be used to model and encode complex information processing in the brain and to provide feedback to the users. This paper provides an overview of the different methods to interface with the brain, presents future applications and discusses the merger of AI and BCIs.

eess.SY

Comparative Analysis of THz Signal Emission from SiO$_2$/CoFeB/Metal Heterostructures: Wideband and High-Frequency THz Signal Advantage of PtBi-based Emitter

Spintronic THz emitters have attracted much attention due to their desirable properties, such as affordability, ultra-wideband capability, high efficiency, and tunable polarization. In this study, we investigate the characteristics of THz signals, including their frequency, bandwidth, and amplitude, emitted from a series of heterostructures with ferromagnetic (FM) and nonmagnetic (NM) materials. The FM layer consists of a wedge-shaped CoFeB layer with a thickness of 0 to 5 nm, while the NM materials include various metals such as Pt, Au, W, Ru, Pt$_{\%92}$Bi$_{\%8}$, and Ag$_{\%90}$Bi$_{\%10}$ alloys. Our experiments show that the emitter with Pt-NM layer has the highest amplitude of the emitted THz signal. However, the PtBi-based emitter exhibits a higher central THz peak and wider bandwidth, making it a promising candidate for broadband THz emitters. These results pave the way for further exploration of the specific compositions of Pt$_{1-x}$Bi$_{x}$ for THz emitter design, especially with the goal of generating higher frequency and wider bandwidth THz signals. These advances hold significant potential for applications in various fields such as high-resolution imaging, spectroscopy, communications, medical diagnostics, and more.

physics.optics

Enhancing Spin Transfer Torque in Magnetic Tunnel Junction Devices: Exploring the Influence of Capping Layer Materials and Thickness on Device Characteristics

We have developed and optimized two categories of spin transfer torque magnetic tunnel junctions (STT-MTJs) that exhibit a high tunnel magnetoresistance (TMR) ratio, low critical current, high outputpower in the micro watt range, and auto-oscillation behavior. These characteristics demonstrate the potential of STT-MTJs for low-power, high-speed, and reliable spintronic applications, including magnetic memory, logic, and signal processing. The only distinguishing factor between the two categories, denoted as A-MTJs and B-MTJs, is the composition of their free layers, 2 CoFeB/0.21 Ta/6 CoFeSiB for A-MTJs and 2 CoFeB/0.21 Ta/7 NiFe for B-MTJs. Our study reveals that B-MTJs exhibit lower critical currents for auto-oscillation than A-MTJs. We found that both stacks have comparable saturation magnetization and anisotropy field, suggesting that the difference in auto-oscillation behavior is due to the higher damping of A-MTJs compared to B-MTJs. To verify this hypothesis, we employed the all-optical time-resolved magneto-optical Kerr effect (TRMOKE) technique, which confirmed that STT-MTJs with lower damping exhibited auto-oscillation at lower critical current values. Additionally, our study aimed to optimize the STT-MTJ performance by investigating the impact of the capping layer on the device's response to electronic and optical stimuli.

physics.optics

Spin-Orbit Torque Flash Analog-to-Digital Converter

Although Analog-to-digital converters (ADCs) are critical components in mixed-signal integrated circuits (IC), their performance has not been improved significantly over the last decade. To achieve a radical improvement (compact, low power and reliable ADCs), spintronics can be considered as a proper candidate due to its compatibility with CMOS and wide applications in storage, neuromorphic computing, and so on. In this paper, a proof-of-concept of a 3-bit spin-CMOS Flash ADC using in-plane-anisotropy magnetic tunnel junctions (i-MTJs) with spin-orbit torque (SOT) switching mechanism is designed, fabricated and characterized. The proposed ADC replaces the current mirrors and power-hungry comparators in the conventional Flash ADC with seven parallel i-MTJs with different heavy metal (HM) widths. Monte-Carlo simulations based on the experimental measurements show the process variations/mismatch limits the accuracy of the proposed ADC to 2 bits. Moreover, the maximum differential nonlinearity (DNL) and integral nonlinearity (INL) are 0.739 LSB (least significant bit) and 0.7319 LSB, respectively.

cs.ET

NET-TEN: a silicon neuromorphic network for low-latency detection of seizures in local field potentials

Therapeutic intervention in neurological disorders still relies heavily on pharmacological solutions, while the treatment of patients with drug resistance remains an open challenge. This is particularly true for patients with epilepsy, 30% of whom are refractory to medications. Implantable devices for chronic recording and electrical modulation of brain activity have proved a viable alternative in such cases. To operate, the device should detect the relevant electrographic biomarkers from Local Field Potentials (LFPs) and determine the right time for stimulation. To enable timely interventions, the ideal device should attain biomarker detection with low latency while operating under low power consumption to prolong the battery life. Neuromorphic networks have progressively gained reputation as low-latency low-power computing systems, which makes them a promising candidate as processing core of next-generation implantable neural interfaces. Here we introduce a fully-analog neuromorphic device implemented in CMOS technology for analyzing LFP signals in an in vitro model of acute ictogenesis. We show that the system can detect ictal and interictal events with ms-latency and with high precision, consuming on average 3.50 nW during the task. Our work paves the way to a new generation of brain implantable devices for personalized closed-loop stimulation for epilepsy treatment.

cs.HC

Analysis and Design of a PMUT-based transducer for Powering Brain Implants

This paper presents an analytical design of an ultrasonic power transfer system based on piezoelectric micro-machined ultrasonic transducer (PMUT) for fully wireless brain implants in mice. The key steps like the material selection of each layer and the top electrode radius to maximize the coupling factor are well-detailed. This approach results in the design of a single cell with a high effective coupling coefficient. Furthermore, compact models are used to make the design process less time-consuming for designers. These models are based on the equivalent circuit theory for the PMUT. A cell of 107 um in radius, 5 um in thickness of Lead Zirconate Titanium (PZT), and 10 um in thickness of silicon (Si) is found to have a 4% of effective coupling coefficient among the highest values for a clamped edge boundary conditions. Simulation results show a frequency of 2.84 MHz as resonance. In case of an array, mutual impedance and numerical modeling are used to estimate the distance between the adjacent cells. In addition, the area of the proposed transducer and the number of cells are computed with the Rayleigh distance and neglecting the cross-talk among cells, respectively. The designed transducer consists of 7x7 cells in an area of 3.24 mm2. The transducer is able to deliver an acoustic intensity of 7.185 mW/mm2 for a voltage of 19.5 V for powering brain implants seated in the motor cortex and striatum of the mice's brain. The maximum acoustic intensity occurs at a distance of 2.5 mm in the near field which was estimated with the Rayleigh length equation.

physics.med-ph

Spin-Orbit-Torque-based Devices, Circuits and Architectures

Spintronics, the use of spin of an electron instead of its charge, has received huge attention from research communities for different applications including memory, interconnects, logic implementation, neuromorphic computing, and many other applications. Here, in this paper, we review the works within spintronics, more specifically on spin-orbit torque (SOT) within different research groups. We also provide researchers an insight into the future potentials of the SOT-based designs. This comprehensive review paper covers different aspects of SOT-based design from device and circuit to architecture level as well as more ambitious and futuristic applications of such technology.

cs.ET