SearcharxivSearch

arXiv subjects

Sandip Lashkare

Publications and source records attributed to Sandip Lashkare.

13 recordsLinked to original sources

A Unified Platform to Evaluate STDP Learning Rule and Synapse Model using Pattern Recognition in a Spiking Neural Network

We develop a unified platform to evaluate Ideal, Linear, and Non-linear $\text{Pr}_{0.7}\text{Ca}_{0.3}\text{MnO}_{3}$ memristor-based synapse models, each getting progressively closer to hardware realism, alongside four STDP learning rules in a two-layer SNN with LIF neurons and adaptive thresholds for five-class MNIST classification. On MNIST with small train set and large test set, our two-layer SNN with ideal, 25-state, and 12-state nonlinear memristor synapses achieves 92.73 %, 91.07 %, and 80 % accuracy, respectively, while converging faster and using fewer parameters than comparable ANN/CNN baselines.

cs.NE

Design Space and Variability Analysis of SOI MOSFET for Ultra-Low Power Band-to-Band Tunneling Neurons

Large spiking neural networks (SNNs) require ultra-low power and low variability hardware for neuromorphic computing applications. Recently, a band-to-band tunneling-based (BTBT) integrator, enabling sub-kHz operation of neurons with area and energy efficiency, was proposed. For an ultra-low power implementation of such neurons, a very low BTBT current is needed, so minimizing current without degrading neuronal properties is essential. Low variability is needed in the ultra-low current integrator to avoid network performance degradation in a large BTBT neuron-based SNN. To address this, we conducted design space and variability analysis in TCAD, utilizing a well-calibrated TCAD deck with experimental data from GlobalFoundries 32nm PD-SOI MOSFET. First, we discuss the physics-based explanation of the tunneling mechanism. Second, we explore the impact of device design parameters on SOI MOSFET performance, highlighting parameter sensitivities to tunneling current. With device parameters' optimization, we demonstrate a ~20x reduction in BTBT current compared to the experimental data. Finally, a variability analysis that includes the effects of random dopant fluctuations (RDF), oxide thickness variability (OTV), and channel-oxide interface traps DIT in the BTBT, SS, and ON regimes of operation is shown. The BTBT regime shows high sensitivity to the RDF and OTV as any variation in them directly modulates the tunnel length or the electric field at the drain-channel junction, whereas minimal sensitivity to DIT is observed.

physics.app-ph

FeFET-based MirrorBit cell for High-density NVM storage

HfO2-based Ferroelectric field-effect transistor (FeFET) has become a center of attraction for non-volatile memory applications because of their low power, fast switching speed, high scalability, and CMOS compatibility. In this work, we show an n-channel FeFET-based Multibit memory, termed MirrorBit, which effectively doubles the chip density via programming the gradient ferroelectric polarizations in the gate using an appropriate biasing scheme. We have experimentally demonstrated MirrorBit on GlobalFoundries HfO2-based FeFET devices fabricated at 28 nm bulk HKMG CMOS technology. Retention of MirrorBit states has been shown up to $10^5$ s at different temperatures. Also, the endurance is found to be more than $10^3$ cycles. A TCAD simulation is also presented to explain the origin and working of MirrorBit states based on the FeFET model calibrated using the GlobalFoundries FeFET device. We have also proposed the array-level implementation and sensing methodology of the MirrorBit memory. Thus, we have converted 1-bit FeFET into 2-bit FeFET using a particular programming scheme in existing FeFET, without needing any notable fabrication process alteration, to double the chip density for high-density non-volatile memory storage.

eess.SY

Ferroelectric MirrorBit-Integrated Field-Programmable Memory Array for TCAM, Storage, and In-Memory Computing Applications

In-memory computing on a reconfigurable architecture is the emerging field which performs an application-based resource allocation for computational efficiency and energy optimization. In this work, we propose a Ferroelectric MirrorBit-integrated field-programmable reconfigurable memory. We show the conventional 1-Bit FeFET, the MirrorBit, and MirrorBit-based Ternary Content-addressable memory (MCAM or MirrorBit-based TCAM) within the same field-programmable array. Apart from the conventional uniform Up and Down polarization states, the additional states in the MirrorBit are programmed by applying a non-uniform electric field along the transverse direction, which produces a gradient in the polarization and the conduction band energy. This creates two additional states, thereby, creating a total of 4 states or 2-bit of information. The gradient in the conduction band resembles a Schottky barrier (Schottky diode), whose orientation can be configured by applying an appropriate field. The TCAM operation is demonstrated using the MirrorBit-based diode on the reconfigurable array. The reconfigurable array architecture can switch from AND-type to NOR-type and vice-versa. The AND-type array is appropriate for programming the conventional bit and the MirrorBit. The MirrorBit-based Schottky diode in the NOR-array resembles a crossbar structure, which is appropriate for diode-based CAM operation. Our proposed memory system can enable fast write via 1-bit FeFET, the dense data storage capability by Mirror-bit technology and the fast search capability of the MCAM. Further, the dual configurability enables power, area and speed optimization making the reconfigurable Fe-Mirrorbit memory a compelling solution for In-memory and associative computing.

eess.SY

Process Voltage Temperature Variability Estimation of Tunneling Current for Band-to-Band-Tunneling based Neuron

Compact and energy-efficient Synapse and Neurons are essential to realize the full potential of neuromorphic computing. In addition, a low variability is indeed needed for neurons in Deep neural networks for higher accuracy. Further, process (P), voltage (V), and temperature (T) variation (PVT) are essential considerations for low-power circuits as performance impact and compensation complexities are added costs. Recently, band-to-band tunneling (BTBT) neuron has been demonstrated to operate successfully in a network to enable a Liquid State Machine. A comparison of the PVT with competing modes of operation (e.g., BTBT vs. sub-threshold and above threshold) of the same transistor is a critical factor in assessing performance. In this work, we demonstrate the PVT variation impact in the BTBT regime and benchmark the operation against the subthreshold slope (SS) and ON-regime (ION) of partially depleted-Silicon on Insulator MOSFET. It is shown that the On-state regime offers the lowest variability but dissipates higher power. Hence, not usable for low-power sources. Among the BTBT and SS regimes, which can enable the low-power neuron, the BTBT regime has shown ~3x variability reduction (σ_I_D/μ_I_D) than the SS regime, considering the cumulative PVT variability. The improvement is due to the well-known weaker P, V, and T dependence of BTBT vs. SS. We show that the BTBT variation is uncorrelated with mutually correlated SS & ION operation - indicating its different origin from the mechanism and location perspectives. Hence, the BTBT regime is promising for low-current, low-power, and low device-to-device variability neuron operation.

physics.app-ph

Interlayer-engineered local epitaxial templating induced enhancement in polarization (2P$_r$ > 70$μ$C/cm$^2$) in Hf$_{0.5}$Zr$_{0.5}$O$_2$ thin films

In this work, we report a high remnant polarization, 2Pr >70$μ$C/cm$^2$ in thermally processed atomic layer deposited Hf0.5Zr0.5O2 (HZO) film on Silicon with NH3 plasma exposed thin TiN interlayer and Tungsten (W) as a top electrode. The effect of interlayer on the ferroelectric properties of HZO is compared with standard Metal-Ferroelectric-Metal and Metal-Ferroelectric-Semiconductor structures. X-Ray Diffraction shows that the Orthorhombic (o) phase increases as TiN is thinned. However, the strain in the o-phase is highest at 2 nm TiN and then relaxes significantly for the no-TiN case. HRTEM images reveal that the ultra-thin TiN acts as a seed layer for the local epitaxy in HZO potentially increasing the strain to produce a 2X improvement in the remnant polarization. Finally, the HZO devices are shown to be wake-up-free, and exhibit endurance >10^6 cycles. This study opens a pathway to achieve epitaxial ferroelectric HZO films on Si with improved memory performance.

cond-mat.mtrl-sci

Evolution of ferroelectricity with annealing temperature and thickness in sputter deposited undoped HfO$_2$ on silicon

Ferroelectricity in sputtered undoped-HfO$_2$ is attractive for composition control for low power and non-volatile memory and logic applications. Unlike doped HfO$_2$, evolution of ferroelectricity with annealing and film thickness effect in sputter deposited undoped HfO$_2$ on Si is not yet reported. In present study, we have demonstrated the impact of post metallization annealing temperature and film thickness on ferroelectric properties in dopant-free sputtered HfO$_2$ on Si-substrate. A rich correlation of polarization with phase, lattice constant, and crystallite size and interface reaction is observed. First, anneal temperature shows o-phase saturation beyond 600 oC followed by interface reaction beyond 700 oC to show an optimal temperature window on 600-700 oC. Second, thickness study at the optimal temperature window shows an alluring o-phase crystallite scaling with thickness till a critical thickness of 20 nm indicating that the films are completely o-phase. However, the lattice constants (volume) are high in the 15-20 nm thickness range which correlates with the enhanced value of 2Pr. Beyond 20 nm, crystallite scaling with thickness saturates with the correlated appearance of m-phase and reduction in 2Pr. The optimal thickness-temperature window range of 15-20 nm films annealed at 600-700 oC show 2Pr of ~35.5 micro-C/cm$^2$ is comparable to state-of-the-art. The robust wakeup-free endurance of ~$10^$8 cycles showcased in the promising temperature-thickness window has been identified systematically for non-volatile memory applications.

cond-mat.mtrl-sci

Schottky Barrier MOSFET Enabled Ultra-Low Power Real-Time Neuron for Neuromorphic Computing

Energy-efficient real-time synapses and neurons are essential to enable large-scale neuromorphic computing. In this paper, we propose and demonstrate the Schottky-Barrier MOSFET-based ultra-low power voltage-controlled current source to enable real-time neurons for neuromorphic computing. Schottky-Barrier MOSFET is fabricated on a Silicon-on-insulator platform with polycrystalline Silicon as the channel and Nickel/Platinum as the source/drain. The Poly-Si and Nickel make the back-to-back Schottky junction enabling ultra-low ON current required for energy-efficient neurons.

cs.ET

India's Rise in Nanoelectronics Research

Modern semiconductors innovation has a strong relation to scale and skill. While India has a significant demand for semiconductors, it has a daunting challenge to create a semiconductor ecosystem. Yet, India has quietly come a long way. Starting with Centers of Excellence in Nanoelectronics (CENs) initiated in 2006 and broad science and technology funding, India has transformed its nanoelectronics research ecosystem. From negligible contributions as late as 2011, India has risen to be a top contributor to IEEE Electron Devices journals today. Our study presents important observations in terms of ecosystem development. First, there is a 6 year incubation time from infrastructure initiation to first papers. Then, 4 more years to become globally competitive. Second, growth in experimental research is essential along with modeling & simulations. Finally, the aspirational goals of translational research to contribute to the global technology roadmap requires cutting-edge manufacturing infrastructure & ecosystem access, which still needs development. The learning informs a call to action for the research ecosystem i.e. academia, industry, and policy-makers. First, sustain and amplify successful strategies of national research infrastructure & funding growth. Second, enhance international collaborations to add further scale & infrastructure to R&D. Finally, strengthen the industry-academia-policy consortium approach to transform to an innovation-based economy. Ultimately, the electron devices community is entering an exciting phase where Beyond Moore offers open opportunities in materials, devices to systems, and algorithms. India must build on its success to play a significant role in this new world of disruptive innovation.

eess.SY

Reaction-Drift Model for Switching Transients in Pr$_{0.7}$Ca$_{0.3}$MnO$_3$-Based Resistive RAM

Pr$_{0.7}$Ca$_{0.3}$MnO$_3$ (PCMO) based RRAM shows promising memory properties like non-volatility, low variability, multiple resistance states and scalability. From a modeling perspective, the charge carrier DC current modeling of PCMO RRAM by drift diffusion (DD) in the presence of fixed oxygen ion vacancy traps and self-heating (SH) in Technology Computer Aided Design (TCAD) (but without oxygen ionic transport) was able to explain the experimentally observed space charge limited conduction (SCLC) characteristics, prior to resistive switching. Further, transient analysis using DD+SH model was able to reproduce the experimentally observed fast current increase at ~100 ns timescale, prior to resistive switching. However, a complete quantitative transient current transport plus resistive switching model requires the inclusion of ionic transport. We propose a Reaction-Drift (RD) model for oxygen ion vacancy related trap density variation, which is combined with the DD+SH model. Earlier we have shown that the Set transient consists of 3 stages and Reset transient consists of 4 stages experimentally. In this work, the DD+SH+RD model is able to reproduce the entire transient behavior over 10 ns - 1 s range in timescale for both the Set and Reset operations for different applied biases and ambient temperatures. Remarkably, a universal Reset experimental behavior, log(I) is proportional to (m X log(t)) where m~-1/10 is reproduced in simulations. This model is the first model for PCMO RRAMs to significantly reproduce transient Set/Reset behavior. This model establishes the presence of self-heating and ionic-drift limited resistive switching as primary physical phenomena in these RRAMs.

physics.app-ph

Understanding the Location of Resistance Change in the Pr0.7Ca0.3MnO3 RRAM

Pr1-xCaxMnO3 (PCMO) based resistance random access memory (RRAM) is attractive in large scale memory and neuromorphic applications as it is non-filamentary, area scalable and has multiple resistance states along with excellent endurance and retention. The PCMO RRAM exhibit area scalable resistive switching when in contact with the reactive electrode. The interface redox reaction based resistance switching is observed electrically. Yet, whether resistance change occurs through partial (close to interface) or entire bulk is largely debated. Essentially, a two-terminal device is unable to provide direct evidence of the resistance change location in the PCMO RRAM. In this paper, we propose and experimentally demonstrate a novel three-terminal RRAM device in which a thin third terminal (~20nm) is inserted laterally in a typical vertical 2 terminal RRAM device of PCMO thickness of ~80nm. Using the 3T-RRAM method, we show that resistance change occurs largely at the upper bulk (near reactive electrode interface) - which is highly asymmetric. Yet it produces SCLC based resistance change with symmetric IV characteristics. It is the first time that an interface redox and bulk SCLC based resistance change has been experimentally shown as correlated and consistent - enabled by the 3rd terminal of the RRAM. Such a study enables a critical understanding of the device which enables the design and development of PCMO RRAM for memory and neuromorphic computing applications.

physics.app-ph

A case for multiple and parallel RRAMs as synaptic model for training SNNs

To enable a dense integration of model synapses in a spiking neural networks hardware, various nano-scale devices are being considered. Such a device, besides exhibiting spike-time dependent plasticity (STDP), needs to be highly scalable, have a large endurance and require low energy for transitioning between states. In this work, we first introduce and empirically determine two new specifications for an synapse in SNNs: number of conductance levels per synapse and maximum learning-rate. To the best of our knowledge, there are no RRAMs that meet the latter specification. As a solution, we propose the use of multiple PCMO-RRAMs in parallel within a synapse. While synaptic reading, all PCMO-RRAMs are simultaneously read and for each synaptic conductance-change event, the mechanism for conductance STDP is initiated for only one RRAM, randomly picked from the set. Second, to validate our solution, we experimentally demonstrate STDP of conductance of a PCMO-RRAM and then show that due to a large learning-rate, a single PCMO-RRAM fails to model a synapse in the training of an SNN. As anticipated, network training improves as more PCMO-RRAMs are added to the synapse. Fourth, we discuss the circuit-requirements for implementing such a scheme, to conclude that the requirements are within bounds. Thus, our work presents specifications for synaptic devices in trainable SNNs, indicates the shortcomings of state-of-art synaptic contenders, and provides a solution to extrinsically meet the specifications and discusses the peripheral circuitry that implements the solution.

cs.ET

A simple and efficient SNN and its performance & robustness evaluation method to enable hardware implementation

Spiking Neural Networks (SNN) are more closely related to brain-like computation and inspire hardware implementation. This is enabled by small networks that give high performance on standard classification problems. In literature, typical SNNs are deep and complex in terms of network structure, weight update rules and learning algorithms. This makes it difficult to translate them into hardware. In this paper, we first develop a simple 2-layered network in software which compares with the state of the art on four different standard data-sets within SNNs and has improved efficiency. For example, it uses lower number of neurons (3 x), synapses (3.5 x) and epochs for training (30 x) for the Fisher Iris classification problem. The efficient network is based on effective population coding and synapse-neuron co-design. Second, we develop a computationally efficient (15000 x) and accurate (correlation of 0.98) method to evaluate the performance of the network without standard recognition tests. Third, we show that the method produces a robustness metric that can be used to evaluate noise tolerance.

cs.NE