SearcharxivSearch

arXiv subjects

Udayan Ganguly

Publications and source records attributed to Udayan Ganguly.

At least 19 recordsLinked to original sources

Symbol Detection in a MIMO Wireless Communication System Using a FeFET-coupled CMOS Ring Oscillator Array

Symbol decoding in multiple-input multiple-output (MIMO) wireless communication systems requires the deployment of fast, energy-efficient computing hardware deployable at the edge. The brute-force, exact maximum likelihood (ML) decoder, solved on conventional classical digital hardware, has exponential time complexity. Approximate classical solvers implemented on the same hardware have polynomial time complexity at the best. In this article, we design an alternative ring-oscillator-based coupled oscillator array to act as an oscillator Ising machine (OIM) and heuristically solve the ML-based MIMO detection problem. Complementary metal oxide semiconductor (CMOS) technology is used to design the ring oscillators, and ferroelectric field effect transistor (FeFET) technology is chosen as the coupling element (X) between the oscillators in this CMOS + X OIM design. For this purpose, we experimentally report high linear range of conductance variation (1 micro-S to 60 micro-S) in a FeFET device fabricated at 28 nm high-K/ metal gate (HKMG) CMOS technology node. We incorporate the conductance modulation characteristic in SPICE simulation of the ring oscillators connected in an all-to-all fashion through a crossbar array of these FeFET devices. We show that the above range of conductance variation of the FeFET device is suitable to obtain optimum OIM performance with no significant performance drop up to a MIMO size of 100 transmitting and 100 receiving antennas, thereby making FeFET a suitable device for this application. Our simulations and associated analysis using the Kuramoto model of oscillators also predict that this designed classical analog OIM, if implemented experimentally, will offer logarithmic scaling of computation time with MIMO size, thereby offering a huge improvement (in terms of computation speed) over aforementioned MIMO decoders run on conventional digital hardware.

physics.app-ph

Regularization-based Framework for Quantization-, Fault- and Variability-Aware Training

Efficient inference is critical for deploying deep learning models on edge AI devices. Low-bit quantization (e.g., 3- and 4-bit) with fixed-point arithmetic improves efficiency, while low-power memory technologies like analog nonvolatile memory enable further gains. However, these methods introduce non-ideal hardware behavior, including bit faults and device-to-device variability. We propose a regularization-based quantization-aware training (QAT) framework that supports fixed, learnable step-size, and learnable non-uniform quantization, achieving competitive results on CIFAR-10 and ImageNet. Our method also extends to Spiking Neural Networks (SNNs), demonstrating strong performance on 4-bit networks on CIFAR10-DVS and N-Caltech 101. Beyond quantization, our framework enables fault and variability-aware fine-tuning, mitigating stuck-at faults (fixed weight bits) and device resistance variability. Compared to prior fault-aware training, our approach significantly improves performance recovery under upto 20% bit-fault rate and 40% device-to-device variability. Our results establish a generalizable framework for quantization and robustness-aware training, enhancing efficiency and reliability in low-power, non-ideal hardware.

cs.LG

Heracles: A HfO2 Ferroelectric Capacitor Compact Model for Efficient Circuit Simulations

The growing use of ferroelectric-based technology, extending beyond conventional memory storage applications, necessitates the development of compact models that can be easily integrated into circuit simulation environments. These models assist circuit designers in the design and the early assessment of the performance of their systems. The Heracles model is a physics-based compact model for circuit simulations in a SPICE environment for HfO2-based ferroelectric capacitors (FeCaps). The model has been calibrated based on experimental data obtained from HfO2-based FeCaps. A thermal model with an accurate description of the device parasitics is included to derive precise device characteristics based on first principles. The incorporation of statistical device data enables Monte Carlo analysis based on realistic distributions, thereby rendering the model particularly well-suited for design-technology co-optimization (DTCO). The model's efficacy is further demonstrated in circuit simulations using an integrated circuit with current programming, wherein partial switching of the ferroelectric polarization is observed. Finally, the model was benchmarked in an array simulation, reaching convergence in 1.8 s with an array size of 100 kb.

cs.ET

Temporal and Spatial Reservoir Ensembling Techniques for Liquid State Machines

Reservoir computing (RC), is a class of computational methods such as Echo State Networks (ESN) and Liquid State Machines (LSM) describe a generic method to perform pattern recognition and temporal analysis with any non-linear system. This is enabled by Reservoir Computing being a shallow network model with only Input, Reservoir, and Readout layers where input and reservoir weights are not learned (only the readout layer is trained). LSM is a special case of Reservoir computing inspired by the organization of neurons in the brain and generally refers to spike-based Reservoir computing approaches. LSMs have been successfully used to showcase decent performance on some neuromorphic vision and speech datasets but a common problem associated with LSMs is that since the model is more-or-less fixed, the main way to improve the performance is by scaling up the Reservoir size, but that only gives diminishing rewards despite a tremendous increase in model size and computation. In this paper, we propose two approaches for effectively ensembling LSM models - Multi-Length Scale Reservoir Ensemble (MuLRE) and Temporal Excitation Partitioned Reservoir Ensemble (TEPRE) and benchmark them on Neuromorphic-MNIST (N-MNIST), Spiking Heidelberg Digits (SHD), and DVSGesture datasets, which are standard neuromorphic benchmarks. We achieve 98.1% test accuracy on N-MNIST with a 3600-neuron LSM model which is higher than any prior LSM-based approach and 77.8% test accuracy on the SHD dataset which is on par with a standard Recurrent Spiking Neural Network trained by Backprop Through Time (BPTT). We also propose receptive field-based input weights to the Reservoir to work alongside the Multi-Length Scale Reservoir ensemble model for vision tasks. Thus, we introduce effective means of scaling up the performance of LSM models and evaluate them against relevant neuromorphic benchmarks

cs.LG

Design Space and Variability Analysis of SOI MOSFET for Ultra-Low Power Band-to-Band Tunneling Neurons

Large spiking neural networks (SNNs) require ultra-low power and low variability hardware for neuromorphic computing applications. Recently, a band-to-band tunneling-based (BTBT) integrator, enabling sub-kHz operation of neurons with area and energy efficiency, was proposed. For an ultra-low power implementation of such neurons, a very low BTBT current is needed, so minimizing current without degrading neuronal properties is essential. Low variability is needed in the ultra-low current integrator to avoid network performance degradation in a large BTBT neuron-based SNN. To address this, we conducted design space and variability analysis in TCAD, utilizing a well-calibrated TCAD deck with experimental data from GlobalFoundries 32nm PD-SOI MOSFET. First, we discuss the physics-based explanation of the tunneling mechanism. Second, we explore the impact of device design parameters on SOI MOSFET performance, highlighting parameter sensitivities to tunneling current. With device parameters' optimization, we demonstrate a ~20x reduction in BTBT current compared to the experimental data. Finally, a variability analysis that includes the effects of random dopant fluctuations (RDF), oxide thickness variability (OTV), and channel-oxide interface traps DIT in the BTBT, SS, and ON regimes of operation is shown. The BTBT regime shows high sensitivity to the RDF and OTV as any variation in them directly modulates the tunnel length or the electric field at the drain-channel junction, whereas minimal sensitivity to DIT is observed.

physics.app-ph

FeFET-based MirrorBit cell for High-density NVM storage

HfO2-based Ferroelectric field-effect transistor (FeFET) has become a center of attraction for non-volatile memory applications because of their low power, fast switching speed, high scalability, and CMOS compatibility. In this work, we show an n-channel FeFET-based Multibit memory, termed MirrorBit, which effectively doubles the chip density via programming the gradient ferroelectric polarizations in the gate using an appropriate biasing scheme. We have experimentally demonstrated MirrorBit on GlobalFoundries HfO2-based FeFET devices fabricated at 28 nm bulk HKMG CMOS technology. Retention of MirrorBit states has been shown up to $10^5$ s at different temperatures. Also, the endurance is found to be more than $10^3$ cycles. A TCAD simulation is also presented to explain the origin and working of MirrorBit states based on the FeFET model calibrated using the GlobalFoundries FeFET device. We have also proposed the array-level implementation and sensing methodology of the MirrorBit memory. Thus, we have converted 1-bit FeFET into 2-bit FeFET using a particular programming scheme in existing FeFET, without needing any notable fabrication process alteration, to double the chip density for high-density non-volatile memory storage.

eess.SY

A temporally and spatially local spike-based backpropagation algorithm to enable training in hardware

Spiking Neural Networks (SNNs) have emerged as a hardware efficient architecture for classification tasks. The challenge of spike-based encoding has been the lack of a universal training mechanism performed entirely using spikes. There have been several attempts to adopt the powerful backpropagation (BP) technique used in non-spiking artificial neural networks (ANN): (1) SNNs can be trained by externally computed numerical gradients. (2) A major advancement towards native spike-based learning has been the use of approximate Backpropagation using spike-time dependent plasticity (STDP) with phased forward/backward passes. However, the transfer of information between such phases for gradient and weight update calculation necessitates external memory and computational access. This is a challenge for standard neuromorphic hardware implementations. In this paper, we propose a stochastic SNN based Back-Prop (SSNN-BP) algorithm that utilizes a composite neuron to simultaneously compute the forward pass activations and backward pass gradients explicitly with spikes. Although signed gradient values are a challenge for spike-based representation, we tackle this by splitting the gradient signal into positive and negative streams. We show that our method approaches BP ANN baseline with sufficiently long spike-trains. Finally, we show that the well-performing softmax cross-entropy loss function can be implemented through inhibitory lateral connections enforcing a Winner Take All (WTA) rule. Our SNN with a 2-layer network shows excellent generalization through comparable performance to ANNs with equivalent architecture and regularization parameters on static image datasets like MNIST, Fashion-MNIST, Extended MNIST, and temporally encoded image datasets like Neuromorphic MNIST datasets. Thus, SSNN-BP enables BP compatible with purely spike-based neuromorphic hardware.

cs.NE

Non-Ideal Program-Time Conservation in Charge Trap Flash for Deep Learning

Training deep neural networks (DNNs) is computationally intensive but arrays of non-volatile memories like Charge Trap Flash (CTF) can accelerate DNN operations using in-memory computing. Specifically, the Resistive Processing Unit (RPU) architecture uses the voltage-threshold program by stochastic encoded pulse trains and analog memory features to accelerate vector-vector outer product and weight update for the gradient descent algorithms. Although CTF, offering high precision, has been regarded as an excellent choice for implementing RPU, the accumulation of charge due to the applied stochastic pulse trains is ultimately of critical significance in determining the final weight update. In this paper, we report the non-ideal program-time conservation in CTF through pulsing input measurements. We experimentally measure the effect of pulse width and pulse gap, keeping the total ON-time of the input pulse train constant, and report three non-idealities: (1) Cumulative V_T shift reduces when total ON-time is fragmented into a larger number of shorter pulses, (2) Cumulative V_T shift drops abruptly for pulse widths < 2 μs, (3) Cumulative V_T shift depends on the gap between consecutive pulses and the V_T shift reduction gets recovered for smaller gaps. We present an explanation based on a transient tunneling field enhancement due to blocking oxide trap-charge dynamics to explain these non-idealities. Identifying and modeling the responsible mechanisms and predicting their system-level effects during learning is critical. This non-ideal accumulation is expected to affect algorithms and architectures relying on devices for implementing mathematically equivalent functions for in-memory computing-based acceleration.

cs.NE

Ferroelectric MirrorBit-Integrated Field-Programmable Memory Array for TCAM, Storage, and In-Memory Computing Applications

In-memory computing on a reconfigurable architecture is the emerging field which performs an application-based resource allocation for computational efficiency and energy optimization. In this work, we propose a Ferroelectric MirrorBit-integrated field-programmable reconfigurable memory. We show the conventional 1-Bit FeFET, the MirrorBit, and MirrorBit-based Ternary Content-addressable memory (MCAM or MirrorBit-based TCAM) within the same field-programmable array. Apart from the conventional uniform Up and Down polarization states, the additional states in the MirrorBit are programmed by applying a non-uniform electric field along the transverse direction, which produces a gradient in the polarization and the conduction band energy. This creates two additional states, thereby, creating a total of 4 states or 2-bit of information. The gradient in the conduction band resembles a Schottky barrier (Schottky diode), whose orientation can be configured by applying an appropriate field. The TCAM operation is demonstrated using the MirrorBit-based diode on the reconfigurable array. The reconfigurable array architecture can switch from AND-type to NOR-type and vice-versa. The AND-type array is appropriate for programming the conventional bit and the MirrorBit. The MirrorBit-based Schottky diode in the NOR-array resembles a crossbar structure, which is appropriate for diode-based CAM operation. Our proposed memory system can enable fast write via 1-bit FeFET, the dense data storage capability by Mirror-bit technology and the fast search capability of the MCAM. Further, the dual configurability enables power, area and speed optimization making the reconfigurable Fe-Mirrorbit memory a compelling solution for In-memory and associative computing.

eess.SY

Process Voltage Temperature Variability Estimation of Tunneling Current for Band-to-Band-Tunneling based Neuron

Compact and energy-efficient Synapse and Neurons are essential to realize the full potential of neuromorphic computing. In addition, a low variability is indeed needed for neurons in Deep neural networks for higher accuracy. Further, process (P), voltage (V), and temperature (T) variation (PVT) are essential considerations for low-power circuits as performance impact and compensation complexities are added costs. Recently, band-to-band tunneling (BTBT) neuron has been demonstrated to operate successfully in a network to enable a Liquid State Machine. A comparison of the PVT with competing modes of operation (e.g., BTBT vs. sub-threshold and above threshold) of the same transistor is a critical factor in assessing performance. In this work, we demonstrate the PVT variation impact in the BTBT regime and benchmark the operation against the subthreshold slope (SS) and ON-regime (ION) of partially depleted-Silicon on Insulator MOSFET. It is shown that the On-state regime offers the lowest variability but dissipates higher power. Hence, not usable for low-power sources. Among the BTBT and SS regimes, which can enable the low-power neuron, the BTBT regime has shown ~3x variability reduction (σ_I_D/μ_I_D) than the SS regime, considering the cumulative PVT variability. The improvement is due to the well-known weaker P, V, and T dependence of BTBT vs. SS. We show that the BTBT variation is uncorrelated with mutually correlated SS & ION operation - indicating its different origin from the mechanism and location perspectives. Hence, the BTBT regime is promising for low-current, low-power, and low device-to-device variability neuron operation.

physics.app-ph

Interlayer-engineered local epitaxial templating induced enhancement in polarization (2P$_r$ > 70$μ$C/cm$^2$) in Hf$_{0.5}$Zr$_{0.5}$O$_2$ thin films

In this work, we report a high remnant polarization, 2Pr >70$μ$C/cm$^2$ in thermally processed atomic layer deposited Hf0.5Zr0.5O2 (HZO) film on Silicon with NH3 plasma exposed thin TiN interlayer and Tungsten (W) as a top electrode. The effect of interlayer on the ferroelectric properties of HZO is compared with standard Metal-Ferroelectric-Metal and Metal-Ferroelectric-Semiconductor structures. X-Ray Diffraction shows that the Orthorhombic (o) phase increases as TiN is thinned. However, the strain in the o-phase is highest at 2 nm TiN and then relaxes significantly for the no-TiN case. HRTEM images reveal that the ultra-thin TiN acts as a seed layer for the local epitaxy in HZO potentially increasing the strain to produce a 2X improvement in the remnant polarization. Finally, the HZO devices are shown to be wake-up-free, and exhibit endurance >10^6 cycles. This study opens a pathway to achieve epitaxial ferroelectric HZO films on Si with improved memory performance.

cond-mat.mtrl-sci

Evolution of ferroelectricity with annealing temperature and thickness in sputter deposited undoped HfO$_2$ on silicon

Ferroelectricity in sputtered undoped-HfO$_2$ is attractive for composition control for low power and non-volatile memory and logic applications. Unlike doped HfO$_2$, evolution of ferroelectricity with annealing and film thickness effect in sputter deposited undoped HfO$_2$ on Si is not yet reported. In present study, we have demonstrated the impact of post metallization annealing temperature and film thickness on ferroelectric properties in dopant-free sputtered HfO$_2$ on Si-substrate. A rich correlation of polarization with phase, lattice constant, and crystallite size and interface reaction is observed. First, anneal temperature shows o-phase saturation beyond 600 oC followed by interface reaction beyond 700 oC to show an optimal temperature window on 600-700 oC. Second, thickness study at the optimal temperature window shows an alluring o-phase crystallite scaling with thickness till a critical thickness of 20 nm indicating that the films are completely o-phase. However, the lattice constants (volume) are high in the 15-20 nm thickness range which correlates with the enhanced value of 2Pr. Beyond 20 nm, crystallite scaling with thickness saturates with the correlated appearance of m-phase and reduction in 2Pr. The optimal thickness-temperature window range of 15-20 nm films annealed at 600-700 oC show 2Pr of ~35.5 micro-C/cm$^2$ is comparable to state-of-the-art. The robust wakeup-free endurance of ~$10^$8 cycles showcased in the promising temperature-thickness window has been identified systematically for non-volatile memory applications.

cond-mat.mtrl-sci

Schottky Barrier MOSFET Enabled Ultra-Low Power Real-Time Neuron for Neuromorphic Computing

Energy-efficient real-time synapses and neurons are essential to enable large-scale neuromorphic computing. In this paper, we propose and demonstrate the Schottky-Barrier MOSFET-based ultra-low power voltage-controlled current source to enable real-time neurons for neuromorphic computing. Schottky-Barrier MOSFET is fabricated on a Silicon-on-insulator platform with polycrystalline Silicon as the channel and Nickel/Platinum as the source/drain. The Poly-Si and Nickel make the back-to-back Schottky junction enabling ultra-low ON current required for energy-efficient neurons.

cs.ET

Stochasticity Invariance Control in Pr$_{1-x}$Ca$_x$MnO$_3$ RRAM to enable Large-Scale Stochastic Recurrent Neural Networks

Emerging non-volatile memories have been proposed for a wide range of applications from easing the von-Neumann bottleneck to neuromorphic applications. Specifically, scalable RRAMs based on Pr$_{1-x}$Ca$_x$MnO$_3$ (PCMO) exhibit analog switching have been demonstrated as an integrating neuron, an analog synapse, and a voltage-controlled oscillator. More recently, the inherent stochasticity of memristors has been proposed for efficient hardware implementations of Boltzmann Machines. However, as the problem size scales, the number of neurons increase and controlling the stochastic distribution tightly over many iterations is necessary. This requires parametric control over stochasticity. Here, we characterize the stochastic Set in PCMO RRAMs. We identify that the Set time distribution depends on the internal state of the device (i.e., resistance) in addition to external input (i.e., voltage pulse). This requires the confluence of contradictory properties like stochastic switching as well as deterministic state control in the same device. Unlike, "stochastic-everywhere" filamentary memristors, in PCMO RRAMs, we leverage the (i) stochastic Set in negative polarity and (ii) deterministic analog Reset in positive polarity to demonstrate 100x reduced Set time distribution drift. The impact on Boltzmann Machines' performance is analyzed and as opposed to the "fixed external input stochasticity", the "state-monitored stochasticity" can solve problems 20x larger in size. State monitoring also tunes out the device-to-device variability effect on distributions providing 10x better performance. In addition to the physical insights, this study establishes the use of experimental stochasticity in PCMO RRAMs in stochastic recurrent neural networks reliably over many iterations.

cs.ET

An Accurate Process Induced Variability Aware Compact Model-based Circuit Performance Estimation for Design-Technology Co-optimization

In sub-10nm FinFETs, Line-edge-roughness (LER) and metal-gate granularity (MGG) are the two most dominant sources of variability and are mostly modeled semi-empirically. In this work, compact models of LER and MGG are used. We show an accurate process-induced variability (PIV) aware compact model-based circuit performance estimation for Design-Technology Co-optimization (DTCO). This work is carried out using an experimentally validated BSIM-CMG model on a 7nm FinFET node. First, we have shown performance bench-marking of LER and MGG models with the state-of-the-art and shown {\textbackslash}4x({\textbackslash}2.3x) accuracy improvement for NMOS(PMOS) in the estimation of device figure of merits (DFoMs). Second, RO and SRAM circuits performance estimation is carried out for LER and MGG variability. Further, {\textbackslash}22\% more optimistic estimate of (σ/μ)\textsubscript{SHM} (Static Hold Margin) compared to the state-of-the-art model with V\textsubscript{DD} variation is shown. Finally, we demonstrate our improved DFoMs accuracy translated to more accurate circuits figure of merits (CFoMs) performance estimation. For worst-case SHM (3(σ/μ)\textsubscript{SHM}@VDD=0.75 V) compared to state-of-the-art, dynamic(standby) power reduction by {\textbackslash}73\%({\textbackslash}61\%) is shown. Thus, our enhanced variability model accuracy enables more credible DTCO with significantly better performance estimates.

physics.app-ph

Exploiting the Electrothermal Timescale in PrMnO3 RRAM for a compact, clock-less neuron exhibiting biological spiking patterns

Spiking Neural Networks (SNNs) are gaining widespread momentum in the field of neuromorphic computing. These network systems integrated with neurons and synapses provide computational efficiency by mimicking the human brain. It is desired to incorporate the biological neuronal dynamics, including complex spiking patterns which represent diverse brain activities within the neural networks. Earlier hardware realization of neurons was (1) area intensive because of large capacitors in the circuit design, (2) neuronal spiking patterns were demonstrated with clocked neurons at the device level. To achieve more realistic biological neuron spiking behavior, emerging memristive devices are considered promising alternatives. In this paper, we propose, PrMnO3(PMO) -RRAM device-based neuron. The voltage-controlled electrothermal timescales of the compact PMO RRAM device replace the electrical timescales of charging a large capacitor. The electrothermal timescale is used to implement an integration block with multiple voltage-controlled timescales coupled with a refractory block to generate biological neuronal dynamics. Here, first, a Verilog-A implementation of the thermal device model is demonstrated, which captures the current-temperature dynamics of the PMO device. Second, a driving circuitry is designed to mimic different spiking patterns of cortical neurons, including Intrinsic bursting (IB) and Chattering (CH). Third, a neuron circuit model is simulated, which includes the PMO RRAM device model and the driving circuitry to demonstrate the asynchronous neuron behavior. Finally, a hardware-software hybrid analysis is done in which the PMO RRAM device is experimentally characterized to mimic neuron spiking dynamics. The work presents a realizable and more biologically comparable hardware-efficient solution for large-scale SNNs.

cs.ET

Algorithm For 3D-Chemotaxis Using Spiking Neural Network

In this work, we aim to devise an end-to-end spiking implementation for contour tracking in 3D media inspired by chemotaxis, where the worm reaches the region which has the given set concentration. For a planer medium, efficient contour tracking algorithms have already been devised, but a new degree of freedom has quite a few challenges. Here we devise an algorithm based on klinokinesis - where the motion of the worm is in response to the stimuli but not proportional to it. Thus the path followed is not the shortest, but we can track the set concentration successfully. We are using simple LIF neurons for the neural network implementation, considering the feasibility of its implementation in the neuromorphic computing hardware.

cs.NE

Spiking-GAN: A Spiking Generative Adversarial Network Using Time-To-First-Spike Coding

Spiking Neural Networks (SNNs) have shown great potential in solving deep learning problems in an energy-efficient manner. However, they are still limited to simple classification tasks. In this paper, we propose Spiking-GAN, the first spike-based Generative Adversarial Network (GAN). It employs a kind of temporal coding scheme called time-to-first-spike coding. We train it using approximate backpropagation in the temporal domain. We use simple integrate-and-fire (IF) neurons with very high refractory period for our network which ensures a maximum of one spike per neuron. This makes the model much sparser than a spike rate-based system. Our modified temporal loss function called 'Aggressive TTFS' improves the inference time of the network by over 33% and reduces the number of spikes in the network by more than 11% compared to previous works. Our experiments show that on training the network on the MNIST dataset using this approach, we can generate high quality samples. Thereby demonstrating the potential of this framework for solving such problems in the spiking domain.

cs.NE