SearcharxivSearch

arXiv subjects

Mrigank Sharad

Publications and source records attributed to Mrigank Sharad.

At least 19 recordsLinked to original sources

Approximate ADCs for In-Memory Computing

In memory computing (IMC) architectures for deep learning (DL) accelerators leverage energy-efficient and highly parallel matrix vector multiplication (MVM) operations, implemented directly in memory arrays. Such IMC designs have been explored based on CMOS as well as emerging non-volatile memory (NVM) technologies like RRAM. IMC architectures generally involve a large number of cores consisting of memory arrays, storing the trained weights of the DL model. Peripheral units like DACs and ADCs are also used for applying inputs and reading out the output values. Recently reported designs reveal that the ADCs required for reading out the MVM results, consume more than 85% of the total compute power and also dominate the area, thereby eschewing the benefits of the IMC scheme. Mitigation of imperfections in the ADCs, namely, non-linearity and variations, incur significant design overheads, due to dedicated calibration units. In this work we present peripheral aware design of IMC cores, to mitigate such overheads. It involves incorporating the non-idealities of ADCs in the training of the DL models, along with that of the memory units. The proposed approach applies equally well to both current mode as well as charge mode MVM operations demonstrated in recent years., and can significantly simplify the design of mixed-signal IMC units.

cs.ET

Acoustic Scene Analysis using Analog Spiking Neural Network

Sensor nodes in a wireless sensor network (WSN) for security surveillance applications should preferably be small, energy-efficient, and inexpensive with in-sensor computational abilities. An appropriate data processing scheme in the sensor node reduces the power dissipation of the transceiver through the compression of information to be communicated. This study attempted a simulation-based analysis of human footstep sound classification in natural surroundings using simple time-domain features. The spiking neural network (SNN), a computationally low-weight classifier derived from an artificial neural network (ANN), was used to classify acoustic sounds. The SNN and required feature extraction schemes are amenable to low-power subthreshold analog implementation. The results show that all analog implementations of the proposed SNN scheme achieve significant power savings over the digital implementation of the same computing scheme and other conventional digital architectures using frequency-domain feature extraction and ANN-based classification. The algorithm is tolerant of the impact of process variations, which are inevitable in analog design, owing to the approximate nature of the data processing involved in such applications. Although SNN provides low-power operation at the algorithm level itself, ANN to SNN conversion leads to an unavoidable loss of classification accuracy of ~5%. We exploited the low-power operation of the analog processing SNN module by applying redundancy and majority voting, which improved the classification accuracy, taking it close to the ANN model.

cs.NE

Adaptive Multi-bit SRAM Topology Based Analog PUF

Physically Unclonable Functions (PUFs) are lightweight cryptographic primitives for generating unique signatures from minuscule manufacturing variations. In this work, we present lightweight, area efficient and low power adaptive multi-bit SRAM topology based Current Mirror Array (CMA) analog PUF design for securing the sensor nodes, authentication and key generation. The proposed Strong PUF increases the complexity of the machine learning attacks thus making it difficult for the adversary. The design is based on scl180 library.

cs.AR

Current Mode Neuron for the Memristor based synapse

Due to many limitations of Von Neumann architecture such as speed, memory bandwidth, efficiency of global interconnects and increase in the application of artificial neural network, researchers have been pushed to look into alternative architectures such as Neuromorphic computing system. Memristors (memristive crossbar memory RCM) are used as synapses due to its high packing density and energy efficiency and CMOS blocks as neurons. The increase in the terminal resistance of the RCM can degrade its energy efficiency and bandwidth. A more energy efficient current mode neuron has been proposed in this paper which can operate at lower voltages as compared to conventional voltage mode neuron circuit.

cs.ET

Mixer-First Receiver with wide-RF Range

In the Passive Mixer first receiver, four mosfet, and four baseband impedance can synthesize High-Q bandpass filter. This on-chip High-Q bandpass filter can replace bulky, expensive and off-chip SAW filters. The impedance which is seen by the antenna is tuned by switch resistance of the Passive Mixer and the input impedance of the Trans-impedance amplifier (baseband impedance) to match antenna impedance. The gain of the feedback loop of the TIA controls the input impedance of the TIA. Further, M baseband impedance can be replaced by M/4 complex impedance. This complex impedance provides a wide RF range by changing the transconductance.

eess.SP

Power efficient Spiking Neural Network Classifier based on memristive crossbar network for spike sorting application

In this paper authors have presented a power efficient scheme for implementing a spike sorting module. Spike sorting is an important application in the field of neural signal acquisition for implantable biomedical systems whose function is to map the Neural-spikes (N-spikes) correctly to the neurons from which it originates. The accurate classification is a pre-requisite for the succeeding systems needed in Brain-Machine-Interfaces (BMIs) to give better performance. The primary design constraint to be satisfied for the spike sorter module is low power with good accuracy. There lies a trade-off in terms of power consumption between the on-chip and off-chip training of the N-spike features. In the former case care has to be taken to make the computational units power efficient whereas in the later the data rate of wireless transmission should be minimized to reduce the power consumption due to the transceivers. In this work a 2-step shared training scheme involving a K-means sorter and a Spiking Neural Network (SNN) is elaborated for on-chip training and classification. Also, a low power SNN classifier scheme using memristive crossbar type architecture is compared with a fully digital implementation. The advantage of the former classifier is that it is power efficient while providing comparable accuracy as that of the digital implementation due to the robustness of the SNN training algorithm which has a good tolerance for variation in memristance.

cs.NE

Single Chip Self-Tunable N-Input N-Output PID Control System with Integrated Analog Front-end for Miniature Robotics

In this work, we explore the design of an integrated, low power single chip multi-channel Proportional-Integral-Derivative (PID) controller for emerging miniature robotics, that includes N inputs and N corresponding outputs thereby resulting in N parallel channels in the control system. It includes analog front-end (AFE) and analog PID controllers for PID parameter tuning based on PSO algorithm. The AFE incorporates adaptive biasing to ensure low power. The PSO is optimized with respect to tuning precision, power and area. This makes it attractive for real-time tuning of multiple miniaturized robotic devices with a single PSO tuning algorithm block assigned for the task. For simulation and testing purposes, we take N as 3 with the channels being defined by their application-ends or plants, namely: dc motor, temperature sensor and gyroscope.

eess.SY

Energy Efficient and High Performance Current-Mode Neural Network Circuit using Memristors and Digitally Assisted Analog CMOS Neurons

Emerging nano-scale programmable Resistive-RAM (RRAM) has been identified as a promising technology for implementing brain-inspired computing hardware. Several neural network architectures, that essentially involve computation of scalar products between input data vectors and stored network weights can be efficiently implemented using high density cross-bar arrays of RRAM integrated with CMOS. In such a design, the CMOS interface may be responsible for providing input excitations and for processing the RRAM output. In order to achieve high energy efficiency along with high integration density in RRAM based neuromorphic hardware, the design of RRAM-CMOS interface can therefore play a major role. In this work we propose design of high performance, current mode CMOS interface for RRAM based neural network design. The use of current mode excitation for input interface and design of digitally assisted current-mode CMOS neuron circuit for the output interface is presented. The proposed technique achieve 10x energy as well as performance improvement over conventional approaches employed in literature. Network level simulations show that the proposed scheme can achieve 2 orders of magnitude lower energy dissipation as compared to a digital ASIC implementation of a feed-forward neural network.

cs.ET

High Sensitivity Biosensor using Injection Locked Spin Torque Nano-Oscillators

With ever increasing research on magnetic nano systems it is shown to have great potential in the areas of magnetic storage, biosensing, magnetoresistive insulation etc. In the field of biosensing specifically Spin Valve sensors coupled with Magnetic Nanolabels is showing great promise due to noise immunity and energy efficiency [1]. In this paper we present the application of injection locked based Spin Torque Nano Oscillator (STNO) suitable for high resolution energy efficient labeled DNA Detection. The proposed STNO microarray consists of 20 such devices oscillating at different frequencies making it possible to multiplex all the signals using capacitive coupling. Frequency Division Multiplexing can be aided with Time division multiplexing to increase the device integration and decrease the readout time while maintaining the same efficiency in presence of constant input referred noise.

cs.ET

Digital LDO with Time-Interleaved Comparators for Fast Response and Low Ripple

On-chip voltage regulation using distributed Digital Low Drop Out (LDO) voltage regulators has been identified as a promising technique for efficient power-management for emerging multi-core processors. Digital LDOs (DLDO) can offer low voltage operation, faster transient response, and higher current efficiency. Response time as well as output voltage ripple can be reduced by increasing the speed of the dynamic comparators. However, the comparator offset steeply increases for high clock frequencies, thereby leading to enhanced variations in output voltage. In this work we explore the design of digital LDOs with multiple dynamic comparators that can overcome this bottleneck. In the proposed topology, we apply time-interleaved comparators with the same voltage threshold and uniform current step in order to accomplish the aforementioned features. Simulation based analysis shows that the DLDO with time-interleaved comparators can achieve better overall performance in terms of current efficiency, ripple and settling time. For a load step of 50mA, a DLDO with 8 time-interleaved comparators could achieve an output ripple of less than 5mV, while achieving a settling time of less than 0.5us. Load current dependant dynamic adjustment of clock frequency is proposed to maintain high current efficiency of ~97%.

cs.AR

Design and Synthesis of Ultra Low Energy Spin-Memristor Threshold Logic

A threshold logic gate (TLG) performs weighted sum of multiple inputs and compares the sum with a threshold. We propose Spin-Memeristor Threshold Logic (SMTL) gates, which employ memristive cross-bar array (MCA) to perform current-mode summation of binary inputs, whereas, the low-voltage fast-switching spintronic threshold devices (STD) carry out the threshold operation in an energy efficient manner. Field programmable SMTL gate arrays can operate at a small terminal voltage of ~50mV, resulting in ultra-low power consumption in gates as well as programmable interconnect networks. We evaluate the performance of SMTL using threshold logic synthesis. Results for common benchmarks show that SMTL based programmable logic hardware can be more than 100x energy efficient than state of the art CMOS FPGA.

cs.ET

Hierarchical Temporal Memory Based on Spin-Neurons and Resistive Memory for Energy-Efficient Brain-Inspired Computing

Hierarchical temporal memory (HTM) tries to mimic the computing in cerebral-neocortex. It identifies spatial and temporal patterns in the input for making inferences. This may require large number of computationally expensive tasks like, dot-product evaluations. Nano-devices that can provide direct mapping for such primitives are of great interest. In this work we show that the computing blocks for HTM can be mapped using low-voltage, fast-switching, magneto-metallic spin-neurons combined with emerging resistive cross-bar network (RCN). Results show possibility of more than 200x lower energy as compared to 45nm CMOS ASIC design

cs.ET

Exploring Ultra Low-Power on-Chip Clocking Using Functionality Enhanced Spin-Torque Switches

Emerging spin-torque (ST) phenomena may lead to ultra-low-voltage, high-speed nano-magnetic switches. Such current-based-switches can be attractive for designing low swing global-interconnects, like, clocking-networks and databuses. In this work we present the basic idea of using such ST-switches for low-power on-chip clocking. For clockingnetworks, Spin-Hall-Effect (SHE) can be used to produce an assist-field for fast ST-switching using global-mesh-clock with less than 100mV swing. The ST-switch acts as a compact-latch, written by ultra-low-voltage input-pulses. The data is read using a high-resistance tunnel-junction. The clock-driven SHE write-assist can be shared among large number of ST-latches, thereby reducing the load-capacitance for clock-distribution. The SHE assist can be activated by a low-swing clock (~150mV) and hence can facilitate ultra-low voltage clock-distribution. Owing to reduced clock-load and low-voltage operation, the proposed scheme can achieve 97% low-power for on-chip clocking as compared to the state of the art CMOS design. Rigorous device-circuit simulations and system-level modelling for the proposed scheme will be addressed in future.

cond-mat.mes-hall

Energy-Efficient and Robust Associative Computing with Electrically Coupled Dual Pillar Spin-Torque Oscillators

Dynamics of coupled spin-torque oscillators can be exploited for non-Boolean information processing. However, the feasibility of coupling large number of STOs with energy-efficiency and sufficient robustness towards parameter-variation and thermal-noise, may be critical for such computing applications. In this work, the impacts of parameter-variation and thermal-noise on two different coupling mechanisms for STOs, namely, magnetic-coupling and electrical-coupling are analyzed. Magnetic coupling is simulated using dipolar-field interactions. For electricalcoupling we employed global RF-injection. In this method, multiple STOs are phase-locked to a common RF-signal that is injected into the STOs along with the DC bias. Results for variation and noise analysis indicate that electrical-coupling can be significantly more robust as compared to magnetic-coupling. For room-temperature simulations, appreciable phase-lock was retained among tens of electrically coupled STOs for up to 20% 3s random variations in critical device parameters. The magnetic-coupling technique however failed to retain locking beyond ~3% 3s parameter-variations, even for small-size STO clusters with near-neighborhood connectivity. We propose and analyze Dual-Pillar STO (DP-STO) for low-power computing using the proposed electrical coupling method. We observed that DP-STO can better exploit the electrical-coupling technique due to separation between the biasing RF signal and its own RF output.

cond-mat.dis-nn

Spin Neurons: A Possible Path to Energy-Efficient Neuromorphic Computers

Recent years have witnessed growing interest in the field of brain-inspired computing based on neural-network architectures. In order to translate the related algorithmic models into powerful, yet energy-efficient cognitive-computing hardware, computing-devices beyond CMOS may need to be explored. The suitability of such devices to this field of computing would strongly depend upon how closely their physical characteristics match with the essential computing primitives employed in such models. In this work we discuss the rationale of applying emerging spin-torque devices for bio-inspired computing. Recent spin-torque experiments have shown the path to low-current, low-voltage and high-speed magnetization switching in nano-scale magnetic devices. Such magneto-metallic, current-mode spin-torque switches can mimic the analog summing and thresholding operation of an artificial neuron with high energy-efficiency. Comparison with CMOS-based analog circuit-model of neuron shows that spin neurons can achieve more than two orders of magnitude lower energy and beyond three orders of magnitude reduction in energy-delay product. The application of spin neurons can therefore be an attractive option for neuromorphic computers of future.

cond-mat.dis-nn

Boolean and Non-Boolean Computation With Spin Devices

Recently several device and circuit design techniques have been explored for applying nano-magnets and spin torque devices like spin valves and domain wall magnets in computational hardware. However, most of them have been focused on digital logic, and, their benefits over robust and high performance CMOS remains debatable. Ultra-low voltage, current-switching operation of magneto-metallic spin torque devices can potentially be more suitable for non-Boolean computation schemes that can exploit current-mode analog processing. Device circuit co-design for different classes of non-Boolean-architectures using spin-torque based neuron models in spin-CMOS hybrid circuits show that the spin-based non-Boolean designs can achieve 15X-100X lower computation energy for applications like, image-processing, data-conversion, cognitive-computing, pattern matching and programmable-logic, as compared to state of art CMOS designs.

cond-mat.dis-nn

Exploring Boolean and Non-Boolean Computing Applications of Spin Torque Devices

In this paper we discuss the potential of emerging spintorque devices for computing applications. Recent proposals for spinbased computing schemes may be differentiated as all-spin vs. hybrid, programmable vs. fixed, and, Boolean vs. non-Boolean. All spin logic-styles may offer high area-density due to small form-factor of nano-magnetic devices. However, circuit and system-level design techniques need to be explored that leverage the specific spin-device characteristics to achieve energy-efficiency, performance and reliability comparable to those of CMOS. The non-volatility of nanomagnets can be exploited in the design of energy and area-efficient programmable logic. In such logic-styles, spin-devices may play the dual-role of computing as well as memory-elements that provide field-programmability. Spin-based threshold logic design is presented as an example (dynamic resisitve threshold logic and magnetic threshold logic). Emerging spintronic phenomena may lead to ultralow- voltage, current-mode, spin-torque switches that can offer attractive computing capabilities, beyond digital switches. Such devices may be suitable for non-Boolean data-processing applications which involve analog processing. Integration of such spin-torque devices with charge-based devices like CMOS and resistive memory can lead to highly energy-efficient information processing hardware for applications like pattern-matching, neuromorphic-computing, image-processing and data-conversion. Towards the end, we discuss the possibility of applying emerging spin-torque switches in the design of energy-efficient global interconnects, for future chip multiprocessors.

cond-mat.dis-nn

Ultra-High Density, High-Performance and Energy-Efficient All Spin Logic

All Spin Logic gates employ multiple nano-magnets interacting through spin-torque using non-magnetic channels. Compactness, non-volatility and ultra-low voltage operation are some of the attractive features of ASL, while, low switching-speed (of nano-magnets as compared to CMOS gates) and static-power dissipation can be identified as the major bottlenecks. In this work we explore design techniques that leverage the specific device characteristics of ASL to overcome the inefficiencies and to enhance the merits of this technology, for a given set of device parameters. We exploit the non-volatility of nano-magnets to model fully-pipelined ASL that can achieve higher performance. Clocking of power supply in pipelined ASL would require CMOS transistors that may consume significantly large voltage headroom and area, as compared to the nano-magnets. We show that the use of leaky transistors can significantly mitigate such bottlenecks, without sacrificing energy-efficiency and robustness. Exploiting the inherent isolation between the biasing charge current and spin-current paths in ASL, we propose to stack multiple ASL metal layers, leading to ultra-high-density and energy-efficient 3-D computation blocks. Results for the design of an FIR filter show that ASL can achieve performance and power consumption comparable to CMOS while the ultra-high-density of ASL can be projected as its main advantage over CMOS.

cond-mat.mes-hall