SearcharxivSearch

arXiv subjects

Supriyo Datta

Publications and source records attributed to Supriyo Datta.

At least 19 recordsLinked to original sources

Exploiting Negative Capacitance for Unconventional Coulomb Engineering

The many-body ground state of a two-dimensional electron system can be tuned by Coulomb engineering through control of the dielectric environment. However, in conventional dielectrics the static permittivity is restricted to positive values, limiting the accessible interaction regimes. Here we argue that the negative capacitance demonstrated in appropriately engineered structures can open new vistas for Coulomb engineering. The associated negative permittivity could transform the natural repulsive interaction of electrons into an attractive one, raising the intriguing possibility of nontrivial ground states, including superconductivity. Using models of two-dimensional electron systems with linear and parabolic dispersion relations coupled to environments with negative capacitance, we estimate the strength and sign of the engineered Coulomb interaction and outline parameter regimes that could stabilize correlated electronic phases.

cond-mat.mes-hall

Improving deep neural network performance through sampling

Energy efficient sampling with probabilistic neurons or p-bits has been demonstrated in the context of Boltzmann machines and it is natural to ask if these approaches can be extended to the field of generative AI where energy costs have become prohibitively large. However, this very active field is dominated by feedforward deep neural networks (DNNs) which primarily use multi-bit deterministic neurons with no role for sampling. In this paper we first show that it is feasible to obtain superior accuracy through the use of multiple samples generated by probabilistic networks. This possibility raises the question of which option is energetically preferable for improving accuracy: generating more samples, or adding more bits to a single deterministic sample. We provide a simple expression that can be used to estimate these energy tradeoffs and illustrate it with results for different algorithms and architectures.

cond-mat.dis-nn

Energy-Efficient Supervised Learning with a Binary Stochastic Forward-Forward Algorithm

Reducing energy consumption has become a pressing need for modern machine learning, which has achieved many of its most impressive results by scaling to larger and more energy-consumptive neural networks. Unfortunately, the main algorithm for training such networks, backpropagation, poses significant challenges for custom hardware accelerators, due to both its serial dependencies and the memory footprint needed to store forward activations for the backward pass. Alternatives to backprop, although less effective, do exist; here the main computational bottleneck becomes matrix multiplication. In this study, we derive forward-forward algorithms for binary, stochastic units. Binarization of the activations transforms matrix multiplications into indexing operations, which can be executed efficiently in hardware. Stochasticity, combined with tied weights across units with different biases, bypasses the information bottleneck imposed by binary units. Furthermore, although slow and expensive in traditional hardware, binary sampling that is very fast can be implemented cheaply with p-bits (probabilistic bits), novel devices made up of unstable magnets. We evaluate our proposed algorithms on the MNIST, Fashion-MNIST, and CIFAR-10 datasets, showing that its performance is close to real-valued forward-forward, but with an estimated energy savings of about one order of magnitude.

cs.LG

Emergent Synaptic Plasticity from Tunable Dynamics of Probabilistic Bits

Probabilistic (p-) computing, which leverages the stochasticity of its building blocks (p-bits) to solve a variety of computationally hard problems, has recently emerged as a promising physics-inspired hardware accelerator platform. A functionality of importance for p-computers is the ability to program-and reprogram-the interaction strength between arbitrary p-bits on-chip. In natural systems subject to random fluctuations, it is known that spatiotemporal noise can interact with the system's nonlinearities to render useful functionalities. Leveraging that principle, here we introduce a novel scheme for tunable coupling that inserts a ''hidden'' p-bit between each pair of computational p-bits. By modulating the fluctuation rate of the hidden p-bit relative to the synapse speed, we demonstrate both numerically and analytically that the effective interaction between the computational p-bits can be continuously tuned. Moreover, this tunability is directional, where the effective coupling from one computational p-bit to another can be made different from the reverse. This synaptic-plasticity mechanism could open new avenues for designing (re-)configurable p-computers and may inspire novel algorithms that leverage dynamic, hardware-level tuning of stochastic interactions.

cond-mat.dis-nn

Connecting physics to systems with modular spin-circuits

An emerging paradigm in modern electronics is that of CMOS + $\sf X$ requiring the integration of standard CMOS technology with novel materials and technologies denoted by $\sf X$. In this context, a crucial challenge is to develop accurate circuit models for $\sf X$ that are compatible with standard models for CMOS-based circuits and systems. In this perspective, we present physics-based, experimentally benchmarked modular circuit models that can be used to evaluate a class of CMOS + $\sf X$ systems, where $\sf X$ denotes magnetic and spintronic materials and phenomena. This class of materials is particularly challenging because they go beyond conventional charge-based phenomena and involve the spin degree of freedom which involves non-trivial quantum effects. Starting from density matrices $-$ the central quantity in quantum transport $-$ using well-defined approximations, it is possible to obtain spin-circuits that generalize ordinary circuit theory to 4-component currents and voltages (1 for charge and 3 for spin). With step-by-step examples that progressively become more complex, we illustrate how the spin-circuit approach can be used to start from the physics of magnetism and spintronics to enable accurate system-level evaluations. We believe the core approach can be extended to include other quantum degrees of freedom like valley and pseudospins starting from corresponding density matrices.

cond-mat.mes-hall

Heisenberg machines with programmable spin-circuits

We show that we can harness two recent experimental developments to build a compact hardware emulator for the classical Heisenberg model in statistical physics. The first is the demonstration of spin-diffusion lengths in excess of microns in graphene even at room temperature. The second is the demonstration of low barrier magnets (LBMs) whose magnetization can fluctuate rapidly even at sub-nanosecond rates. Using experimentally benchmarked circuit models, we show that an array of LBMs driven by an external current source has a steady-state distribution corresponding to a classical system with an energy function of the form $E = -1/2\sum_{i,j} J_{ij} (\hat{m}_i \cdot \hat{m}_j$). This may seem surprising for a non-equilibrium system but we show that it can be justified by a Lyapunov function corresponding to a system of coupled Landau-Lifshitz-Gilbert (LLG) equations. The Lyapunov function we construct describes LBMs interacting through the spin currents they inject into the spin neutral substrate. We suggest ways to tune the coupling coefficients $J_{ij}$ so that it can be used as a hardware solver for optimization problems involving continuous variables represented by vector magnetizations, similar to the role of the Ising model in solving optimization problems with binary variables. Finally, we train a Heisenberg XOR gate based on a network of four coupled stochastic LLG equations, illustrating the concept of probabilistic computing with a programmable Heisenberg model.

cond-mat.mes-hall

A full-stack view of probabilistic computing with p-bits: devices, architectures and algorithms

The transistor celebrated its 75${}^\text{th}$ birthday in 2022. The continued scaling of the transistor defined by Moore's Law continues, albeit at a slower pace. Meanwhile, computing demands and energy consumption required by modern artificial intelligence (AI) algorithms have skyrocketed. As an alternative to scaling transistors for general-purpose computing, the integration of transistors with unconventional technologies has emerged as a promising path for domain-specific computing. In this article, we provide a full-stack review of probabilistic computing with p-bits as a representative example of the energy-efficient and domain-specific computing movement. We argue that p-bits could be used to build energy-efficient probabilistic systems, tailored for probabilistic algorithms and applications. From hardware, architecture, and algorithmic perspectives, we outline the main applications of probabilistic computers ranging from probabilistic machine learning and AI to combinatorial optimization and quantum simulation. Combining emerging nanodevices with the existing CMOS ecosystem will lead to probabilistic computers with orders of magnitude improvements in energy efficiency and probabilistic sampling, potentially unlocking previously unexplored regimes for powerful probabilistic algorithms.

cs.ET

Roadmap for Unconventional Computing with Nanotechnology

In the "Beyond Moore's Law" era, with increasing edge intelligence, domain-specific computing embracing unconventional approaches will become increasingly prevalent. At the same time, adopting a variety of nanotechnologies will offer benefits in energy cost, computational speed, reduced footprint, cyber resilience, and processing power. The time is ripe for a roadmap for unconventional computing with nanotechnologies to guide future research, and this collection aims to fill that need. The authors provide a comprehensive roadmap for neuromorphic computing using electron spins, memristive devices, two-dimensional nanomaterials, nanomagnets, and various dynamical systems. They also address other paradigms such as Ising machines, Bayesian inference engines, probabilistic computing with p-bits, processing in memory, quantum memories and algorithms, computing with skyrmions and spin waves, and brain-inspired computing for incremental learning and problem-solving in severely resource-constrained environments. These approaches have advantages over traditional Boolean computing based on von Neumann architecture. As the computational requirements for artificial intelligence grow 50 times faster than Moore's Law for electronics, more unconventional approaches to computing and signal processing will appear on the horizon, and this roadmap will help identify future needs and challenges. In a very fertile field, experts in the field aim to present some of the dominant and most promising technologies for unconventional computing that will be around for some time to come. Within a holistic approach, the goal is to provide pathways for solidifying the field and guiding future impactful discoveries.

cs.ET

An Efficient MCMC Approach to Energy Function Optimization in Protein Structure Prediction

Protein structure prediction is a critical problem linked to drug design, mutation detection, and protein synthesis, among other applications. To this end, evolutionary data has been used to build contact maps which are traditionally minimized as energy functions via gradient descent based schemes like the L-BFGS algorithm. In this paper we present what we call the Alternating Metropolis-Hastings (AMH) algorithm, which (a) significantly improves the performance of traditional MCMC methods, (b) is inherently parallelizable allowing significant hardware acceleration using GPU, and (c) can be integrated with the L-BFGS algorithm to improve its performance. The algorithm shows an improvement in energy of found structures of 8.17% to 61.04% (average 38.9%) over traditional MH and 0.53% to 17.75% (average 8.9%) over traditional MH with intermittent noisy restarts, tested across 9 proteins from recent CASP competitions. We go on to map the Alternating MH algorithm to a GPGPU which improves sampling rate by 277x and improves simulation time to a low energy protein prediction by 7.5x to 26.5x over CPU. We show that our approach can be incorporated into state-of-the-art protein prediction pipelines by applying it to both trRosetta2's energy function and the distogram component of Alphafold1's energy function. Finally, we note that specially designed probabilistic computers (or p-computers) can provide even better performance than GPUs for MCMC algorithms like the one discussed here.

q-bio.BM

Accelerated Quantum Monte Carlo with Probabilistic Computers

Quantum Monte Carlo (QMC) techniques are widely used in a variety of scientific problems and much work has been dedicated to developing optimized algorithms that can accelerate QMC on standard processors (CPU). With the advent of various special purpose devices and domain specific hardware, it has become increasingly important to establish clear benchmarks of what improvements these technologies offer compared to existing technologies. In this paper, we demonstrate 2 to 3 orders of magnitude acceleration of a standard QMC algorithm using a specially designed digital processor, and a further 2 to 3 orders of magnitude by mapping it to a clockless analog processor. Our demonstration provides a roadmap for 5 to 6 orders of magnitude acceleration for a transverse field Ising model (TFIM) and could possibly be extended to other QMC models as well. The clockless analog hardware can be viewed as the classical counterpart of the quantum annealer and provides performance within a factor of $<10$ of the latter. The convergence time for the clockless analog hardware scales with the number of qubits as $\sim N$, improving the $\sim N^2$ scaling for CPU implementations, but appears worse than that reported for quantum annealers by D-Wave.

quant-ph

Can Negative Capacitance Induce Superconductivity?

Superconductivity was originally observed in 3D metals caused by an effective attraction between electrons mediated by the electron-phonon interaction. Since then there has been a lot of work on 2D conductors including the possibility of alternative mechanisms that can lead to an effective attractive interaction. Inspired by the experimental demonstration of both steady-state and transient negative capacitance in a variety of structures, this paper investigates the possibility of superconductivity in a two-dimensional conductor embedded in a negative permittivity medium whose role is to turn the normally repulsive Coulomb interaction into an attractive one. A weak coupling BCS theory is used to identify the key parameters that have to be optimized to observe a superconducting transition, especially the need for a small effective negative permittivity, which could be obtained by balancing a negative permittivity medium with a positive permittivity one.

cond-mat.mes-hall

Hardware-aware $in \ situ$ Boltzmann machine learning using stochastic magnetic tunnel junctions

One of the big challenges of current electronics is the design and implementation of hardware neural networks that perform fast and energy-efficient machine learning. Spintronics is a promising catalyst for this field with the capabilities of nanosecond operation and compatibility with existing microelectronics. Considering large-scale, viable neuromorphic systems however, variability of device properties is a serious concern. In this paper, we show an autonomously operating circuit that performs hardware-aware machine learning utilizing probabilistic neurons built with stochastic magnetic tunnel junctions. We show that $in \ situ$ learning of weights and biases in a Boltzmann machine can counter device-to-device variations and learn the probability distribution of meaningful operations such as a full adder. This scalable autonomously operating learning circuit using spintronics-based neurons could be especially of interest for standalone artificial-intelligence devices capable of fast and efficient learning at the edge.

cond-mat.mes-hall

Probabilistic computing with p-bits

Digital computers store information in the form of bits that can take on one of two values 0 and 1, while quantum computers are based on qubits that are described by a complex wavefunction, whose squared magnitude gives the probability of measuring either 0 or 1. Here, we make the case for a probabilistic computer based on p-bits, which take on values 0 and 1 with controlled probabilities and can be implemented with specialized compact energy-efficient hardware. We propose a generic architecture for such p-computers and emulate systems with thousands of p-bits to show that they can significantly accelerate randomized algorithms used in a wide variety of applications including but not limited to Bayesian networks, optimization, Ising models, and quantum Monte Carlo.

cs.ET

Benchmarking a Probabilistic Coprocessor

Computation in the past decades has been driven by deterministic computers based on classical deterministic bits. Recently, alternative computing paradigms and domain-based computing like quantum computing and probabilistic computing have gained traction. While quantum computers based on q-bits utilize quantum effects to advance computation, probabilistic computers based on probabilistic (p-)bits are naturally suited to solve problems that require large amount of random numbers utilized in Monte Carlo and Markov Chain Monte Carlo algorithms. These Monte Carlo techniques are used to solve important problems in the fields of optimization, numerical integration or sampling from probability distributions. However, to efficiently implement Monte Carlo algorithms the generation of random numbers is crucial. In this paper, we present and benchmark a probabilistic coprocessor based on p-bits that are naturally suited to solve these problems. We present multiple examples and project that a nanomagnetic implementation of our probabilistic coprocessor can outperform classical CPU and GPU implementations by multiple orders of magnitude.

cs.ET

Multifunctional Spin Logic Gates In Graphene Spin Circuits

All-spin-based computing combining logic and nonvolatile magnetic memory is promising for emerging information technologies. However, the realization of a universal spin logic operation representing a reconfigurable building block with all-electrical spin current communication has so far remained challenging. Here, we experimentally demonstrate a reprogrammable all-electrical multifunctional spin logic gate in a nanoelectronic device architecture utilizing graphene buses for spin communication and multiplexing and nanomagnets for writing and reading information at room temperature. This gate realizes a multistate majority spin logic operation (sMAJ), which is reconfigured to achieve XNOR, (N)AND, and (N)OR Boolean operations depending on the magnetization of inputs. Physics-based spin circuit model is developed to understand the underlying mechanisms of the multifunctional spin logic gate and its operations. These demonstrations provide a platform for scalable all-electric spin logic and neuromorphic computing in the all-spin domain logic-in-memory architecture.

cond-mat.mes-hall

Quantitative Evaluation of Hardware Binary Stochastic Neurons

Recently there has been increasing activity to build dedicated Ising Machines to accelerate the solution of combinatorial optimization problems by expressing these problems as a ground-state search of the Ising model. A common theme of such Ising Machines is to tailor the physics of underlying hardware to the mathematics of the Ising model to improve some aspect of performance that is measured in speed to solution, energy consumption per solution or area footprint of the adopted hardware. One such approach to build an Ising spin, or a binary stochastic neuron (BSN), is a compact mixed-signal unit based on a low-barrier nanomagnet based design that uses a single magnetic tunnel junction (MTJ) and three transistors (3T-1MTJ) where the MTJ functions as a stochastic resistor (1SR). Such a compact unit can drastically reduce the area footprint of BSNs while promising massive scalability by leveraging the existing Magnetic RAM (MRAM) technology that has integrated 1T-1MTJ cells in ~Gbit densities. The 3T-1SR design however can be realized using different materials or devices that provide naturally fluctuating resistances. Extending previous work, we evaluate hardware BSNs from this general perspective by classifying necessary and sufficient conditions to design a fast and energy-efficient BSN that can be used in scaled Ising Machine implementations. We connect our device analysis to systems-level metrics by emphasizing hardware-independent figures-of-merit such as flips per second and dissipated energy per random bit that can be used to classify any Ising Machine.

cs.ET

Unified Framework for Charge-Spin Interconversion in Spin-Orbit Materials

Materials with spin-orbit coupling are of great interest for various spintronics applications due to the efficient electrical generation and detection of spin-polarized electrons. Over the past decade, many materials have been studied, including topological insulators, transition metals, Kondo insulators, semimetals, semiconductors, and oxides; however, there is no unifying physical framework for understanding the physics and therefore designing a material system and devices with the desired properties. We present a model that binds together the experimental data observed on the wide variety of materials in a unified manner. We show that in a material with a given spin-momentum locking, the density of states plays a crucial role in determining the charge-spin interconversion efficiency, and a simple inverse relationship can be obtained. Remarkably, experimental data obtained over the last decade on many different materials closely follow such an inverse relationship. We further deduce two figure-of-merits of great current interest: the spin-orbit torque (SOT) efficiency (for the direct effect) and the inverse Rashba-Edelstein effect length (for the inverse effect), which statistically show good agreement with the existing experimental data on wide varieties of materials. Especially, we identify a scaling law for the SOT efficiency with respect to the carrier concentration in the sample, which agrees with existing data. Such an agreement is intriguing since our transport model includes only Fermi surface contributions and fundamentally different from the conventional views of the SOT efficiency that includes contributions from all the occupied states.

cond-mat.mes-hall

Autonomous Probabilistic Coprocessing with Petaflips per Second

In this paper we present a concrete design for a probabilistic (p-) computer based on a network of p-bits, robust classical entities fluctuating between -1 and +1, with probabilities that are controlled through an input constructed from the outputs of other p-bits. The architecture of this probabilistic computer is similar to a stochastic neural network with the p-bit playing the role of a binary stochastic neuron, but with one key difference: there is no sequencer used to enforce an ordering of p-bit updates, as is typically required. Instead, we explore \textit{sequencerless} designs where all p-bits are allowed to flip autonomously and demonstrate that such designs can allow ultrafast operation unconstrained by available clock speeds without compromising the solution's fidelity. Based on experimental results from a hardware benchmark of the autonomous design and benchmarked device models, we project that a nanomagnetic implementation can scale to achieve petaflips per second with millions of neurons. A key contribution of this paper is the focus on a hardware metric $-$ flips per second $-$ as a problem and substrate-independent figure-of-merit for an emerging class of hardware annealers known as Ising Machines. Much like the shrinking feature sizes of transistors that have continually driven Moore's Law, we believe that flips per second can be continually improved in later technology generations of a wide class of probabilistic, domain specific hardware.

cs.ET