SearcharxivSearch

arXiv subjects

Mircea R. Stan

Publications and source records attributed to Mircea R. Stan.

13 recordsLinked to original sources

High-Performance Low-Power Adiabatic Systolic Array Design in Advanced FinFET Nodes

Adiabatic logic has traditionally been recognized as a low-power solution but constrained to low clock speeds to preserve adiabatic behavior. For advanced FinFET nodes, however, clock frequencies have plateaued due to power/thermal concerns (dark silicon) even as the intrinsic device speeds have continued to scale. This convergence opens an opportunity for adiabatic logic to maintain adiabatic behavior even at GHz clocks. We demonstrate an adiabatic logic (AL) design methodology through a MAC systolic array implemented in commercial 16 nm FinFET technology with a resonant 4-phase power clock (PCK) generator, including digital-to-AL and AL-to-digital interfaces. Simulations show that the AL MAC systolic array at 1 GHz achieves power reductions of up to 42% and 36% at the core and system levels, respectively, compared to digital counterparts. Scaling to more advanced nodes should provide even better power/performance metrics.

cs.AR

Guarded Equivalence Predicates for Scalable Formal Hardware Information-Flow Verification

Formal hardware information-flow verification is a principled way to rule out secret-dependent functional or timing observations, but scaling such proofs remains difficult. Self-composition reduces information-flow verification to safety checking over two circuit copies, creating relational proof obligations that are hard for a generic PDR engine to discover from bit-level logic alone. Recent PDR-based techniques exploit this duplicated structure through copy symmetry and global cross-copy equivalence predicates. These predicates are effective when corresponding internal signals agree throughout the reachable state space, but they do not capture equalities that are relevant only in a specific control context. We observe that such contextual relations arise naturally in hardware IFV proofs: an internal signal pair may need to agree only within a control phase, transaction window, loop state, or protocol region. We introduce guarded equivalence predicates to expose these relations to PDR. Rather than treating a proposed contextual equality as an assumption, the verifier submits the corresponding mismatch condition as an auxiliary blocking obligation. Guards are proposed from relational counterexamples-to-induction using CTI-local extraction and state-split search; only candidates proved unreachable by the backend affect the proof. Across 12 IFV benchmarks and two PDR backends, guarded predicates convert two contextual baseline timeouts into completed proofs within 34.2--89.5s under an 1800s limit, while reducing proof time by up to 10.8x on additional benchmarks.

cs.CR

Cool-3D: An End-to-End Thermal-Aware Framework for Early-Phase Design Space Exploration of Microfluidic-Cooled 3DICs

The rapid advancement of three-dimensional integrated circuits (3DICs) has heightened the need for early-phase design space exploration (DSE) to minimize design iterations and unexpected challenges. Emphasizing the pre-register-transfer level (Pre-RTL) design phase is crucial for reducing trial-and-error costs. However, 3DIC design introduces additional complexities due to thermal constraints and an expanded design space resulting from vertical stacking and various cooling strategies. Despite this need, existing Pre-RTL DSE tools for 3DICs remain scarce, with available solutions often lacking comprehensive design options and full customization support. To bridge this gap, we present Cool-3D, an end-to-end, thermal-aware framework for 3DIC design that integrates mainstream architectural-level simulators, including gem5, McPAT, and HotSpot 7.0, with advanced cooling models. Cool-3D enables broad and fine-grained design space exploration, built-in microfluidic cooling support for thermal analysis, and an extension interface for non-parameterizable customization, allowing designers to model and optimize 3DIC architectures with greater flexibility and accuracy. To validate the Cool-3D framework, we conduct three case studies demonstrating its ability to model various hardware design options and accurately capture thermal behaviors. Cool-3D serves as a foundational framework that not only facilitates comprehensive 3DIC design space exploration but also enables future innovations in 3DIC architecture, cooling strategies, and optimization techniques. The entire framework, along with the experimental data, is in the process of being released on GitHub.

cs.AR

Strained topological insulator spin-orbit torque random access memory (STI-SOTRAM) bit cell for energy-efficient Processing in Memory

We present a novel design of a strained topological insulator spin-orbit torque random access memory (STI-SOTRAM) bit cell comprising a piezoelectric/magnet (gating)/topological insulator (TI)/magnet (storage) heterostructure that leverages the TI's high charge-to-spin conversion efficiency coupled with the piezo-induced strain-based gating mechanism for low-power in-memory computing. The piezo-induced strain effectively modulates the conductivity of the topological surface state (TSS) by altering the gating magnet's magnetization from out-to-in-plane, facilitating the storage magnet's spin-orbit torque (SOT) switching. Through comprehensive coupled stochastic Landau-Lifshitz-Gilbert (LLG) simulations, we explore the device dynamics, anisotropy-stress phase space for switching, and write conditions and demonstrate a significant reduction in energy dissipation compared to conventional heavy metal (HM)-based SOT switching. Additionally, we project the energy consumption for in-memory Boolean operations (AND and OR). Our findings suggest the promise of the STI-SOTRAM for low-power, high-performance edge computing.

cond-mat.mes-hall

Memristive Learning Cellular Automata: Theory and Applications

Memristors are novel non volatile devices that manage to combine storing and processing capabilities in the same physical place.Their nanoscale dimensions and low power consumption enable the further design of various nanoelectronic processing circuits and corresponding computing architectures, like neuromorhpic, in memory, unconventional, etc.One of the possible ways to exploit the memristor's advantages is by combining them with Cellular Automata (CA).CA constitute a well known non von Neumann computing architecture that is based on the local interconnection of simple identical cells forming N-dimensional grids.These local interconnections allow the emergence of global and complex phenomena.In this paper, we propose a hybridization of the CA original definition coupled with memristor based implementation, and, more specifically, we focus on Memristive Learning Cellular Automata (MLCA), which have the ability of learning using also simple identical interconnected cells and taking advantage of the memristor devices inherent variability.The proposed MLCA circuit level implementation is applied on optimal detection of edges in image processing through a series of SPICE simulations, proving its robustness and efficacy.

cs.ET

Reservoir Computing based Neural Image Filters

Clean images are an important requirement for machine vision systems to recognize visual features correctly. However, the environment, optics, electronics of the physical imaging systems can introduce extreme distortions and noise in the acquired images. In this work, we explore the use of reservoir computing, a dynamical neural network model inspired from biological systems, in creating dynamic image filtering systems that extracts signal from noise using inverse modeling. We discuss the possibility of implementing these networks in hardware close to the sensors.

cs.CV

Hardware based Spatio-Temporal Neural Processing Backend for Imaging Sensors: Towards a Smart Camera

In this work we show how we can build a technology platform for cognitive imaging sensors using recent advances in recurrent neural network architectures and training methods inspired from biology. We demonstrate learning and processing tasks specific to imaging sensors, including enhancement of sensitivity and signal-to-noise ratio (SNR) purely through neural filtering beyond the fundamental limits sensor materials, and inferencing and spatio-temporal pattern recognition capabilities of these networks with applications in object detection, motion tracking and prediction. We then show designs of unit hardware cells built using complementary metal-oxide semiconductor (CMOS) and emerging materials technologies for ultra-compact and energy-efficient embedded neural processors for smart cameras.

cs.CV

Tolerating Soft Errors in Processor Cores Using CLEAR (Cross-Layer Exploration for Architecting Resilience)

We present CLEAR (Cross-Layer Exploration for Architecting Resilience), a first of its kind framework which overcomes a major challenge in the design of digital systems that are resilient to reliability failures: achieve desired resilience targets at minimal costs (energy, power, execution time, area) by combining resilience techniques across various layers of the system stack (circuit, logic, architecture, software, algorithm). This is also referred to as cross-layer resilience. In this paper, we focus on radiation-induced soft errors in processor cores. We address both single-event upsets (SEUs) and single-event multiple upsets (SEMUs) in terrestrial environments. Our framework automatically and systematically explores the large space of comprehensive resilience techniques and their combinations across various layers of the system stack (586 cross-layer combinations in this paper), derives cost-effective solutions that achieve resilience targets at minimal costs, and provides guidelines for the design of new resilience techniques. Our results demonstrate that a carefully optimized combination of circuit-level hardening, logic-level parity checking, and micro-architectural recovery provides a highly cost-effective soft error resilience solution for general-purpose processor cores. For example, a 50x improvement in silent data corruption rate is achieved at only 2.1% energy cost for an out-of-order core (6.1% for an in-order core) with no speed impact. However, (application-aware) selective circuit-level hardening alone, guided by a thorough analysis of the effects of soft errors on application benchmarks, provides a cost-effective soft error resilience solution as well (with ~1% additional energy cost for a 50x improvement in silent data corruption rate).

cs.AR

CLEAR: Cross-Layer Exploration for Architecting Resilience - Combining Hardware and Software Techniques to Tolerate Soft Errors in Processor Cores

We present a first of its kind framework which overcomes a major challenge in the design of digital systems that are resilient to reliability failures: achieve desired resilience targets at minimal costs (energy, power, execution time, area) by combining resilience techniques across various layers of the system stack (circuit, logic, architecture, software, algorithm). This is also referred to as cross-layer resilience. In this paper, we focus on radiation-induced soft errors in processor cores. We address both single-event upsets (SEUs) and single-event multiple upsets (SEMUs) in terrestrial environments. Our framework automatically and systematically explores the large space of comprehensive resilience techniques and their combinations across various layers of the system stack (586 cross-layer combinations in this paper), derives cost-effective solutions that achieve resilience targets at minimal costs, and provides guidelines for the design of new resilience techniques. We demonstrate the practicality and effectiveness of our framework using two diverse designs: a simple, in-order processor core and a complex, out-of-order processor core. Our results demonstrate that a carefully optimized combination of circuit-level hardening, logic-level parity checking, and micro-architectural recovery provides a highly cost-effective soft error resilience solution for general-purpose processor cores. For example, a 50x improvement in silent data corruption rate is achieved at only 2.1% energy cost for an out-of-order core (6.1% for an in-order core) with no speed impact. However, selective circuit-level hardening alone, guided by a thorough analysis of the effects of soft errors on application benchmarks, provides a cost-effective soft error resilience solution as well (with ~1% additional energy cost for a 50x improvement in silent data corruption rate).

cs.AR

Computing with Non-equilibrium Ratchets

Electronic ratchets transduce local spatial asymmetries into directed currents in the absence of a global drain bias, by rectifying temporal signals that reside far from thermal equilibrium. We show that the absence of a drain bias can provide distinct energy advantages for computation, specifically, reducing static dissipation in a logic circuit. Since the ratchet functions as a gate voltage-controlled current source, it also potentially reduces the dynamic dissipation associated with charging/discharging capacitors. In addition, the unique charging mechanism eliminates timing related constraints on logic inputs, in principle allowing for adiabatic charging. We calculate the ratchet currents in classical and quantum limits, and show how a sequence of ratchets can be cascaded to realize universal Boolean logic.

cond-mat.mes-hall

Graphene Nanoribbons: from chemistry to circuits

The Y-chart is a powerful tool for understanding the relationship between various views (behavioral, structural, physical) of a system, at different levels of abstraction, from high-level, architecture and circuits, to low-level, devices and materials. We thus use the Y-chart adapted for graphene to guide the chapter and explore the relationship among the various views and levels of abstraction. We start with the innermost level, namely, the structural and chemical view. The edge chemistry of patterned graphene nanoribbons (GNR) lies intermediate between graphene and benzene, and the corresponding strain lifts the degeneracy that otherwise promotes metallicity in bulk graphene. At the same time, roughness at the edges washes out chiral signatures, making the nanoribbon width the principal arbiter of metallicity. The width-dependent conductivity allows the design of a monolithically patterned wide-narrow-wide all graphene interconnect-channel heterostructure. In a three-terminal incarnation, this geometry exhibits superior electrostatics, a correspondingly benign short-channel effect and a reduction in the contact Schottky barrier through covalent bonding. However, the small bandgaps make the devices transparent to band-to-band tunneling. Increasing the gap with width confinement (or other ways to break the sublattice symmetry) is projected to reduce the mobility even for very pure samples, through a fundamental asymptotic constraint on the bandstructure. An analogous trade-off, ultimately between error rate (reliability) and delay (switching speed) can be projected to persist for all graphitic derivatives. Proceeding thus to a higher level, a compact model is presented to capture the complex nanoribbon circuits, culminating in inverter characteristics, design metrics and layout diagrams.

cond-mat.mes-hall

Monolithically Patterned Wide-Narrow-Wide All-Graphene Devices

We investigate theoretically the performance advantages of all-graphene nanoribbon field-effect transistors (GNRFETs) whose channel and source/drain (contact) regions are patterned monolithically from a two-dimensional single sheet of graphene. In our simulated devices, the source/drain and interconnect regions are composed of wide graphene nanoribbon (GNR) sections that are semimetallic, while the channel regions consist of narrow GNR sections that open semiconducting bandgaps. Our simulation employs a fully atomistic model of the device, contact and interfacial regions using tight-binding theory. The electronic structures are coupled with a self-consistent three-dimensional Poisson's equation to capture the nontrivial contact electrostatics, along with a quantum kinetic formulation of transport based on non-equilibrium Green's functions (NEGF). Although we only consider a specific device geometry, our results establish several general performance advantages of such monolithic devices (besides those related to fabrication and patterning), namely the improved electrostatics, suppressed short-channel effects, and Ohmic contacts at the narrow-to-wide interfaces.

cond-mat.mes-hall

Diluted chirality dependence in edge rough graphene nanoribbon field-effect transistors

We investigate the role of various structural nonidealities on the performance of armchair-edge graphene nanoribbon field effect transistors (GNRFETs). Our results show that edge roughness dilutes the chirality dependence often predicted by theory but absent experimentally. Instead, GNRs are classifiable into wide (semi-metallic) vs narrow (semiconducting) strips, defining thereby the building blocks for wide-narrow-wide all-graphene devices and interconnects. Small bandgaps limit drain bias at the expense of band-to-band tunneling in GNRFETs. We outline the relation between device performance metrics and non-idealities such as width modulation, width dislocations and surface step, and non-ideality parameters such as roughness amplitude and correlation length.

cond-mat.mes-hall