SearcharxivSearch

arXiv subjects

Jeongmin Hong

Publications and source records attributed to Jeongmin Hong.

7 recordsLinked to original sources

Low-overhead General-purpose Near-Data Processing in CXL Memory Expanders

Emerging Compute Express Link (CXL) enables cost-efficient memory expansion beyond the local DRAM of processors. While its CXL$.$mem protocol provides minimal latency overhead through an optimized protocol stack, frequent CXL memory accesses can result in significant slowdowns for memory-bound applications whether they are latency-sensitive or bandwidth-intensive. The near-data processing (NDP) in the CXL controller promises to overcome such limitations of passive CXL memory. However, prior work on NDP in CXL memory proposes application-specific units that are not suitable for practical CXL memory-based systems that should support various applications. On the other hand, existing CPU or GPU cores are not cost-effective for NDP because they are not optimized for memory-bound applications. In addition, the communication between the host processor and CXL controller for NDP offloading should achieve low latency, but existing CXL$.$io/PCIe-based mechanisms incur $μ$s-scale latency and are not suitable for fine-grained NDP. To achieve high-performance NDP end-to-end, we propose a low-overhead general-purpose NDP architecture for CXL memory referred to as Memory-Mapped NDP (M$^2$NDP), which comprises memory-mapped functions (M$^2$func) and memory-mapped $μ$threading (M$^2μ$thread). M$^2$func is a CXL$.$mem-compatible low-overhead communication mechanism between the host processor and NDP controller in CXL memory. M$^2μ$thread enables low-cost, general-purpose NDP unit design by introducing lightweight $μ$threads that support highly concurrent execution of kernels with minimal resource wastage. Combining them, M$^2$NDP achieves significant speedups for various workloads by up to 128x (14.5x overall) and reduces energy by up to 87.9% (80.3% overall) compared to baseline CPU/GPU hosts with passive CXL memory.

cs.AR

Bandwidth-Effective DRAM Cache for GPUs with Storage-Class Memory

We propose overcoming the memory capacity limitation of GPUs with high-capacity Storage-Class Memory (SCM) and DRAM cache. By significantly increasing the memory capacity with SCM, the GPU can capture a larger fraction of the memory footprint than HBM for workloads that oversubscribe memory, achieving high speedups. However, the DRAM cache needs to be carefully designed to address the latency and BW limitations of the SCM while minimizing cost overhead and considering GPU's characteristics. Because the massive number of GPU threads can thrash the DRAM cache, we first propose an SCM-aware DRAM cache bypass policy for GPUs that considers the multi-dimensional characteristics of memory accesses by GPUs with SCM to bypass DRAM for data with low performance utility. In addition, to reduce DRAM cache probes and increase effective DRAM BW with minimal cost, we propose a Configurable Tag Cache (CTC) that repurposes part of the L2 cache to cache DRAM cacheline tags. The L2 capacity used for the CTC can be adjusted by users for adaptability. Furthermore, to minimize DRAM cache probe traffic from CTC misses, our Aggregated Metadata-In-Last-column (AMIL) DRAM cache organization co-locates all DRAM cacheline tags in a single column within a row. The AMIL also retains the full ECC protection, unlike prior DRAM cache's Tag-And-Data (TAD) organization. Additionally, we propose SCM throttling to curtail power and exploiting SCM's SLC/MLC modes to adapt to workload's memory footprint. While our techniques can be used for different DRAM and SCM devices, we focus on a Heterogeneous Memory Stack (HMS) organization that stacks SCM dies on top of DRAM dies for high performance. Compared to HBM, HMS improves performance by up to 12.5x (2.9x overall) and reduces energy by up to 89.3% (48.1% overall). Compared to prior works, we reduce DRAM cache probe and SCM write traffic by 91-93% and 57-75%, respectively.

cs.AR

Nanoscale three-dimensional magnetic sensing with a probabilistic nanomagnet driven by spin-orbit torque

Detection of vector magnetic fields at nanoscale dimensions is critical in applications ranging from basic material science, to medical diagnostic. Meanwhile, an all-electric operation is of great significance for achieving a simple and compact sensing system. Here, we propose and experimentally demonstrate a simple approach to sensing a vector magnetic field at nanoscale dimensions, by monitoring a probabilistic nanomagnet's transition probability from a metastable state, excited by a driving current due to SOT, to a settled state. We achieve sensitivities for Hx, Hy, and Hz of 1.02%/Oe, 1.09%/Oe and 3.43%/Oe, respectively, with a 200 x 200 nm^2 nanomagnet. The minimum detectable field is dependent on the driving pulse events N, and is expected to be as low as 1 uT if N = 3 x 10^6.

cond-mat.mtrl-sci

Large magnetoresistance in an electric field controlled antiferromagnetic tunnel junction

Large magnetoresistance effect controlled by electric field rather than magnetic field or electric current is a preferable routine for designing low power consumption magnetoresistance-based spintronic devices. Here we propose an electric-field controlled antiferromagnetic (AFM) tunnel junction with structure of piezoelectric substrate/Mn3Pt/SrTiO3/Pt operating by the magnetic phase transition (MPT) of antiferromagnet Mn3Pt through its magneto-volume effect. The transport properties of the proposed AFM tunnel junction have been investigated by employing first-principles calculations. Our results show that a magnetoresistance over hundreds of percent is achievable when Mn3Pt undergoes MPT from a collinear AFM state to a non-collinear AFM state. Band structure analysis based on density functional calculations shows that the large TMR can be attributed to the joint effect of significant different Fermi surface of Mn3Pt at two AFM phases and the band symmetry filtering effect of the SrTiO3 tunnel barrier. In addition, other than single-crystalline tunnel barrier, we also discuss the robustness of the proposed magnetoresistance effect by considering amorphous AlOx barrier. Our results may open perspective way for effectively electrical writing and reading of the AFM state and its application in energy efficient magnetic memory devices.

cond-mat.mes-hall

Switching of Perpendicularly Polarized Nanomagnets with Spin Orbit Torque without an External Magnetic Field by Engineering a Tilted Anisotropy

Spin orbit torque (SOT) provides an efficient way of generating spin current that promises to significantly reduce the current required for switching nanomagnets. However, an in-plane current generated SOT cannot deterministically switch a perpendicularly polarized magnet due to symmetry reasons. On the other hand, perpendicularly polarized magnets are preferred over in-plane magnets for high-density data storage applications due to their significantly larger thermal stability in ultra-scaled dimensions. Here we show that it is possible switch a perpendicularly polarized magnet by SOT without needing an external magnetic field. This is accomplished by engineering an anisotropy in the magnets such that the magnetic easy axis slightly tilts away from the film-normal. Such a tilted anisotropy breaks the symmetry of the problem and makes it possible to switch the magnet deterministically. Using a simple Ta/CoFeB/MgO/Ta heterostructure, we demonstrate reversible switching of the magnetization by reversing the polarity of the applied current. This demonstration presents a new approach for controlling nanomagnets with spin orbit torque.

cond-mat.mtrl-sci

Sub-nanosecond signal propagation in anisotropy engineered nanomagnetic logic chains

Energy efficient nanomagnetic logic (NML) computing architectures propagate and process binary information by relying on dipolar field coupling to reorient closely-spaced nanoscale magnets. Signal propagation in nanomagnet chains of various sizes, shapes, and magnetic orientations has been previously characterized by static magnetic imaging experiments with low-speed adiabatic operation; however the mechanisms which determine the final state and their reproducibility over millions of cycles in high-speed operation (sub-ns time scale) have yet to be experimentally investigated. Monitoring NML operation at its ultimate intrinsic speed reveals features undetectable by conventional static imaging including individual nanomagnetic switching events and systematic error nucleation during signal propagation. Here, we present a new study of NML operation in a high speed regime at fast repetition rates. We perform direct imaging of digital signal propagation in permalloy nanomagnet chains with varying degrees of shape-engineered biaxial anisotropy using full-field magnetic soft x-ray transmission microscopy after applying single nanosecond magnetic field pulses. Further, we use time-resolved magnetic photo-emission electron microscopy to evaluate the sub-nanosecond dipolar coupling signal propagation dynamics in optimized chains with 100 ps time resolution as they are cycled with nanosecond field pulses at a rate of 3 MHz. An intrinsic switching time of 100 ps per magnet is observed. These experiments, and accompanying macro-spin and micromagnetic simulations, reveal the underlying physics of NML architectures repetitively operated on nanosecond timescales and identify relevant engineering parameters to optimize performance and reliability.

cond-mat.mes-hall

Speed and Reliability of Nanomagnetic Logic Technology

Nanomagnetic logic is an energy efficient computing architecture that relies on the dipole field coupling of neighboring magnets to transmit and process binary information. In this architecture, nanomagnet chains act as local interconnects. To assess the merits of this technology, the speed and reliability of magnetic signal transmission along these chains must be experimentally determined. In this work, time-resolved pump-probe x-ray photo-emission electron microscopy is used to observe magnetic signal transmission along a chain of nanomagnets. We resolve successive error-free switching events in a single nanomagnet chain at speeds on the order of 100 ps per nanomagnet, consistent with predictions based on micromagnetic modeling. Errors which disrupt transmission are also observed. We discuss the nature of these errors, and approaches for achieving reliable operation.

cond-mat.mes-hall