SearcharxivSearch

arXiv subjects

Jung-Hoon Kim

Publications and source records attributed to Jung-Hoon Kim.

5 recordsLinked to original sources

LPU: A Latency-Optimized and Highly Scalable Processor for Large Language Model Inference

The explosive arrival of OpenAI's ChatGPT has fueled the globalization of large language model (LLM), which consists of billions of pretrained parameters that embodies the aspects of syntax and semantics. HyperAccel introduces latency processing unit (LPU), a latency-optimized and highly scalable processor architecture for the acceleration of LLM inference. LPU perfectly balances the memory bandwidth and compute logic with streamlined dataflow to maximize performance and efficiency. LPU is equipped with expandable synchronization link (ESL) that hides data synchronization latency between multiple LPUs. HyperDex complements LPU as an intuitive software framework to run LLM applications. LPU achieves 1.25 ms/token and 20.9 ms/token for 1.3B and 66B model, respectively, which is 2.09x and 1.37x faster than the GPU. LPU, synthesized using Samsung 4nm process, has total area of 0.824 mm2 and power consumption of 284.31 mW. LPU-based servers achieve 1.33x and 1.32x energy efficiency over NVIDIA H100 and L4 servers, respectively.

cs.AR

Exploration of Systolic-Vector Architecture with Resource Scheduling for Dynamic ML Workloads

As artificial intelligence (AI) and machine learning (ML) technologies disrupt a wide range of industries, cloud datacenters face ever-increasing demand in inference workloads. However, conventional CPU-based servers cannot handle excessive computational requirements of deep neural network (DNN) models, while GPU-based servers suffer from huge power consumption and high operating cost. In this paper, we present a scalable systolic-vector architecture that can cope with dynamically changing DNN workloads in cloud datacenters. We first devise a lightweight DNN model description format called unified model format (UMF) that enables general model representation and fast decoding in hardware accelerator. Based on this model format, we propose a heterogeneous architecture that features a load balancer that performs a high-level workload distribution and multiple systolic-vector clusters, in which each cluster consists of a programmable scheduler, throughput-oriented systolic arrays, and function-oriented vector processors. We also propose a heterogeneity-aware scheduling algorithm that enables concurrent execution of multiple DNN workloads while maximizing heterogeneous hardware utilization based on computation and memory access time estimation. Finally, we build an architecture simulation framework based on actual synthesis and place-and-route implementation results and conduct design space exploration for the proposed architecture. As a result, the proposed systolic-vector architecture achieves 10.9x higher throughput performance and 30.17x higher energy efficiency than a compatible GPU on realistic ML workloads. The proposed heterogeneity-aware scheduling algorithm improves the throughput and energy efficiency by 81% and 20%, respectively, compared to a standard round-robin scheduling.

cs.AR

Error-correcting codes on scale-free networks

We investigate the potential of scale-free networks as error-correcting codes. We find that irregular low-density parity-check codes with highest performance known to date have degree distributions well fitted by a power-law function $p(k)\sim k^{-γ}$ with $γ$ close to 2, which suggests that codes built on scale-free networks with appropriate power exponents can be good error-correcting codes, with performance possibly approaching the Shannon limit. We demonstrate for an erasure channel that codes with power-law degree distribution of the form $p(k)=C(k+α)^{-γ}$, with $k \geq 2$ and suitable selection of the parameters $α$ and $γ$, indeed have very good error-correction capabilities.

cond-mat.stat-mech

Plasma-Induced Frequency Chirp of Intense Femtosecond Lasers and Its Role in Shaping High-Order Harmonic Spectral Lines

We investigate the self-phase modulation of intense femtosecond laser pulses propagating in an ionizing gas and its effects on collective properties of high-order harmonics generated in the medium. Plasmas produced in the medium are shown to induce a positive frequency chirp on the leading edge of the propagating laser pulse, which subsequently drives high harmonics to become positively chirped. In certain parameter regimes, the plasma-induced positive chirp can help to generate sharply peaked high harmonics, by compensating for the dynamically-induced negative chirp that is caused by the steep intensity profile of intense short laser pulses.

physics.optics

Canonical Transformations and the Hamilton-Jacobi Theory in Quantum Mechanics

Canonical transformations using the idea of quantum generating functions are applied to construct a quantum Hamilton-Jacobi theory, based on the analogy with the classical case. An operator and a c-number forms of the time-dependent quantum Hamilton-Jacobi equation are derived and used to find dynamical solutions of quantum problems. The phase-space picture of quantum mechanics is discussed in connection with the present theory.

quant-ph