SearcharxivSearch

arXiv subjects

Ching-Yi Lin

Publications and source records attributed to Ching-Yi Lin.

6 recordsLinked to original sources

ReRAM-aware Model Finetuning addressing I-V Non-linearity and Retention Errors

Traditional CPU, GPU, and NPU architectures are increasingly limited by the von Neumann bottleneck. While In-Memory Computing (IMC) using ReRAM crossbar arrays offers a high-density, energy-efficient alternative, its practical deployment is constrained through their non-idealities. Existing hardware-aware training frameworks often require training from scratch, which is computationally prohibitive for modern large-scale models. In this work, we propose a finetuning-based hardware-aware training algorithm that enables robust DNN deployment on ReRAM with minimal training overhead. Our approach mitigates I-V non-linearity by applying a range-shrunk sinh transformation and incorporates retention errors directly into a regularization loss during the finetuning process. We evaluate our framework across models and tasks such as image classification and question-answering (QA). Experimental results demonstrate that our method achieves similar accuracy on large-scale models like ResNet18 and DeiT-Tiny as the base model. In-case of ImageNet for MobileNetV3 families the technique has only less than 2% accuracy degradation. Further, applying the technique on the SQuAD v2 dataset results in only 1 point degradation of F-1 score.

cs.LG

Design Space Exploration for ReRAM-based Architectures to Address Scaling Non-idealities

ReRAM-based in-memory computing (IMC) architectures are promising candidates for energy-efficient matrix-vector multiplication. While scaling the size of ReRAM arrays allows for the amortization of power-hungry peripheral circuits like DACs and ADCs, it simultaneously introduces more parasitic along the signal path. Because of these challenges, current design methodologies often lack practical guidelines to balance these effects at early design stage, forcing designers to rely on time-consuming, iterative transistor-level simulations. In this work, we propose a comprehensive framework for design space exploration that enables the selection of optimal array size, ADC resolution, and system frequency without requiring exhaustive simulations. The framework utilizes a specialized testbench to extract parameters from a limited set of representative transistor-level simulations. These parameters are then used to accurately predict the performance of arbitrary architectures. We demonstrate the effectiveness of this framework through two realistic design cases aimed at maximizing energy efficiency (TOPs/s/W). The results show that the framework successfully identifies optimal architectural configurations under strict power and error constraints, providing an efficient path for high-performance IMC design.

eess.SY

An Asynchronous Delta Modulator for Spike Encoding in Event-Driven Brain-Machine Interface

This paper presents the design and implementation of an asynchronous delta modulator as a spike encoder for event-driven neural recording in a 65nm CMOS process. The proposed neuromorphic front-end converts analog signals into discrete, asynchronous ON and OFF spikes, effectively compressing continuous biopotentials into spike trains compatible with spiking neural networks (SNNs). Its asynchronous operation enables seamless integration with neuromorphic architectures for real-time decoding in closed-loop brain-machine interfaces (BMIs). Measurement results from silicon demonstrate an energy consumption of 60.73 nJ/spike, an F1-score of 80% compared to a behavioral model of the asynchronous delta modulator, and a compact pixel area of 73.45 um $\times$ 73.64 um.

eess.SY

Ising-ReRAM: A Low Power Ising Machine ReRAM Crossbar for NP Problems

Computational workloads are growing exponentially, driving power consumption to unsustainable levels. Efficiently distributing large-scale networks is an NP-Complete problem equivalent to Boolean satisfiability (SAT), making it one of the core challenges in modern computation. To address this, physics and device inspired methods such as Ising systems have been explored for solving SAT more efficiently. In this work, we implement an Ising model equivalence of the 3-SAT problem using a ReRAM crossbar fabricated in the Skywater 130 nm CMOS process. Our ReRAM-based algorithm achieves $91.0\%$ accuracy in matrix representation across iterative reprogramming cycles. Additionally, we establish a foundational energy profile by measuring the energy costs of small sub-matrix structures within the problem space, demonstrating under linear growth trajectory for combining sub-matrices into larger problems. These results demonstrate a promising platform for developing scalable architectures to accelerate NP-Complete problem solving.

eess.SY

Systolic Array-based Architecture for Low-Bit Integerized Vision Transformers

Transformer-based models are becoming more and more intelligent and are revolutionizing a wide range of human tasks. To support their deployment, AI labs offer inference services that consume hundreds of GWh of energy annually and charge users based on the number of tokens processed. Under this cost model, minimizing power consumption and maximizing throughput have become key design goals for the inference hardware. While graphics processing units (GPUs) are commonly used, their flexibility comes at the cost of low operational intensity and limited efficiency, especially under the high query-per-model ratios of modern inference services. In this work, we address these challenges by proposing a low-bit, model-specialized accelerator that strategically selects tasks with high operation (OP) reuse and minimal communication overhead for offloading. Our design incorporates multiple systolic arrays with deep, fine-grained pipelines and array-compatible units that support essential operations in multi-head self-attention (MSA) module. At the accelerator-level, each self-attention (SA) head is pipelined within a single accelerator to increase data reuse and further minimize bandwidth. Our 3-bit integerized model achieves 96.83% accuracy on CIFAR-10 and 77.81% top-1 accuracy on ImageNet. We validate the hardware design on a 16nm FPGA (Alveo U250), where it delivers 13,568 GigaOps/second (GOPs/s) and 219.4 GOPs/s/W. Compared to a same-technology GPU (GTX 1080), our design offers 1.50x higher throughput and 4.47x better power efficiency. Even against a state-of-the-art GPU (RTX 5090), we still achieve 20% better power efficiency despite having 87% lower throughput.

eess.SY

Low-Bit Integerization of Vision Transformers using Operand Reordering for Efficient Hardware

Pre-trained vision transformers have achieved remarkable performance across various visual tasks but suffer from expensive computational and memory costs. While model quantization reduces memory usage by lowering precision, these models still incur significant computational overhead due to the dequantization before matrix operations. In this work, we analyze the computation graph and propose an integerization process based on operation reordering. Specifically, the process delays dequantization until after matrix operations. This enables integerized matrix multiplication and linear module by directly processing the quantized input. To validate our approach, we synthesize the self-attention module of ViT on a systolic array-based hardware. Experimental results show that our low-bit inference reduces per-PE power consumption for linear layer and matrix multiplication, bridging the gap between quantized models and efficient inference.

cs.LG