Searcharxiv⌕ Search

arXiv · 2610.02288

AdaptViT: Runtime-Adaptive Vision Transformer Deployment on Custom RISC-V

Abstract

Deploying Vision Transformers (ViTs) on low-power edge devices is challenging due to high computational demands. Conventional pruning frameworks require a separate compiled binary for each sparsity level, increasing storage overhead and limiting runtime adaptability. This paper presents an end-to-end deployment pipeline that transforms pretrained ViTs into a single runtime-configurable binary, enabling dynamic compute-budget switching on embedded CPUs. This is achieved by restructuring generated C kernels with modified loop bounds and binary-mask control logic, allowing execution to switch across discrete sparsity levels via compact external configuration files. Compared to multi-binary deployment, the proposed runtime-adaptive approach reduces on-device storage by up to 4.86x, requiring only 163 MB for ViT-Base instead of nearly 800 MB. To maximize pruning efficiency, we introduce a hardware-aligned block pruning strategy for Multi-Layer Perceptron (MLP) layers. In addition, a custom ISA extension is proposed to exploit input-reuse patterns in linear projection kernels. On a Synopsys TRV32P3FX RISC-V processor, the full system achieves up to 2.8x speedup at 65% MLP and 50% attention-head pruning for ViT-Base. The ISA extension alone provides a 1.56x speedup and 33% lower inference energy, with a 24.7% area overhead in a TSMC 28 nm implementation.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Vishnu PS, Ajay Kumar M, Yike Li, Robert Bogdan Staszewski, Deepu John. 2026-10-01. AdaptViT: Runtime-Adaptive Vision Transformer Deployment on Custom RISC-V. https://arxiv.org/abs/2610.02288

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster

The electric power supply for AI data centers is now the most significant bottleneck in the race toward Artificial General Intelligence, surpassing even the constraint of AI accelerator availability. To our knowledge, this paper is the first to describe the end-to-end power management process for a hyper-scale AI datacenter; from early power planning to accommodate next-generation accelerators 6--12 months before their general availability, to tuning power settings after large scale deployment, and finally to dynamic, runtime power management for evolving workloads. We present detailed power measurements for a 150 MW datacenter hosting a cluster of 83K GB200 GPUs. We share insights from building this state-of-the-art AI cluster. We hope this work encourages practitioners across the industry to share their own experiences as well.

cs.AR↗

PEEK: Heterogeneous Parallelism for Privileged Error Detection in Safety-Critical Processors

Heterogeneous parallel error detection architecture has been widely studied for safeguarding OoO superscalar processors in safetycritical systems, as it achieves significantly lower hardware overhead compared to traditional LockStep, by exploiting the parallelism that exists in a secondary execution. However, previous works do not cover the protection of privileged-mode execution, impeding their effectiveness in real-world deployment. Moreover, naive extension to privileged-mode can cause a litany of issues, from abysmal performance due to high synchronization costs, to full deadlocks. Here, we present PEEK, the first privileged parallel error detection architecture. Based on a deep analysis of privileged execution, we redesign the verification pipeline, addressing all the bottlenecks and bugs identified in privileged-mode protection. Evaluated using various metrics on an RTL-level full system running Linux, PEEK achieves full-privilege protection on Linux with negligible performance slowdown and affordable hardware overhead. PEEK has been taped out using a 28nm process, and its source is available at https://anonymous.4open.science/r/PEEK-3000.

cs.AR↗

Divide and conquer: Scalable performance and energy in MCM GPUs

Multi-chip-module (MCM) GPUs offer a promising path to scale compute capability beyond monolithic designs by integrating multiple chiplets on a common package. However, the impact of disaggregation on performance scalability and energy consumption remains underexplored. The design space grows rapidly across dimensions such as SMs per chiplet, chiplet count, and interconnection network. The inter-chiplet network is particularly critical, as it determines whether additional compute resources translate into performance gains. This limited understanding leaves industry and research without clear guidance on the performance and energy trade-offs of MCM GPU scaling. In this work, we investigate whether distributing compute and memory capability across multiple chiplets offers a more scalable alternative to concentrating resources. We quantify their effects on performance, energy, and efficiency and examine how inter-chiplet topology influences scalability at different system sizes. Our results demonstrate that a 16 chiplet Torus configuration with 256 SMs delivers a remarkable $2.40\times$ performance improvement over a state-of-the-art MCM architecture with the same compute capability, while simultaneously reducing energy consumption by $4.45\times$. These substantial gains provide evidence that disaggregation is a first-order architectural factor and will be critical to unlocking the performance and energy-efficiency potential of next-generation GPUs.

cs.AR↗