arXiv · 2607.13649
CIMERA: Compute-in-Interconnect and Memory with Reconfigurable Precision for LLM Inference
Abstract
LLM impose significant computational and memory demands, creating challenges for energy-efficient inference across platforms ranging from data centers to power-constrained edge devices. Weight precision plays a critical role in balancing inference accuracy, throughput, and energy consumption, while modern LLM workloads exhibit pronounced heterogeneity and tolerance that favors adaptive precision execution. This paper presents CIMERA, a reconfigurable-precision LLM inference accelerator that integrates compute-in-interconnect and memory to mitigate the memory wall and enable precision-aware execution. Compared to Nvidia H100, CIMERA delivers up to $25\times$ and $10\times$ higher energy efficiency for 1B and 13B models, respectively.
Explore related subjects
Keep this discovery
Yue Jiet Chong, Yimin Wang, Wei Zhang, Xuanyao Fong. 2026-07-15. CIMERA: Compute-in-Interconnect and Memory with Reconfigurable Precision for LLM Inference. https://arxiv.org/abs/2607.13649
Cite the original work for its findings. Save a collection to share your selection of sources.