Searcharxiv⌕ Search

arXiv · 2610.02502

RAPID: Row-Parallel Arithmetic Processing in DRAM

Abstract

Processing-using-memory (PUM) architectures perform computation directly within DRAM to reduce costly data movement between memory and processors. Because charge-sharing operations are confined to individual bitlines, existing DRAM-PUM architectures reorganize data into column-oriented, bit-serial representations. This organization is fundamentally incompatible with the row-oriented, word-parallel layouts used by conventional processors and accelerators, requiring expensive data-layout transformations whenever computation transitions between PUM and conventional execution. In this paper, we present RAPID, a Row-parallel Arithmetic Processing-In-DRAM architecture. RAPID augments the DRAM subarray with two lightweight extensions: migration cells that enable localized horizontal data movement between neighboring bitlines and inversion cells that provide efficient in-array logical inversion. These primitives enable RAPID to operate directly on row-parallel, bit-parallel data, preserving CPU-compatible layouts while exploiting the massive parallelism of the DRAM subarray. In particular, RAPID demonstrates that localized horizontal communication is sufficient to realize shallow arithmetic networks and efficient parallel reduction for multiplication, reducing arithmetic latency while preserving throughput and eliminating costly data-layout transformations, all while maintaining the conventional DRAM array organization. We demonstrate the feasibility and overhead of augmenting DRAM subarrays with migration and inversion cells through detailed transistor-level layout and SPICE-validated circuit simulations. Using the RAPID compiler it is possible to evaluate the performance and data reorganization tradeoffs to ensure the best execution across combined CPU and PUM. Evaluating RAPID on 19 MLPerf benchmarks, there is a 5.9x higher end-to-end performance compared to SIMDRAM for DDR4 PUM execution.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

William C. Tegge, João Paulo Cardoso de Lima, Shouzhi Fang, Jeronimo Castrillon, Alex K. Jones. 2026-10-01. RAPID: Row-Parallel Arithmetic Processing in DRAM. https://arxiv.org/abs/2610.02502

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A Self-Calibrating Framework for Analog Circuit Sizing Using LLM-Derived Analytical Equations

We present a design automation framework for analog circuit sizing that produces calibrated, topology-specific analytical equations from raw circuit netlists. A large language model (LLM) derives a complete Python sizing function in which each device dimension is traceable to a specific design rationale - a form of interpretable output absent from existing optimization-based and LLM-based sizing methods. A deterministic calibration loop extracts process-dependent parameters from a single DC operating point simulation, while a prediction-error feedback mechanism compensates for analytical inaccuracies. We validate the framework on circuits ranging from 6 to 30 transistors - spanning single-stage, current-mirror (simple and cascoded), folded-cascode, gain-boosted folded-cascode, two-stage Miller-compensated, nested-Miller-compensated, and complementary class-AB output topologies - across six process nodes from 32 nm to 180 nm. On matched-specification benchmarks, including the class-AB opamp case, the framework converges within a few simulations. Despite large initial prediction errors, convergence depends on the measurement-feedback architecture, not prediction accuracy. The one-shot calibration automatically captures process-dependent variations, enabling cross-node portability without modification, retraining, or per-process characterization.

cs.AR↗

Towards Enabling Distance-Based Memory Addressing

Approximate Nearest-Neighbor Search (ANNS) in high dimensional vector datasets is an application of significant prevalence across different AI applications. However, such an operation is significantly bandwidth limited at large workingset sizes owing to the curse of dimensionality. Traditional indices used to accelerate ANNS rely on search-space pruning as a preprocessing step to alleviate such bandwidth requirement, but such optimization occurs either at the cost of increased bandwidth-inefficiency and/or degradation of search quality. This paper proposes a data-parallel hardware/software mechanism for performing large-scale similarity search in-memory. We propose a novel algorithm to simplify the computation requirement for similarity search across various distance metrics through lightweight primitives to perform a fast and approximate data-parallel brute-force search on the entire vector space. We further build a memory system capable of executing the required operations to generate a distance metric per datapoints, which is then used to enable pruning as a post-processing step. We offer adequate software support for user control over the proposed system. By enabling such search-space pruning as a post-processing step, we achieve near-perfect recall across representative workloads while achieving orders of magnitude performance and energy improvement over state-of-the-art algorithmic approaches on million and billion-scale workloads.

cs.AR↗

Reinforcement Learning-Based Optimization of Workload-Aware Power Delivery Networks

Power Delivery Networks (PDNs) are critical components of modern VLSI chips, providing stable voltage levels while satisfying electromigration (EM) and IR-drop constraints. Conventional PDN design methodologies typically rely on worst-case assumptions, often resulting in over-provisioned networks and inefficient use of resources. This paper presents a reinforcement learning-based framework for the optimization of workload-aware PDNs. The proposed methodology first generates workload-aware PDNs using architectural power traces obtained from system-level simulations. These power traces are mapped to spatial power density distributions, enabling adaptive allocation of PDN resources according to local current demand. A reinforcement learning agent then performs wire-width optimization to minimize PDN area while maintaining EM and voltage integrity constraints. Electrical and reliability metrics are obtained using SPICE-based circuit analysis and EM lifetime estimation. Experimental evaluation is performed on a dataset of workload-aware PDNs generated from 4-, 8-, and 16-core multiprocessor floorplans using PARSEC and SPLASH-2 benchmark workloads. Furthermore, the proposed Deep Q-Network (DQN)-based optimizer reduces the average normalized PDN area by 47\% while satisfying all EM and IR-drop constraints. Compared to simulated annealing, the proposed approach achieves comparable optimization quality while providing approximately 26$\times$ faster optimization.

cs.AR↗