Searcharxiv⌕ Search

arXiv subjects

Emad Jacob Maroun

Publications and source records attributed to Emad Jacob Maroun.

2 recordsLinked to original sources

Compiler and Hardware Co-Design for Accelerator Architectures

Heterogeneous accelerator architectures offer an efficient path to performance for compute-intensive workloads. However, full-stack integration remains difficult. We present EAAC (Extensible Accelerator Architecture), a flexible and extensible compiler and hardware architecture designed to lower the overhead of hardware-compiler co-development for rapid prototyping of hardware accelerators. EAAC targets static data-flow workloads with predictable memory access patterns. By using MLIR, we enable possible integration with a range of different frontends that emit MLIR. And by providing a minimal set of compiler functionality that enable necessary data-orchestration we lower the effort needed to get a simple implementation of a hardware acceleration unit up and running, while providing a relatively blank and un-opinionated starting point for further work. We validate this approach with a prototype GEMM accelerator combining a systolic array unit with an embedded RISC-V core, compiled end-to-end through the EAAC MLIR pipeline. On a synthetic fully-connected-layer workload, the accelerator achieves a 28x speedup in execution time over a RISC-V-only baseline, with the GEMM operation itself accounting for only 1.64\% of total execution time. We further characterize the compiler's hardware-semaphore allocation, showing that the number of semaphores required scales linearly with instruction count in the worst case, and identify this as a concrete target for future optimization.

cs.AR↗

A Dense and Efficient Instruction Set Architecture Encoding

Instruction density and encoding efficiency are some of the few things directly affected by an instruction set architecture's design. In contrast, a processor's implementation often significantly influences performance, power efficiency, and area usage. Therefore, a major goal of instruction set design should be maximizing instruction density and encoding efficiency. This paper introduces the design elements of the Scry instruction set architecture that most significantly affect instruction density and encoding efficiency. Scry is a novel and experimental instruction set that revisits first principles to design an instruction set fit for modern processor implementations. Scry uses forward-temporal referencing as a means of data flow, where instructions refer to which future instructions consume their outputs. It also uses internal tagging, where the processors track data types internally, to reduce the number of instructions needed and increase flexibility. Combining these two methods, Scry achieves instruction-feature parity with RISC-V's RV64IMC using only 2-byte instructions compared to RISC-V's 4 bytes. Scry's instructions occupy only 28% of the 2-byte encoding space, where RV64IMC instructions occupy 68% of the 4-byte encoding space. We show that hand-compiled Scry's static instruction density is comparable to RV64IMC for small functions and improves as functions grow in size.

cs.AR↗