SearcharxivSearch

arXiv subjects

Johnson Umeike

Publications and source records attributed to Johnson Umeike.

2 recordsLinked to original sources

Arcal\'{i}s: Accelerating Remote Procedure Calls Using a L\'{i}ghtweight Near-Cache Solution

Modern microservices increasingly depend on high-performance remote procedure calls (RPCs) to coordinate fine-grained, distributed computation. As network bandwidths continue to scale, the CPU overhead associated with RPC processing, particularly serialization, deserialization, and protocol handling, has become a critical bottleneck. This challenge is exacerbated by fast user-space networking stacks such as DPDK, which expose RPC processing as the dominant performance limiter. While prior hardware accelerators have explored NIC-attached and FPGA-based offload, these approaches remain farther from the cache hierarchy, so the frequent data accesses during RPC processing each pay an extra interconnect traversal cost that inflates RPC time. Therefore, RPC handling should occur as close as possible to the cache; however, a near-cache solution must be small, hence practical and deployable. Our key insight to enable such a solution is taking advantage of a reconfigurable accelerator that can be configured specifically for the services currently running on the CPUs. We present Arcal\'{\i}s, a near-cache RPC accelerator that positions a lightweight hardware engine adjacent to the last-level cache (LLC). Arcal\'{\i}s offloads RPC processing to dedicated microengines that operate with cache-line latency while preserving programmability. By decoupling RPC processing logic, enabling microservice-specific execution, and positioning itself near the LLC, Arcal\'{\i}s achieves a 1.72-4.91$\times$ end-to-end speedup compared to the CPU baseline, significantly reduces microarchitectural overhead by up to 88\%, and achieves up to a 1.62$\times$ higher throughput than prior solutions. These results highlight the potential of near-cache RPC acceleration as a practical solution for high-performance microservice deployment.

cs.AR

Belenos: Bottleneck Evaluation to Link Biomechanics to Novel Computing Optimizations

Finite element simulations are essential in biomechanics, enabling detailed modeling of tissues and organs. However, architectural inefficiencies in current hardware and software stacks limit performance and scalability, especially for iterative tasks like material parameter identification. As a result, workflows often sacrifice fidelity for tractability. Reconfigurable hardware, such as FPGAs, offers a promising path to domain-specific acceleration without the cost of ASICs, but its potential in biomechanics remains underexplored. This paper presents Belenos, a comprehensive workload characterization of finite element biomechanics using FEBio, a widely adopted simulator, gem5 sensitivity studies, and VTune analysis. VTune results reveal that smaller workloads experience moderate front-end stalls, typically around 13.1%, whereas larger workloads are dominated by significant back-end bottlenecks, with backend-bound cycles ranging from 59.9% to over 82.2%. Complementary gem5 sensitivity studies identify optimal hardware configurations for Domain-Specific Accelerators (DSA), showing that suboptimal pipeline, memory, or branch predictor settings can degrade performance by up to 37.1%. These findings underscore the need for architecture-aware co-design to efficiently support biomechanical simulation workloads.

cs.AR