Searcharxiv⌕ Search

arXiv subjects

Mostafa Eghbali Zarch

Publications and source records attributed to Mostafa Eghbali Zarch.

4 recordsLinked to original sources

Fused Breadth-First Probabilistic Traversals on Distributed GPU Systems

Probabilistic breadth-first traversals (BPTs) are used in many network science and graph machine learning applications. In this paper, we are motivated by the application of BPTs in stochastic diffusion-based graph problems such as influence maximization. These applications heavily rely on BPTs to implement a Monte-Carlo sampling step for their approximations. Given the large sampling complexity, stochasticity of the diffusion process, and the inherent irregularity in real-world graph topologies, efficiently parallelizing these BPTs remains significantly challenging. In this paper, we present a new algorithm to fuse massive number of concurrently executing BPTs with random starts on the input graph. Our algorithm is designed to fuse BPTs by combining separate traversals into a unified frontier on distributed multi-GPU systems. To show the general applicability of the fused BPT technique, we have incorporated it into two state-of-the-art influence maximization parallel implementations (gIM and Ripples). Our experiments on up to 4K nodes of the OLCF Frontier supercomputer ($32,768$ GPUs and $196$K CPU cores) show strong scaling behavior, and that fused BPTs can improve the performance of these implementations up to 34$\times$ (for gIM) and ~360$\times$ (for Ripples).

cs.DC↗

DC: Depth Control on Quantum Classical Circuit

The growing prevalence of near-term intermediate-scale quantum (NISQ) systems has brought forth a heightened focus on the issue of circuit reliability. Several quantum computing activities, such as circuit design and multi-qubit mapping, are focused on enhancing reliability via the use of different optimization techniques. The optimization of quantum classical circuits has been the subject of substantial research, with a focus on techniques such as ancilla-qubit reuse and tactics aimed at minimizing circuit size and depth. Nevertheless, the reliability of bigger and more complex circuits remains a difficulty due to potential failures or the need for time-consuming compilation processes, despite the use of modern optimization strategies. This study presents a revolutionary Depth Control (DC) methodology that involves slicing and lowering the depth of conventional circuits. This strategy aims to improve the reliability and decrease the mapping costs associated with quantum hardware. DC provides reliable outcomes for circuits of indefinite size on any Noisy Intermediate-Scale Quantum (NISQ) system. The experimental findings demonstrate that the use of DC leads to a substantial improvement in the Probability of Success Threshold (PST), with an average increase of 11x compared to non-DC baselines. Furthermore, DC exhibits a notable superiority over the next best outcome by ensuring accurate outputs with a considerable margin. In addition, the utilization of Design Compiler (DC) enables the execution of mapping and routing optimizations inside a polynomial-time complexity, which represents an advancement compared to previously suggested methods that need exponential time.

quant-ph↗

Improving the Efficiency of OpenCL Kernels through Pipes

In an effort to lower the barrier to the adoption of FPGAs by a broader community, today major FPGA vendors offer compiler toolchains for OpenCL code. While using these toolchain allows porting existing code to FPGAs, ensuring performance portability across devices (i.e., CPUs, GPUs and FPGAs) is not a trivial task. This is in part due to the different hardware characteristics of these devices, including the nature of the hardware parallelism and the memory bandwidth they offer. In particular, global memory accesses are known to be one of the main performance bottlenecks for OpenCL kernels deployed on FPGA. In this paper, we investigate the use of pipes to improve memory bandwidth utilization and performance of OpenCL kernels running on FPGA. This is done by separating the global memory accesses from the computation, enabling better use of the load units required to access global memory. We perform experiments on a set of broadly used benchmark applications with various compute and memory access patterns. Our experiments, conducted on an Intel Arria GX board, show that the proposed method is effective in improving the memory bandwidth utilization of most kernels, particularly those exhibiting irregular memory access patterns. This, in turn, leads to performance improvements, in some cases significant.

cs.DC↗

Exploring Thread Coarsening on FPGA

Over the past few years, there has been an increased interest in including FPGAs in data centers and high-performance computing clusters along with GPUs and other accelerators. As a result, it has become increasingly important to have a unified, high-level programming interface for CPUs, GPUs and FPGAs. This has led to the development of compiler toolchains to deploy OpenCL code on FPGA. However, the fundamental architectural differences between GPUs and FPGAs have led to performance portability issues: it has been shown that OpenCL code optimized for GPU does not necessarily map well to FPGA, often requiring manual optimizations to improve performance. In this paper, we explore the use of thread coarsening - a compiler technique that consolidates the work of multiple threads into a single thread - on OpenCL code running on FPGA. While this optimization has been explored on CPU and GPU, the architectural features of FPGAs and the nature of the parallelism they offer lead to different performance considerations, making an analysis of thread coarsening on FPGA worthwhile. Our evaluation, performed on our microbenchmarks and on a set of applications from open-source benchmark suites, shows that thread coarsening can yield performance benefits (up to 3-4x speedups) to OpenCL code running on FPGA at a limited resource utilization cost.

cs.DC↗