SearcharxivSearch

arXiv · 1907.02290

Accelerated FDPS --- Algorithms to Use Accelerators with FDPS

Abstract

In this paper, we describe the algorithms we implemented in FDPS to make efficient use of accelerator hardware such as GPGPUs. We have developed FDPS to make it possible for many researchers to develop their own high-performance parallel particle-based simulation programs without spending large amount of time for parallelization and performance tuning. The basic idea of FDPS is to provide a high-performance implementation of parallel algorithms for particle-based simulations in a "generic" form, so that researchers can define their own particle data structure and interparticle interaction functions and supply them to FDPS. FDPS compiled with user-supplied data type and interaction function provides all necessary functions for parallelization, and using those functions researchers can write their programs as though they are writing simple non-parallel program. It has been possible to use accelerators with FDPS, by writing the interaction function that uses the accelerator. However, the efficiency was limited by the latency and bandwidth of communication between the CPU and the accelerator and also by the mismatch between the available degree of parallelism of the interaction function and that of the hardware parallelism. We have modified the interface of user-provided interaction function so that accelerators are more efficiently used. We also implemented new techniques which reduce the amount of work on the side of CPU and amount of communication between CPU and accelerators. We have measured the performance of N-body simulations on a systems with NVIDIA Volta GPGPU using FDPS and the achieved performance is around 27 \% of the theoretical peak limit. We have constructed a detailed performance model, and found that the current implementation can achieve good performance on systems with much smaller memory and communication bandwidth.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Masaki Iwasawa, Daisuke Namekata, Keigo Nitadori, Kentaro Nomura, Long Wang, Miyuki Tsubouchi, Junichiro Makino. 2019-07-04. Accelerated FDPS --- Algorithms to Use Accelerators with FDPS. https://doi.org/10.1093/pasj%2Fpsz133

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

The EDD Radio Astronomy Backend Framework

Modern digital radio astronomy receivers produce increasingly wide-bandwidth, high bit-rate data streams that necessitate the development of flexible, scalable, and maintainable backend processing and recording systems. Historically, such backend instrumentation has been tightly coupled to telescope observing modes, limiting reuse between observatories and science cases. We present the Effelsberg Direct Digitisation (EDD) backend framework, a software-defined architecture for constructing real-time radio astronomy backends on commodity off-the-shelf computing infrastructure. We describe its design, implementation, supported observing modes, and operational deployments. EDD separates a common core framework from plugin-provided observing capabilities. The core provides orchestration, telescope interfaces, pipeline lifecycle management, monitoring, and deployment tooling, while plugins implement processing pipelines for specific observing modes. The framework is designed to support both single-dish and interferometric instruments through site-specific configuration and plugin selection. EDD currently supports spectroscopy and spectropolarimetry, pulsar timing and searching, baseband recording, very long baseline interferometry, correlation, and beamforming. Operational deployments include the Effelsberg 100-m telescope, the SKA-MPI prototype dish, the Thai National Radio Telescope, and the ARGOS interferometric prototype array. By separating common services, observing-mode plugins, and site-specific configuration, it allows backend capabilities to be deployed across heterogeneous telescope environments and provides a community resource for broadband radio astronomy instrumentation.

astro-ph.IM

Bayesian Superiority in On/Off analysis

We present a detailed comparison of Bayesian criteria with three non-informative priors - flat, Jeffreys, and scale-invariant - for testing a signal against an unknown background and compare them with the classical frequentist Li-Ma approach in the On/Off problem. We perform Monte Carlo simulations for various background levels and evaluate the Li-Ma and Bayesian criteria by their Type I error rates. We then simulate a nonzero signal and compare the criteria in terms of Type II error rates. We find that the Bayesian criterion with the Jeffreys prior yields lower Type I and Type II error rates than the Li-Ma criterion. In addition, we show that the Bayesian criteria are more robust than the Li-Ma criterion when the background distribution is overdispersed relative to the Poisson distribution.

astro-ph.IM

An RFSoC-based Backend and Timing System for the Balloon-borne Very Long Baseline Interferometry Experiment

We present the design and performance characterization of the digital backend and precision-timing system for the Balloon-borne Very Long Baseline Interferometry Experiment (BVEX), a pathfinder for high-frequency stratospheric VLBI at 22 GHz. The backend uses one of the four 14-bit analog-to-digital converter inputs on an AMD-Xilinx RFSoC 4x2. Although the converters support sampling rates up to 5 GSPS, the flight configuration digitizes the 2-4 GHz intermediate frequency at 4.096 GSPS. CASPER firmware provides both a high-resolution spectrometer for pointing and receiver verification, and a VLBI acquisition chain with two-bit requantization that records at a rate of about 8.2 Gbps. The timestamped data packets are sent over 100 Gigabit Ethernet (GbE) to a 16 TB NVMe array in a storage computer that draws approximately 70-80 W. The timing chain uses a Rakon oven-controlled crystal oscillator as a timing reference while a time-interval counter measures its drift relative to a GPS reference with approximately 60 ps resolution. This is the first deployment of an RFSoC-based VLBI backend and precision-timing system on a stratospheric balloon. Ground tests validated the backend, spectrometer, and timing chain. The August 2025 CSA STRATOS flight ended before reaching the target float altitude because of a balloon failure, and as a result no science observations were obtained. For the planned 2027 reflight, we are developing a conduction-cooled data storage computer with 24 TB of NVMe capacity and a direct data path from the 100 GbE interface to the NVMe array.

astro-ph.IM