arXiv · 1508.01847
Big Data Analytics on Traditional HPC Infrastructure Using Two-Level Storage
Abstract
Data-intensive computing has become one of the major workloads on traditional high-performance computing (HPC) clusters. Currently, deploying data-intensive computing software framework on HPC clusters still faces performance and scalability issues. In this paper, we develop a new two-level storage system by integrating Tachyon, an in-memory file system with OrangeFS, a parallel file system. We model the I/O throughputs of four storage structures: HDFS, OrangeFS, Tachyon and two-level storage. We conduct computational experiments to characterize I/O throughput behavior of two-level storage and compare its performance to that of HDFS and OrangeFS, using TeraSort benchmark. Theoretical models and experimental tests both show that the two-level storage system can increase the aggregate I/O throughputs. This work lays a solid foundation for future work in designing and building HPC systems that can provide a better support on I/O intensive workloads with preserving existing computing resources.
Explore related subjects
Keep this discovery
Pengfei Xuan, Jeffrey Denton, Rong Ge, Pradip K. Srimani, Feng Luo. 2015-08-08. Big Data Analytics on Traditional HPC Infrastructure Using Two-Level Storage. https://doi.org/10.1145/2831244.2831253
Cite the original work for its findings. Save a collection to share your selection of sources.