arXiv · 2011.01349
Distributed Machine Learning for Computational Engineering using MPI
Abstract
We propose a framework for training neural networks that are coupled with partial differential equations (PDEs) in a parallel computing environment. Unlike most distributed computing frameworks for deep neural networks, our focus is to parallelize both numerical solvers and deep neural networks in forward and adjoint computations. Our parallel computing model views data communication as a node in the computational graph for numerical simulations. The advantage of our model is that data communication and computing are cleanly separated and thus provide better flexibility, modularity, and testability. We demonstrate using various large-scale problems that we can achieve substantial acceleration by using parallel solvers for PDEs in training deep neural networks that are coupled with PDEs.
Explore related subjects
Keep this discovery
Kailai Xu, Weiqiang Zhu, Eric Darve. 2020-11-02. Distributed Machine Learning for Computational Engineering using MPI. https://arxiv.org/abs/2011.01349
Cite the original work for its findings. Save a collection to share your selection of sources.