TY - RPRT TI - Layered gradient accumulation and modular pipeline parallelism: fast and efficient training of large language models AU - Joel Lamy-Poirier PY - 2021 UR - https://arxiv.org/abs/2106.02679 ID - 2106.02679 ER -