TY - RPRT TI - Simple and Scalable Strategies to Continually Pre-train Large Language Models AU - Adam Ibrahim AU - Benjamin Thérien AU - Kshitij Gupta AU - Mats L. Richter AU - Quentin Anthony AU - Timothée Lesort AU - Eugene Belilovsky AU - Irina Rish PY - 2024 UR - https://arxiv.org/abs/2403.08763 ID - 2403.08763 ER -