TY - RPRT TI - Continual Pre-Training of Large Language Models: How to (re)warm your model? AU - Kshitij Gupta AU - Benjamin Thérien AU - Adam Ibrahim AU - Mats L. Richter AU - Quentin Anthony AU - Eugene Belilovsky AU - Irina Rish AU - Timothée Lesort PY - 2023 UR - https://arxiv.org/abs/2308.04014 ID - 2308.04014 ER -