Trans. Mach. Learn. Res. · 2024 · Journal article

Simple and Scalable Strategies to Continually Pre-train Large Language Models

Adam Ibrahim, Benjamin Therien, Kshitij Gupta, Mats L. Richter, Quentin Anthony, Timothée Lesort, Eugene Belilovsky, Irina Rish

Date
2024-03-13
Citations
130
arXiv
2403.08763
Cite
@article{ibrahim2024simple,
  title = {Simple and Scalable Strategies to Continually Pre-train Large Language Models},
  author = {Adam Ibrahim and Benjamin Therien and Kshitij Gupta and Mats L. Richter and Quentin Anthony and Timothée Lesort and Eugene Belilovsky and Irina Rish},
  year = {2024},
  journal = {Trans. Mach. Learn. Res.},
  eprint = {2403.08763},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2403.08763},
}