Trans. Mach. Learn. Res. · 2024 · Journal article

Simple and Scalable Strategies to Continually Pre-train Large Language Models

Adam Ibrahim, Benjamin Therien, Kshitij Gupta, Mats L. Richter, Quentin Anthony, Timothée Lesort, Eugene Belilovsky, Irina Rish

Rendered by arXiv from the LaTeX source. If anything looks wrong, switch to the PDF.HTML · PDF · Open in a new tab ↗ · Close
Date
2024-03-13
Citations
130
arXiv
2403.08763
Cite
@article{ibrahim2024simple,
  title = {Simple and Scalable Strategies to Continually Pre-train Large Language Models},
  author = {Adam Ibrahim and Benjamin Therien and Kshitij Gupta and Mats L. Richter and Quentin Anthony and Timothée Lesort and Eugene Belilovsky and Irina Rish},
  year = {2024},
  journal = {Trans. Mach. Learn. Res.},
  eprint = {2403.08763},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2403.08763},
}