arXiv · 2025 · Preprint

Mitigating Forgetting Between Supervised and Reinforcement Learning Yields Stronger Reasoners

Xiangchi Yuan, Xiang Chen, Tong Yu, Dachuan Shi, Can Jin, Wenke Lee, Saayan Mitra

Rendered by arXiv from the LaTeX source. If anything looks wrong, switch to the PDF.HTML · PDF · Open in a new tab ↗ · Close
Date
2025-10-06
Citations
18
arXiv
2510.04454
Cite
@misc{yuan2025mitigating,
  title = {Mitigating Forgetting Between Supervised and Reinforcement Learning Yields Stronger Reasoners},
  author = {Xiangchi Yuan and Xiang Chen and Tong Yu and Dachuan Shi and Can Jin and Wenke Lee and Saayan Mitra},
  year = {2025},
  eprint = {2510.04454},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2510.04454},
}