arXiv · 2023 · Preprint · Author with a top-venue record

COPR: Continual Learning Human Preference through Optimal Policy Regularization

Han Zhang, Lin Gui, Yuanzhao Zhai, Hui Wang, Yu Lei, Ruifeng Xu

Rendered by arXiv from the LaTeX source. If anything looks wrong, switch to the PDF.HTML · PDF · Open in a new tab ↗ · Close
Date
2023-10-24
Citations
2
arXiv
2310.15694
Cite
@misc{zhang2023copr,
  title = {COPR: Continual Learning Human Preference through Optimal Policy Regularization},
  author = {Han Zhang and Lin Gui and Yuanzhao Zhai and Hui Wang and Yu Lei and Ruifeng Xu},
  year = {2023},
  eprint = {2310.15694},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2310.15694},
}