NeurIPS · 2024 · Journal article · Top venue

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers

Gautham Vasan, Mohamed Elsayed, Alireza Azimi, Junjie He, Fahim Shariar, Colin Bellinger, Martha White, A. Rupam Mahmood

PDF from arXiv.HTML · PDF · Open in a new tab ↗ · Close
Published in
Neural Information Processing Systems
Date
2024-11-22
Citations
25
arXiv
2411.15370
Cite
@article{vasan2024deep,
  title = {Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers},
  author = {Gautham Vasan and Mohamed Elsayed and Alireza Azimi and Junjie He and Fahim Shariar and Colin Bellinger and Martha White and A. Rupam Mahmood},
  year = {2024},
  journal = {Neural Information Processing Systems},
  eprint = {2411.15370},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2411.15370},
}