NeurIPS · 2024 · Journal article

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers

G. Vasan, Mohamed Elsayed, Alireza Azimi, Jiamin He, Fahim Shariar, Colin Bellinger, Martha White, A. Mahmood

Published in
Neural Information Processing Systems
Date
2024-11-22
Citations
25
arXiv
2411.15370
Cite
@article{vasan2024deep,
  title = {Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers},
  author = {G. Vasan and Mohamed Elsayed and Alireza Azimi and Jiamin He and Fahim Shariar and Colin Bellinger and Martha White and A. Mahmood},
  year = {2024},
  journal = {Neural Information Processing Systems},
  eprint = {2411.15370},
  archivePrefix = {arXiv},
  url = {https://arxiv.org/abs/2411.15370},
}