← Search

Fahim Shahriar

1 accepted papers

2024

Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay Buffers

NeurIPS 2024poster

Modern deep policy gradient methods achieve effective performance on simulated robotic tasks, but they all require large replay buffers or expensive batch updates, or both, making them incompatible for real systems with resource-limited computers. We show that these methods fail catastrophically whe…