FlashSAC: Fast and Stable Off-Policy Reinforcement Learning for High-Dimensional Robot Control
Simulation-based reinforcement learning (RL) is central for robotic control when expert demonstrations are unavailable. However, scaling RL to high-dimensional robots remains challenging. On-policy methods such as PPO are reliable but require large amounts of simulation because they discard past dat…