Q-learning with Long-term Action-space Shaping to Model Complex Behavior for Autonomous Lane Changes
In autonomous driving applications, reinforcement learning agents often have to perform complex behavior, which can translate into optimizing multiple objectives while following certain rules. Encoding traffic rules and desires such as safety and comfort via classical methods based on reward shaping…