Peak-Return Greedy Slicing: Subtrajectory Selection for Transformer-based Offline RL
Offline reinforcement learning enables policy learning solely from fixed datasets, without costly or risky environment interactions, making it highly valuable for real-world applications. While Transformer-based approaches have recently demonstrated strong sequence modeling capabilities, they typica…