2026
Latent Policy Steering through One-Step Flow Policies
RSS 2026poster
Offline reinforcement learning (RL) should be ideal for robotics, allowing learning from dataset without risky exploration. Yet, offline RL’s performance often hinges on a brittle trade-off between (1) return maximization, which can push policies outside dataset support, and (2) behavioral constrain…