2026
Mean Flow Policy with Instantaneous Velocity Constraint for One-step Action Generation
ICLR 2026oral
Learning expressive and efficient policy functions is a promising direction in reinforcement learning (RL). While flow-based policies have recently proven effective in modeling complex action distributions with a fast deterministic sampling process, they still face a trade-off between expressiveness…