2025
Action abstractions for amortized sampling
Oussama Boussif, Lena Nehale Ezzine, Joseph D Viviano, Michał Koziarski, Moksh Jain, Nikolay Malkin +3
ICLR 2025poster
As trajectories sampled by policies used by reinforcement learning (RL) and generative flow networks (GFlowNets) grow longer, credit assignment and exploration become more challenging, and the long planning horizon hinders mode discovery and generalization. The challenge is particularly pronounced i…