2026
Learning More with Less: A Dynamic Dual-Level Down-Sampling Framework for Efficient Policy Optimization
ICLR 2026poster
Critic-free methods like GRPO reduce memory demands by estimating advantages from multiple rollouts but tend to converge slowly, as critical learning signals are diluted by an abundance of uninformative samples and tokens. To tackle this challenge, we propose the **Dynamic Dual-Level Down-Sampling (…