2026
GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
ICLR 2026oral
Large language models (LLMs) are increasingly adapted to downstream tasks via reinforcement learning (RL) methods like Group Relative Policy Optimization (GRPO), which often require thousands of rollouts to learn new tasks. We argue that the interpretable nature of language often provides a much ric…