2026
Natural Language Actor–Critic Is Bilevel: Learning to Reason with Textual Feedback
ICML 2026poster
Reinforcement learning with verifiable rewards can improve LLM reasoning, but learning is sample-inefficient under sparse terminal rewards. Prior work mitigates this by adding natural language critiques, yet it typically treats critique generation as fixed or auxiliary, so correct-sounding feedback …