← Search

Sidhaarth Murali

1 accepted papers

2026

Natural Language Actor–Critic Is Bilevel: Learning to Reason with Textual Feedback

ICML 2026poster

Reinforcement learning with verifiable rewards can improve LLM reasoning, but learning is sample-inefficient under sparse terminal rewards. Prior work mitigates this by adding natural language critiques, yet it typically treats critique generation as fixed or auxiliary, so correct-sounding feedback …

Cited by 0SourceScholar