← Search

Dhruv Gautam

1 accepted papers

2025

RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code

ICLR 2025poster

Recent advances in language model (LM) agents and function calling have enabled autonomous, feedback-driven systems to solve problems across various digital domains. To better understand the unique limitations of LM agents, we introduce RefactorBench, a benchmark consisting of 100 large handcrafted…