2025
When Benchmarks Talk: Re-Evaluating Code LLMs with Interactive Feedback
ACL 2025finding
Programming with a coding assistant is a fundamentally interactive process, yet existing static benchmarks fail to capture key features of model-user collaboration. We introduce an interactive evaluation pipeline to examine how LLMs incorporate different types of feedback in a collaborative setting,…