2025
ConvCodeWorld: Benchmarking Conversational Code Generation in Reproducible Feedback Environments
ICLR 2025poster
Large language models (LLMs) have proven invaluable for code generation, particularly in interactive settings. However, existing code generation benchmarks fail to capture the diverse feedback encountered in multi-turn interactions, limiting our ability to evaluate LLMs in these contexts. To address…