2026
Talk2Code: A Multi-Turn Interaction Benchmark with Dual-Track Evaluation for Code Generation
AAAI 2026technical
While large language models (LLMs) have demonstrated strong capabilities in code generation, current benchmarks primarily focus on single-turn scenarios, neglecting the complexity of multi-turn interactions and user diversity. To address this gap, we introduce Talk2Code, the first benchmark for user