2026
SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models
ICML 2026poster
Evaluating large language models (LLMs) for software engineering has been limited by narrow task coverage, language bias, and insufficient alignment with real-world developer workflows. Existing benchmarks often focus on algorithmic problems or Python-centric bug fixing, leaving critical dimensions …