2025
CodeJudge-Eval: Can Large Language Models be Good Judges in Code Understanding?
COLING 2025main
Recent advancements in large language models (LLMs) have showcased impressive code generation capabilities, primarily evaluated through language-to-code benchmarks. However, these benchmarks may not fully capture a model’s code understanding abilities. We introduce CodeJudge-Eval (CJ-Eval), a novel…