AAAI 2026technical0 citations

CLEAR: Error Analysis via LLM-as-a-Judge Made Easy

Asaf Yehudai, Lilach Eden, Yotam Perlitz, Roy Bar-Haim, Michal Shmueli-Scheuer

Abstract

The evaluation of Large Language Models (LLMs) increasingly relies on other LLMs acting as judges. However, current evaluation paradigms typically yield a single score or ranking, answering which model is better but not why. While essential for benchmarking, these top-level scores obscure the specific, actionable reasons behind a model

BibTeX
@inproceedings{aaai2026_clearerroranalys,
  title = {CLEAR: Error Analysis via LLM-as-a-Judge Made Easy},
  author = {Asaf Yehudai and Lilach Eden and Yotam Perlitz and Roy Bar-Haim and Michal Shmueli-Scheuer},
  booktitle = {AAAI 2026},
  year = {2026}
}
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy · AAAI 2026