2025
Unbiased Evaluation of Large Language Models from a Causal Perspective
ICML 2025poster
Benchmark contamination has become a significant concern in the LLM evaluation community. Previous Agents-as-an-Evaluator address this issue by involving agents in the generation of questions. Despite their success, the biases in Agents-as-an-Evaluator methods remain largely unexplored. In this pape…