2025
GroUSE: A Benchmark to Evaluate Evaluators in Grounded Question Answering
COLING 2025main
Retrieval-Augmented Generation (RAG) has emerged as a common paradigm to use Large Language Models (LLMs) alongside private and up-to-date knowledge bases. In this work, we address the challenges of using LLM-as-a-Judge when evaluating grounded answers generated by RAG systems. To assess the calibra…