2024
Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation
EMNLP 2024main
As large language models (LLMs) evolve, evaluating their output reliably becomes increasingly difficult due to the high cost of human evaluation. To address this, we introduce FLAMe, a family of Foundational Large Autorater Models. FLAMe is trained on a diverse set of over 100 quality assessment tas…