2024
Likelihood-based Mitigation of Evaluation Bias in Large Language Models
ACL 2024findings
Large Language Models (LLMs) are widely used to evaluate natural language generation tasks as automated metrics.However, the likelihood, a measure of LLM’s plausibility for a sentence, can vary due to superficial differences in sentences, such as word order and sentence structure.It is therefore pos…