2025
CAMIEval: Enhancing NLG Evaluation through Multidimensional Comparative Instruction-Following Analysis
NAACL 2025long
With the rapid development of large language models (LLMs), due to their strong performance across various fields, LLM-based evaluation methods (LLM-as-a-Judge) have become widely used in natural language generation (NLG) evaluation. However, these methods encounter the following challenges: (1) dis…