2025
Co-Eval: Augmenting LLM-based Evaluation with Machine Metrics
EMNLP 2025
Large language models (LLMs) are increasingly used as evaluators in natural language generation tasks, offering advantages in scalability and interpretability over traditional evaluation methods. However, existing LLM-based evaluations often suffer from biases and misalignment, particularly in domai