Auto-PRE: An Automatic and Cost-Efficient Peer-Review Framework for Language Generation Evaluation
The rapid development of large language models (LLMs) has highlighted the need for efficient and reliable methods to evaluate their performance. Traditional evaluation methods often face challenges like high costs, limited task formats, dependence on human references, and systematic biases. To addre