2023
FELM: Benchmarking Factuality Evaluation of Large Language Models
NeurIPS 2023poster
Assessing factuality of text generated by large language models (LLMs) is an emerging yet crucial research area, aimed at alerting users to potential errors and guiding the development of more reliable LLMs. Nonetheless, the evaluators assessing factuality necessitate suitable evaluation themselves…