← Search

Sarah Kwan

1 accepted papers

2023

Evaluating the Factual Consistency of Large Language Models Through News Summarization

ACL 2023findings

While large language models (LLMs) have proven to be effective on a large variety of tasks, they are also known to hallucinate information. To measure whether an LLM prefers factually consistent continuations of its input, we propose a new benchmark called FIB (Factual Inconsistency Benchmark) that…