2024
DiaHalu: A Dialogue-level Hallucination Evaluation Benchmark for Large Language Models
EMNLP 2024finding
Though large language models (LLMs) achieve significant success in recent years, the hallucination issue remains a challenge, and numerous benchmarks are proposed for hallucination detection. Nevertheless, some of these benchmarks are not naturally generated by LLMs but are intentionally induced. Al…