← Search

Forrest Sheng Bao

5 accepted papers

2025

FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs

NAACL 2025short

Summarization is one of the most common tasks performed by large language models (LLMs), especially in applications like Retrieval-Augmented Generation (RAG). However, existing evaluations of hallucinations in LLM-generated summaries, and evaluations of hallucination detection models both suffer fro…

2024

SummaCoz: A Dataset for Improving the Interpretability of Factual Consistency Detection for Summarization

EMNLP 2024finding

Summarization is an important application of Large Language Models (LLMs). When judging the quality of a summary, factual consistency holds a significant weight. Despite numerous efforts dedicated to building factual inconsistency detectors, the exploration of explanability remains limited among exi…

2023

DocAsRef: An Empirical Study on Repurposing Reference-based Summary Quality Metrics as Reference-free Metrics

EMNLP 2023short findings

Automated summary quality assessment falls into two categories: reference-based and reference-free. Reference-based metrics, historically deemed more accurate due to the additional information provided by human-written references, are limited by their reliance on human input. In this paper, we hypot…

Cited by 0SourceScholar
2022

PrefScore: Pairwise Preference Learning for Reference-free Summarization Quality Assessment

COLING 2022main

Evaluating machine-generated summaries without a human-written reference summary has been a need for a long time. Inspired by preference labeling in existing work of summarization evaluation, we propose to judge summary quality by learning the preference rank of summaries using the Bradley-Terry pow…