← Search

Ruixuan Tu

3 accepted papers

2025

FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs

NAACL 2025short

Summarization is one of the most common tasks performed by large language models (LLMs), especially in applications like Retrieval-Augmented Generation (RAG). However, existing evaluations of hallucinations in LLM-generated summaries, and evaluations of hallucination detection models both suffer fro…

2023

DocAsRef: An Empirical Study on Repurposing Reference-based Summary Quality Metrics as Reference-free Metrics

EMNLP 2023short findings

Automated summary quality assessment falls into two categories: reference-based and reference-free. Reference-based metrics, historically deemed more accurate due to the additional information provided by human-written references, are limited by their reliance on human input. In this paper, we hypot…

Cited by 0SourceScholar