← Search

Alyssa Hwang

2 accepted papers

2024

FanOutQA: A Multi-Hop, Multi-Document Question Answering Benchmark for Large Language Models

ACL 2024short

One type of question that is commonly found in day-to-day scenarios is “fan-out” questions, complex multi-hop, multi-document reasoning questions that require finding information about a large number of entities. However, there exist few resources to evaluate this type of question-answering capabili…

2024

RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors

ACL 2024long

Many commercial and open-source models claim to detect machine-generated text with extremely high accuracy (99% or more). However, very few of these detectors are evaluated on shared benchmark datasets and even when they are, the datasets used for evaluation are insufficiently challenging—lacking va…