← Search

Liam Dugan

7 accepted papers

2024

FanOutQA: A Multi-Hop, Multi-Document Question Answering Benchmark for Large Language Models

ACL 2024short

One type of question that is commonly found in day-to-day scenarios is “fan-out” questions, complex multi-hop, multi-document reasoning questions that require finding information about a large number of entities. However, there exist few resources to evaluate this type of question-answering capabili…

2024

MiRAGeNews: Multimodal Realistic AI-Generated News Detection

EMNLP 2024finding

The proliferation of inflammatory or misleading “fake” news content has become increasingly common in recent years. Simultaneously, it has become easier than ever to use AI tools to generate photorealistic images depicting any scene imaginable. Combining these two—AI-generated fake news content—is p…

2024

RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors

ACL 2024long

Many commercial and open-source models claim to detect machine-generated text with extremely high accuracy (99% or more). However, very few of these detectors are evaluated on shared benchmark datasets and even when they are, the datasets used for evaluation are insufficiently challenging—lacking va…

2024

ReDel: A Toolkit for LLM-Powered Recursive Multi-Agent Systems

EMNLP 2024system demonstrations

Recently, there has been increasing interest in using Large Language Models (LLMs) to construct complex multi-agent systems to perform tasks such as compiling literature reviews, drafting consumer reports, and planning vacations. Many tools and libraries exist for helping create such systems, howeve…

2023

Real or Fake Text?: Investigating Human Ability to Detect Boundaries between Human-Written and Machine-Generated Text

AAAI 2023technical

As text generated by large language models proliferates, it becomes vital to understand how humans engage with such text, and whether or not they are able to detect when the text they are reading did not originate with a human writer. Prior work on human detection of generated text focuses on the ca…

2022

A Feasibility Study of Answer-Agnostic Question Generation for Education

ACL 2022findings

We conduct a feasibility study into the applicability of answer-agnostic question generation models to textbook passages. We show that a significant portion of errors in such systems arise from asking irrelevant or un-interpretable questions and that such errors can be ameliorated by providing summa…

2022

The Case for a Single Model that can Both Generate Continuations and Fill-in-the-Blank

NAACL 2022findings

The task of inserting text into a specified position in a passage, known as fill in the blank (FitB), is useful for a variety of applications where writers interact with a natural language generation (NLG) system to craft text. While previous work has tackled this problem with models trained specifi…

Cited by 3SourcePDFScholar