← Search

Jordan Boyd-Graber

25 accepted papers

2025

A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users

EMNLP 2025

To assist users in complex tasks, LLMs generate plans: step-by-step instructions towards a goal. While alignment methods aim to ensure LLM plans are helpful, they train (RLHF) or evaluate (ChatbotArena) on what users prefer, assuming this reflects what helps them. We test this with Planorama: an int

Cited by 0SourcePDFScholar
2025

ADAPTIVE IE: Investigating the Complementarity of Human-AI Collaboration to Adaptively Extract Information on-the-fly

COLING 2025main

Information extraction (IE) needs vary over time, where a flexible information extraction (IE) system can be useful. Despite this, existing IE systems are either fully supervised, requiring expensive human annotations, or fully unsupervised, extracting information that often do not cater to user’s n…

Cited by 1SourcePDFScholar
2025

Discrepancy Detection at the Data Level: Toward Consistent Multilingual Question Answering

EMNLP 2025

Multilingual question answering (QA) systems must ensure factual consistency across languages, especially for objective queries such as What is jaundice?, while also accounting for cultural variation in subjective responses. We propose MIND, a user-in-the-loop fact-checking pipeline to detect factua

2024

Large Language Models Help Humans Verify Truthfulness – Except When They Are Convincingly Wrong

NAACL 2024long

Large Language Models (LLMs) are increasingly used for accessing information on the web. Their truthfulness and factuality are thus of great interest. To help users make the right decisions about the information they get, LLMs should not only provide information but also help users fact-check it. We…

Cited by 36SourcePDFScholar
2024

Pregnant Questions: The Importance of Pragmatic Awareness in Maternal Health Question Answering

NAACL 2024long

Questions posed by information-seeking users often contain implicit false or potentially harmful assumptions. In a high-risk domain such as maternal and infant health, a question-answering system must recognize these pragmatic constraints and go beyond simply answering user questions, examining them…

2024

Rapidly Piloting Real-time Linguistic Assistance for Simultaneous Interpreters with Untrained Bilingual Surrogates

COLING 2024main

Simultaneous interpretation is a cognitively taxing task, and even seasoned professionals benefit from real-time assistance. However, both recruiting professional interpreters and evaluating new assistance techniques are difficult. We present a novel, realistic simultaneous interpretation task that…

Cited by 0SourcePDFScholar
2023

Program Chairs’ Report on Peer Review at ACL 2023

ACL 2023long

We present a summary of the efforts to improve conference peer review that were implemented at ACL’23. This includes work with the goal of improving review quality, clearer workflow and decision support for the area chairs, as well as our efforts to improve paper-reviewer matching for various kinds…

2022

Adapting Coreference Resolution Models through Active Learning

ACL 2022long

Neural coreference resolution models trained on one dataset may not transfer to new, low-resource domains. Active learning mitigates this problem by sampling a small subset of data for annotators to label. While active learning is well-defined for classification tasks, its application to coreference…

2022

Automatic Song Translation for Tonal Languages

ACL 2022findings

This paper develops automatic song translation (AST) for tonal languages and addresses the unique challenge of aligning words’ tones with melody of a song in addition to conveying the original meaning. We propose three criteria for effective AST—preserving meaning, singability and intelligibility—an…

Cited by 16SourcePDFScholar
2022

Cheater’s Bowl: Human vs. Computer Search Strategies for Open-Domain QA

EMNLP 2022finding

For humans and computers, the first step in answering an open-domain question is retrieving a set of relevant documents from a large corpus. However, the strategies that computers use fundamentally differ from those of humans. To better understand these differences, we design a gamified interface fo…

2022

Match the Script, Adapt if Multilingual: Analyzing the Effect of Multilingual Pretraining on Cross-lingual Transferability

ACL 2022long

Pretrained multilingual models enable zero-shot learning even for unseen languages, and that performance can be further improved via adaptation prior to finetuning. However, it is unclear how the number of pretraining languages influences a model’s zero-shot learning for languages unseen during pret…

2022

Re-Examining Calibration: The Case of Question Answering

EMNLP 2022finding

For users to trust model predictions, they need to understand model outputs, particularly their confidence — calibration aims to adjust (calibrate) models’ confidence to match expected accuracy. We argue that the traditional calibration evaluation does not promote effective calibrations: for example…

2022

SimQA: Detecting Simultaneous MT Errors through Word-by-Word Question Answering

EMNLP 2022main

Detractors of neural machine translation admit that while its translations are fluent, it sometimes gets key facts wrong. This is particularly important in simultaneous interpretation where translations have to be provided as fast as possible: before a sentence is complete. Yet, evaluations of simul…

Cited by 7SourcePDFScholar
2021

Adapting Entities across Languages and Cultures

EMNLP 2021finding

How would you explain Bill Gates to a German? He is associated with founding a company in the United States, so perhaps the German founder Carl Benz could stand in for Gates in those contexts. This type of translation is called adaptation in the translation community. Until now, this task has not be…

Cited by 17SourcePDFScholar
2021

Distantly-Supervised Dense Retrieval Enables Open-Domain Question Answering without Evidence Annotation

EMNLP 2021main

Open-domain question answering answers a question based on evidence retrieved from a large corpus. State-of-the-art neural approaches require intermediate evidence annotations for training. However, such intermediate annotations are expensive, and methods that rely on them cannot transfer to the mor…

2021

Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?

ACL 2021long

Leaderboards are widely used in NLP and push the field forward. While leaderboards are a straightforward ranking of NLP models, this simplicity can mask nuances in evaluation items (examples) and subjects (NLP models). Rather than replace leaderboards, we advocate a re-imagining so that they better…

2021

Fool Me Twice: Entailment from Wikipedia Gamification

NAACL 2021long

We release FoolMeTwice (FM2 for short), a large dataset of challenging entailment pairs collected through a fun multi-player game. Gamification encourages adversarial examples, drastically lowering the number of examples that can be solved using “shortcuts” compared to other popular entailment datas…

2021

Multi-Step Reasoning Over Unstructured Text with Beam Dense Retrieval

NAACL 2021long

Complex question answering often requires finding a reasoning chain that consists of multiple evidence pieces. Current approaches incorporate the strengths of structured knowledge and unstructured text, assuming text corpora is semi-structured. Building on dense retrieval methods, we propose a new m…

2021

Toward Deconfounding the Effect of Entity Demographics for Question Answering Accuracy

EMNLP 2021main

The goal of question answering (QA) is to answer _any_ question. However, major QA datasets have skewed distributions over gender, profession, and nationality. Despite that skew, an analysis of model accuracy reveals little evidence that accuracy is lower for people based on gender or nationality; i…

Cited by 6SourcePDFScholar
2021

What’s in a Name? Answer Equivalence For Open-Domain Question Answering

EMNLP 2021main

A flaw in QA evaluation is that annotations often only provide one gold answer. Thus, model predictions semantically equivalent to the answer but superficially different are considered incorrect. This work explores mining alias entities from knowledge bases and using them as additional gold answers…

2017

The Amazing Mysteries of the Gutter: Drawing Inferences Between Panels in Comic Book Narratives

CVPR 2017spotlight

Visual narrative is often a combination of explicit information and judicious omissions, relying on the viewer to supply missing details. In comics, most movements in time and space are hidden in the "gutters" between panels. To follow the story, readers logically connect panels together by inferrin…

Cited by 132PDFScholar
2016

Opponent Modeling in Deep Reinforcement Learning

ICML 2016poster

Opponent modeling is necessary in multi-agent settings where secondary agents with competing goals also adapt their strategies, yet it remains challenging because of strategies’ complex interaction and the non-stationary nature. Most previous work focuses on developing probabilistic models or parame…

2015

Paired-Dual Learning for Fast Training of Latent Variable Hinge-Loss MRFs

ICML 2015poster

Latent variables allow probabilistic graphical models to capture nuance and structure in important domains such as network science, natural language processing, and computer vision. Naive approaches to learning such complex models can be prohibitively expensive—because they require repeated inferenc…

Cited by 18SourcePDFScholar