← Search

Joel Tetreault

16 accepted papers

2025

CEHA: A Dataset of Conflict Events in the Horn of Africa

COLING 2025main

Natural Language Processing (NLP) of news articles can play an important role in understanding the dynamics and causes of violent conflict. Despite the availability of datasets categorizing various conflict events, the existing labels often do not cover all of the fine-grained violent conflict event…

2025

Uchaguzi-2022: A Dataset of Citizen Reports on the 2022 Kenyan Election

COLING 2025main

Online reporting platforms have enabled citizens around the world to collectively share their opinions and report in real time on events impacting their local communities. Systematically organizing (e.g., categorizing by attributes) and geotagging large amounts of crowdsourced information is crucial…

Cited by 0SourcePDFScholar
2023

A New Task and Dataset on Detecting Attacks on Human Rights Defenders

ACL 2023findings

The ability to conduct retrospective analyses of attacks on human rights defenders over time and by location is important for humanitarian organizations to better understand historical or ongoing human rights violations and thus better manage the global impact of such events. We hypothesize that NLP…

2023

BUMP: A Benchmark of Unfaithful Minimal Pairs for Meta-Evaluation of Faithfulness Metrics

ACL 2023long

The proliferation of automatic faithfulness metrics for summarization has produced a need for benchmarks to evaluate them. While existing benchmarks measure the correlation with human judgements of faithfulness on model-generated summaries, they are insufficient for diagnosing whether metrics are: 1…

2022

An Exploration of Post-Editing Effectiveness in Text Summarization

NAACL 2022long

Automatic summarization methods are efficient but can suffer from low quality. In comparison, manual summarization is expensive but produces higher quality. Can humans and AI collaborate to improve summarization performance? In similar text generation tasks (e.g., machine translation), human-AI coll…

2022

CrisisLTLSum: A Benchmark for Local Crisis Event Timeline Extraction and Summarization

EMNLP 2022finding

Social media has increasingly played a key role in emergency response: first responders can use public posts to better react to ongoing crisis events and deploy the necessary resources where they are most needed. Timeline extraction and abstractive summarization are critical technical tasks to lever…

2022

Mapping the Design Space of Human-AI Interaction in Text Summarization

NAACL 2022long

Automatic text summarization systems commonly involve humans for preparing data or evaluating model performance, yet, there lacks a systematic understanding of humans’ roles, experience, and needs when interacting with or being assisted by AI. From a human-centered perspective, we map the design opp…

Cited by 34SourcePDFScholar
2022

XLTime: A Cross-Lingual Knowledge Transfer Framework for Temporal Expression Extraction

NAACL 2022findings

Temporal Expression Extraction (TEE) is essential for understanding time in natural language. It has applications in Natural Language Processing (NLP) tasks such as question answering, information retrieval, and causal inference. To date, work in this area has mostly focused on English as there is a…

2021

Evaluating the Evaluation Metrics for Style Transfer: A Case Study in Multilingual Formality Transfer

EMNLP 2021main

While the field of style transfer (ST) has been growing rapidly, it has been hampered by a lack of standardized practices for automatic evaluation. In this paper, we evaluate leading automatic metrics on the oft-researched task of formality style transfer. Unlike previous evaluations, which focus so…

2021

Journalistic Guidelines Aware News Image Captioning

EMNLP 2021main

The task of news article image captioning aims to generate descriptive and informative captions for news article images. Unlike conventional image captions that simply describe the content of the image in general terms, news image captions follow journalistic guidelines and rely heavily on named ent…

2021

Olá, Bonjour, Salve! XFORMAL: A Benchmark for Multilingual Formality Style Transfer

NAACL 2021long

We take the first step towards multilingual style transfer by creating and releasing XFORMAL, a benchmark of multiple formal reformulations of informal text in Brazilian Portuguese, French, and Italian. Results on XFORMAL suggest that state-of-the-art style transfer approaches perform close to simpl…

2020

Multimodal Categorization of Crisis Events in Social Media

CVPR 2020poster

Recent developments in image classification and natural language processing, coupled with the rapid growth in social media usage, have enabled fundamental advances in detecting breaking events around the world in real-time. Emergency response is one such area that stands to gain from these advances.…

Cited by 138PDFScholar
2020

Rhetoric, Logic, and Dialectic: Advancing Theory-based Argument Quality Assessment in Natural Language Processing

COLING 2020main

Though preceding work in computational argument quality (AQ) mostly focuses on assessing overall AQ, researchers agree that writers would benefit from feedback targeting individual dimensions of argumentation theory. However, a large-scale theory-based corpus and corresponding computational models a…

2020

The ApposCorpus: a new multilingual, multi-domain dataset for factual appositive generation

COLING 2020main

News articles, image captions, product reviews and many other texts mention people and organizations whose name recognition could vary for different audiences. In such cases, background information about the named entities could be provided in the form of an appositive noun phrase, either written by…

2016

TGIF: A New Dataset and Benchmark on Animated GIF Description

CVPR 2016spotlight

With the recent popularity of animated GIFs on social media, there is need for ways to index them with rich metadata. To advance research on animated GIF understanding, we collected a new dataset, Tumblr GIF (TGIF), with 100K animated GIFs from Tumblr and 120K natural language descriptions obtained…

Cited by 330PDFcodeScholar