← Search

Ana Marasović

2 accepted papers

2021

Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus

EMNLP 2021main

Large language models have led to remarkable progress on many NLP tasks, and researchers are turning to ever-larger text corpora to train them. Some of the largest corpora available are made by scraping significant portions of the internet, and are frequently introduced with only minimal documentati…

2021

Measuring Association Between Labels and Free-Text Rationales

EMNLP 2021main

In interpretable NLP, we require faithful rationales that reflect the model’s decision-making process for an explained instance. While prior work focuses on extractive rationales (a subset of the input words), we investigate their less-studied counterpart: free-text natural language rationales. We d…