← Search

Ana Marasovic

11 accepted papers

2025

BriefMe: A Legal NLP Benchmark for Assisting with Legal Briefs

ACL 2025finding

A core part of legal work that has been underexplored in Legal NLP is the writing and editing of legal briefs. This requires not only a thorough understanding of the law of a jurisdiction, from judgments to statutes, but also the ability to make new arguments to try to expand the law in a new direct…

Cited by 0SourcePDFScholar
2025

Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps

EMNLP 2025

When prompted to think step-by-step, language models (LMs) produce a chain of thought (CoT), a sequence of reasoning steps that the model supposedly used to produce its prediction. Despite much work on CoT prompting, it is unclear if reasoning verbalized in a CoT is faithful to the models’ parametri

Cited by 0SourcePDFScholar
2024

On Evaluating Explanation Utility for Human-AI Decision Making in NLP

EMNLP 2024finding

Is explainability a false promise? This debate has emerged from the insufficient evidence that explanations help people in situations they are introduced for. More human-centered, application-grounded evaluations of explanations are needed to settle this. Yet, with no established guidelines for such…

2024

Whispers of Doubt Amidst Echoes of Triumph in NLP Robustness

NAACL 2024long

*Do larger and more performant models resolve NLP’s longstanding robustness issues?* We investigate this question using over 20 models of different sizes spanning different architectural choices and pretraining objectives. We conduct evaluations using (a) out-of-domain and challenge test sets, (b) b…

2023

Do Androids Laugh at Electric Sheep? Humor “Understanding” Benchmarks from The New Yorker Caption Contest

ACL 2023long

Large neural networks can now generate jokes, but do they really “understand” humor? We challenge AI models with three tasks derived from the New Yorker Cartoon Caption Contest: matching a joke to a cartoon, identifying a winning caption, and explaining why a winning caption is funny. These tasks en…

2022

CONDAQA: A Contrastive Reading Comprehension Dataset for Reasoning about Negation

EMNLP 2022main

The full power of human language-based communication cannot be realized without negation. All human languages have some form of negation. Despite this, negation remains a challenging phenomenon for current natural language understanding systems. To facilitate the future development of models that ca…

2022

Does Self-Rationalization Improve Robustness to Spurious Correlations?

EMNLP 2022main

Rationalization is fundamental to human reasoning and learning. NLP models trained to produce rationales along with predictions, called self-rationalization models, have been investigated for their interpretability and utility to end-users. However, the extent to which training with human-written ra…

Cited by 15SourcePDFScholar
2022

Few-Shot Self-Rationalization with Natural Language Prompts

NAACL 2022findings

Self-rationalization models that predict task labels and generate free-text elaborations for their predictions could enable more intuitive interaction with NLP systems. These models are, however, currently trained with a large amount of human-written free-text explanations for each task which hinder…

2022

On Advances in Text Generation from Images Beyond Captioning: A Case Study in Self-Rationalization

EMNLP 2022finding

Combining the visual modality with pretrained language models has been surprisingly effective for simple descriptive tasks such as image captioning. More general text generation however remains elusive. We take a step back and ask: How do these models work for more complex generative tasks, i.e. con…

Cited by 3SourcePDFScholar
2021

Teach Me to Explain: A Review of Datasets for Explainable Natural Language Processing

NeurIPS 2021poster

Explainable Natural Language Processing (ExNLP) has increasingly focused on collecting human-annotated textual explanations. These explanations are used downstream in three ways: as data augmentation to improve performance on a predictive task, as supervision to train models to produce explanations…

Cited by 137SourceScholar