← Search

Sandro Pezzelle

12 accepted papers

2025

From Tools to Teammates: Evaluating LLMs in Multi-Session Coding Interactions

ACL 2025long

Large Language Models (LLMs) are increasingly used in working environments for a wide range of tasks, excelling at solving individual problems in isolation. However, are they also able to effectively collaborate over long-term interactions? To investigate this, we introduce MemoryCode, a synthetic m…

2025

LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

ACL 2025short

There is an increasing trend towards evaluating NLP models with LLMs instead of human judgments, raising questions about the validity of these evaluations, as well as their reproducibility in the case of proprietary models. We provide JUDGE-BENCH, an extensible collection of 20 NLP datasets with hum…

2025

They want to pretend not to understand: The Limits of Current LLMs in Interpreting Implicit Content of Political Discourse

ACL 2025finding

Implicit content plays a crucial role in political discourse, where systematically employ pragmatic strategies such as implicatures and presuppositions to influence their audiences. Large Language Models (LLMs) have demonstrated strong performance in tasks requiring complex semantic and pragmatic un…

2024

Do Pre-Trained Language Models Detect and Understand Semantic Underspecification? Ask the DUST!

ACL 2024findings

In everyday language use, speakers frequently utter and interpret sentences that are semantically underspecified, namely, whose content is insufficient to fully convey their message or interpret them univocally. For example, to interpret the underspecified sentence “Don’t spend too much”, which leav…

2024

Naming, Describing, and Quantifying Visual Objects in Humans and LLMs

ACL 2024short

While human speakers use a variety of different expressions when describing the same object in an image, giving rise to a distribution of plausible labels driven by pragmatic constraints, the extent to which current Vision & Language Large Language Models (VLLMs) can mimic this crucial feature of la…

2024

Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition

EMNLP 2024finding

Visual storytelling consists in generating a natural language story given a temporally ordered sequence of images. This task is not only challenging for models, but also very difficult to evaluate with automatic metrics since there is no consensus about what makes a story ‘good’. In this paper, we i…

2023

GROOViST: A Metric for Grounding Objects in Visual Storytelling

EMNLP 2023short main

A proper evaluation of stories generated for a sequence of images---the task commonly referred to as visual storytelling---must consider multiple aspects, such as coherence, grammatical correctness, and visual grounding. In this work, we focus on evaluating the degree of grounding, that is, the exte…

Cited by 0SourcecodeScholar
2023

Speaking the Language of Your Listener: Audience-Aware Adaptation via Plug-and-Play Theory of Mind

ACL 2023findings

Dialogue participants may have varying levels of knowledge about the topic under discussion. In such cases, it is essential for speakers to adapt their utterances by taking their audience into account. Yet, it is an open question how such adaptation can be modelled in computational agents. In this p…

2023

The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models

EMNLP 2023long main

Despite the impressive performance achieved by pre-trained language-and-vision models in downstream tasks, it remains an open question whether this reflects a proper understanding of image-text interaction. In this work, we explore to what extent they handle basic linguistic constructions---active-p…

Cited by 0SourcecodeScholar
2023

When Language Models Fall in Love: Animacy Processing in Transformer Language Models

EMNLP 2023long main

Animacy—whether an entity is alive and sentient—is fundamental to cognitive processing, impacting areas such as memory, vision, and language. However, animacy is not always expressed directly in language: in English it often manifests indirectly, in the form of selectional constraints on verbs and a…

Cited by 0SourcecodeScholar