← Search

Raquel Fernández

14 accepted papers

2026

Optimizing Language Models for Crosslingual Knowledge Consistency

ICML 2026poster

Large language models are known to often exhibit inconsistent knowledge. This is particularly problematic in multilingual scenarios, where models are likely to be asked similar questions in different languages, and inconsistent responses can undermine their reliability. In this work, we show that th…

Cited by 0SourceScholar
2025

Cross-Lingual Transfer of Debiasing and Detoxification in Multilingual LLMs: An Extensive Investigation

ACL 2025finding

Recent generative large language models (LLMs) show remarkable performance in non-English languages, but when prompted in those languages they tend to express higher harmful social biases and toxicity levels. Prior work has shown that finetuning on specialized datasets can mitigate this behavior, an…

2025

I see what you mean: Co-Speech Gestures for Reference Resolution in Multimodal Dialogue

ACL 2025finding

In face-to-face interaction, we use multiple modalities, including speech and gestures, to communicate information and resolve references to objects. However, how representational co-speech gestures refer to objects remains understudied from a computational perspective. In this work, we address this…

2025

LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks

ACL 2025short

There is an increasing trend towards evaluating NLP models with LLMs instead of human judgments, raising questions about the validity of these evaluations, as well as their reproducibility in the case of proprietary models. We provide JUDGE-BENCH, an extensible collection of 20 NLP datasets with hum…

2024

Don’t Buy it! Reassessing the Ad Understanding Abilities of Contrastive Multimodal Models

ACL 2024short

Image-based advertisements are complex multimodal stimuli that often contain unusual visual elements and figurative language. Previous research on automatic ad understanding has reported impressive zero-shot accuracy of contrastive vision-and-language models (VLMs) on an ad-explanation retrieval tas…

2024

Model Internals-based Answer Attribution for Trustworthy Retrieval-Augmented Generation

EMNLP 2024main

Ensuring the verifiability of model answers is a fundamental challenge for retrieval-augmented generation (RAG) in the question answering (QA) domain. Recently, self-citation prompting was proposed to make large language models (LLMs) generate citations to supporting documents along with their answe…

2024

Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition

EMNLP 2024finding

Visual storytelling consists in generating a natural language story given a temporally ordered sequence of images. This task is not only challenging for models, but also very difficult to evaluate with automatic metrics since there is no consensus about what makes a story ‘good’. In this paper, we i…

2023

Cross-Lingual Consistency of Factual Knowledge in Multilingual Language Models

EMNLP 2023long main

Multilingual large-scale Pretrained Language Models (PLMs) have been shown to store considerable amounts of factual knowledge, but large variations are observed across languages. With the ultimate goal of ensuring that users with different language backgrounds obtain consistent feedback from the sam…

Cited by 0SourcecodeScholar
2023

GROOViST: A Metric for Grounding Objects in Visual Storytelling

EMNLP 2023short main

A proper evaluation of stories generated for a sequence of images---the task commonly referred to as visual storytelling---must consider multiple aspects, such as coherence, grammatical correctness, and visual grounding. In this work, we focus on evaluating the degree of grounding, that is, the exte…

Cited by 0SourcecodeScholar
2023

Information Value: Measuring Utterance Predictability as Distance from Plausible Alternatives

EMNLP 2023long main

We present information value, a measure which quantifies the predictability of an utterance relative to a set of plausible alternatives. We introduce a method to obtain interpretable estimates of information value using neural text generators, and exploit their psychometric predictive power to inves…

Cited by 0SourcecodeScholar
2023

The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models

EMNLP 2023long main

Despite the impressive performance achieved by pre-trained language-and-vision models in downstream tasks, it remains an open question whether this reflects a proper understanding of image-text interaction. In this work, we explore to what extent they handle basic linguistic constructions---active-p…

Cited by 0SourcecodeScholar
2023

What Comes Next? Evaluating Uncertainty in Neural Text Generators Against Human Production Variability

EMNLP 2023long main

In Natural Language Generation (NLG) tasks, for any input, multiple communicative goals are plausible, and any goal can be put into words, or produced, in multiple ways. We characterise the extent to which human production varies lexically, syntactically, and semantically across four NLG tasks, conn…

Cited by 0SourcecodeScholar
2021

Is Information Density Uniform in Task-Oriented Dialogues?

EMNLP 2021main

The Uniform Information Density principle states that speakers plan their utterances to reduce fluctuations in the density of the information transmitted. In this paper, we test whether, and within which contextual units this principle holds in task-oriented dialogues. We show that there is evidence…

2020

Words are the Window to the Soul: Language-based User Representations for Fake News Detection

COLING 2020main

Cognitive and social traits of individuals are reflected in language use. Moreover, individuals who are prone to spread fake news online often share common traits. Building on these ideas, we introduce a model that creates representations of individuals on social media based only on the language the…