← Search

Irene Baucells

5 accepted papers

2025

IberoBench: A Benchmark for LLM Evaluation in Iberian Languages

COLING 2025main

The current best practice to measure the performance of base Large Language Models is to establish a multi-task benchmark that covers a range of capabilities of interest. Currently, however, such benchmarks are only available in a few high-resource languages. To address this situation, we present Ib…

Cited by 2SourcePDFScholar
2025

Multi-LMentry: Can Multilingual LLMs Solve Elementary Tasks Across Languages?

EMNLP 2025

As large language models (LLMs) continue to improve, their evaluation increasingly centers on complex, high-level tasks, often at the expense of systematically assessing fundamental capabilities. To address this gap, recent work proposed LMentry, a compact benchmark comprising tasks that are trivial

Cited by 0SourcePDFScholar
2024

Building a Data Infrastructure for a Mid-Resource Language: The Case of Catalan

COLING 2024main

Current LLM-based applications are becoming steadily available for everyone with a reliable access to technology and the internet. These applications offer benefits to their users that leave those without access to them at a serious disadvantage. Given the vastly large amount of data needed to train…

2024

FLOR: On the Effectiveness of Language Adaptation

COLING 2024main

Large language models have amply proven their great capabilities, both in downstream tasks and real-life settings. However, low- and mid-resource languages do not have access to the necessary means to train such models from scratch, and often have to rely on multilingual models despite being underre…

2023

Dynamic Stance: Modeling Discussions by Labeling the Interactions

EMNLP 2023long findings

Stance detection is an increasingly popular task that has been mainly modeled as a static task, by assigning the expressed attitude of a text toward a given topic. Such a framing presents limitations, with trained systems showing poor generalization capabilities and being strongly topic-dependent. I…

Cited by 0SourceScholar