← Search

Júlia Falcão

5 accepted papers

2025

Cognitive Biases, Task Complexity, and Result Intepretability in Large Language Models

COLING 2025main

In humans, cognitive biases are systematic deviations from rationality in judgment that simplify complex decisions. They typically manifest as a consequence of learned behaviors or limitations on information processing capabilities. Recent work has shown that these biases can percolate through train…

Cited by 2SourcePDFScholar
2025

IberoBench: A Benchmark for LLM Evaluation in Iberian Languages

COLING 2025main

The current best practice to measure the performance of base Large Language Models is to establish a multi-task benchmark that covers a range of capabilities of interest. Currently, however, such benchmarks are only available in a few high-resource languages. To address this situation, we present Ib…

Cited by 2SourcePDFScholar
2025

La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin America

ACL 2025long

Leaderboards showcase the current capabilities and limitations of Large Language Models (LLMs). To motivate the development of LLMs that represent the linguistic and cultural diversity of the Spanish-speaking community, we present La Leaderboard, the first open-source leaderboard to evaluate generat…

Cited by 0SourcePDFScholar
2025

VeritasQA: A Truthfulness Benchmark Aimed at Multilingual Transferability

COLING 2025main

As Large Language Models (LLMs) become available in a wider range of domains and applications, evaluating the truthfulness of multilingual LLMs is an issue of increasing relevance. TruthfulQA (Lin et al., 2022) is one of few benchmarks designed to evaluate how models imitate widespread falsehoods. H…

2024

COMET for Low-Resource Machine Translation Evaluation: A Case Study of English-Maltese and Spanish-Basque

COLING 2024main

Trainable metrics for machine translation evaluation have been scoring the highest correlations with human judgements in the latest meta-evaluations, outperforming traditional lexical overlap metrics such as BLEU, which is still widely used despite its well-known shortcomings. In this work we look a…