← Search

Eduardo Sánchez

4 accepted papers

2025

LCFO: Long Context and Long Form Output Dataset and Benchmarking

ACL 2025finding

This paper presents the Long Context and Form Output (LCFO) benchmark, a novel evaluation framework for assessing gradual summarization and summary expansion capabilities across diverse domains. LCFO consists of long input documents (5k words average length), each of which comes with three summaries…

2025

Linguini: A benchmark for language-agnostic linguistic reasoning

NeurIPS 2025poster

We propose a new benchmark to measure a language model's linguistic reasoning skills without relying on pre-existing language-specific knowledge. The test covers 894 questions grouped in 160 problems across 75 (mostly) extremely low-resource languages, extracted from the International Linguistic Oly…

Cited by 0SourcecodeScholar
2025

On the Role of Speech Data in Reducing Toxicity Detection Bias

NAACL 2025long

Text toxicity detection systems exhibit significant biases, producing disproportionate rates of false positives on samples mentioning demographic groups. But what about toxicity detection in speech? To investigate the extent to which text-based biases are mitigated by speech-based systems, we produc…

Cited by 0SourcePDFScholar
2024

Machine Translation Hallucination Detection for Low and High Resource Languages using Large Language Models

EMNLP 2024finding

Recent advancements in massively multilingual machine translation systems have significantly enhanced translation accuracy; however, even the best performing systems still generate hallucinations, severely impacting user trust. Detecting hallucinations in Machine Translation (MT) remains a critical…