← Search

Arina Turkatenko

4 accepted papers

2025

2M-BELEBELE: Highly Multilingual Speech and American Sign Language Comprehension Dataset Download PDF

ACL 2025finding

We introduce the first highly multilingual speech and American Sign Language (ASL) comprehension dataset by extending BELEBELE. Our dataset covers 91 spoken languages at the intersection of BELEBELE and FLEURS, and one sign language (ASL). As a by-product we also extend the Automatic Speech Recognit…

2025

BOUQuET : dataset, Benchmark and Open initiative for Universal Quality Evaluation in Translation

EMNLP 2025

BOUQuET is a multi-way, multicentric and multi-register/domain dataset and benchmark, and a broader collaborative initiative. This dataset is handcrafted in 8 non-English languages (i.e. Egyptian Arabic and Modern Standard Arabic, French, German, Hindi, Indonesian, Mandarin Chinese, Russian, and Spa

Cited by 0SourcePDFScholar
2025

LCFO: Long Context and Long Form Output Dataset and Benchmarking

ACL 2025finding

This paper presents the Long Context and Form Output (LCFO) benchmark, a novel evaluation framework for assessing gradual summarization and summary expansion capabilities across diverse domains. LCFO consists of long input documents (5k words average length), each of which comes with three summaries…

2025

Linguini: A benchmark for language-agnostic linguistic reasoning

NeurIPS 2025poster

We propose a new benchmark to measure a language model's linguistic reasoning skills without relying on pre-existing language-specific knowledge. The test covers 894 questions grouped in 160 problems across 75 (mostly) extremely low-resource languages, extracted from the International Linguistic Oly…

Cited by 0SourcecodeScholar