2024
The Belebele Benchmark: a Parallel Reading Comprehension Dataset in 122 Language Variants
ACL 2024long
We present Belebele, a multiple-choice machine reading comprehension (MRC) dataset spanning 122 language variants. Significantly expanding the language coverage of natural language understanding (NLU) benchmarks, this dataset enables the evaluation of text models in high-, medium-, and low-resource…