NAACL 2024long19 citations

BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual Transfer

Akari Asai, Sneha Kudugunta, Xinyan Yu, Terra Blevins, Hila Gonen, Machel Reid, Yulia Tsvetkov, Sebastian Ruder

Abstract

Despite remarkable advancements in few-shot generalization in natural language processing, most models are developed and evaluated primarily in English. To establish a rigorous and equitable evaluation framework for few-shot cross-lingual transfer, we introduce a new benchmark, called BUFFET, which unifies 15 diverse tasks across 54 languages in a sequence-to-sequence format and provides a fixed set of few-shot examples and instructions. Using BUFFET, we perform thorough evaluations of ten state-of-the-art multilingual large language models with different transfer methods, namely in-context learning and fine-tuning. Our findings reveal significant room for improvement in few-shot in-context cross-lingual transfer. Strong multilingual pre-trained or instruction-tuned models such as BLOOM or ChatGPT often lag behind much smaller mT5-base models given the same number of few-shot samples, particularly in low-resource languages. Our analysis suggests avenues for future research in few-shot cross-lingual transfer.

BibTeX
@inproceedings{asai-etal-2024-buffet,
    title = "{BUFFET}: Benchmarking Large Language Models for Few-shot Cross-lingual Transfer",
    author = "Asai, Akari  and
      Kudugunta, Sneha  and
      Yu, Xinyan  and
      Blevins, Terra  and
      Gonen, Hila  and
      Reid, Machel  and
      Tsvetkov, Yulia  and
      Ruder, Sebastian  and
      Hajishirzi, Hannaneh",
    editor = "Duh, Kevin  and
      Gomez, Helena  and
      Bethard, Steven",
    booktitle = "Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)",
    month = jun,
    year = "2024",
    address = "Mexico City, Mexico",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.naacl-long.100/",
    doi = "10.18653/v1/2024.naacl-long.100",
    pages = "1771--1800"
}
BUFFET: Benchmarking Large Language Models for Few-shot Cross-lingual Transfer · NAACL 2024