← Search

Youssef Benchekroun

1 accepted papers

2025

LongTail-Swap: benchmarking language models’ abilities on rare words

EMNLP 2025

Children learn to speak with a low amount of data and can be taught new words on a few-shot basis, making them particularly data-efficient learners. The BabyLM challenge aims at exploring language model (LM) training in the low-data regime but uses metrics that concentrate on the head of the word di