2025
MASSIVE-Agents: A Benchmark for Multilingual Function-Calling in 52 Languages
EMNLP 2025
We present MASSIVE-Agents, a new benchmark for assessing multilingual function calling across 52 languages. We created MASSIVE-Agents by cleaning the original MASSIVE dataset and then reformatting it for evaluation within the Berkeley Function-Calling Leaderboard (BFCL) framework. The full benchmark