2025
The Invalsi Benchmarks: measuring the Linguistic and Mathematical understanding of Large Language Models in Italian
COLING 2025main
While Italian is a high-resource language, there are few Italian-native benchmarks to evaluate generative Large Language Models (LLMs) in this language. This work presents three new benchmarks: Invalsi MATE to evaluate models performance on mathematical understanding in Italian, Invalsi ITA to evalu…