2025
Measuring the Effect of Disfluency in Multilingual Knowledge Probing Benchmarks
EMNLP 2025
For multilingual factual knowledge assessment of LLMs, benchmarks such as MLAMA use template translations that do not take into account the grammatical and semantic information of the named entities inserted in the sentence. This leads to numerous instances of ungrammaticality or wrong wording of th