← Search

Andrea Esuli

4 accepted papers

2025

Optimizing LLMs for Italian: Reducing Token Fertility and Enhancing Efficiency Through Vocabulary Adaptation

NAACL 2025findings

The number of pretrained Large Language Models (LLMs) is increasing steadily, though the majority are designed predominantly for the English language. While state-of-the-art LLMs can handle other languages, due to language contamination or some degree of multilingual pretraining data, they are not o…

2025

Stress-testing Machine Generated Text Detection: Shifting Language Models Writing Style to Fool Detectors

ACL 2025finding

Recent advancements in Generative AI and Large Language Models (LLMs) have enabled the creation of highly realistic synthetic content, raising concerns about the potential for malicious use, such as misinformation and manipulation. Moreover, detecting Machine-Generated Text (MGT) remains challenging…

2025

The Invalsi Benchmarks: measuring the Linguistic and Mathematical understanding of Large Language Models in Italian

COLING 2025main

While Italian is a high-resource language, there are few Italian-native benchmarks to evaluate generative Large Language Models (LLMs) in this language. This work presents three new benchmarks: Invalsi MATE to evaluate models performance on mathematical understanding in Italian, Invalsi ITA to evalu…

Cited by 0SourcePDFScholar
2024

AI ‘News’ Content Farms Are Easy to Make and Hard to Detect: A Case Study in Italian

ACL 2024long

Large Language Models (LLMs) are increasingly used as ‘content farm’ models (CFMs), to generate synthetic text that could pass for real news articles. This is already happening even for languages that do not have high-quality monolingual LLMs. We show that fine-tuning Llama (v1), mostly trained on E…