← Search

Felice Dell’Orletta

15 accepted papers

2025

Beyond the Spelling Miracle: Investigating Substring Awareness in Character-Blind Language Models

ACL 2025finding

Correctly identifying characters and substrings of words should be a basic but essential ability of any Language Model that aims to proficiently understand and produce language. Despite so, the majority of Pre-trained Language Models (PLMs) are “character-blind” and struggle in spelling tasks, altho…

2025

Evaluating Lexical Proficiency in Neural Language Models

ACL 2025long

We present a novel evaluation framework designed to assess the lexical proficiency and linguistic creativity of Transformer-based Language Models (LMs). We validate the framework by analyzing the performance of a set of LMs of different sizes, in both mono- and multilingual configuration, across tas…

2025

From Human Reading to NLM Understanding: Evaluating the Role of Eye-Tracking Data in Encoder-Based Models

ACL 2025long

Cognitive signals, particularly eye-tracking data, offer valuable insights into human language processing. Leveraging eye-gaze data from the Ghent Eye-Tracking Corpus, we conducted a series of experiments to examine how integrating knowledge of human reading behavior impacts Neural Language Models (…

Cited by 0SourcePDFScholar
2025

Optimizing LLMs for Italian: Reducing Token Fertility and Enhancing Efficiency Through Vocabulary Adaptation

NAACL 2025findings

The number of pretrained Large Language Models (LLMs) is increasing steadily, though the majority are designed predominantly for the English language. While state-of-the-art LLMs can handle other languages, due to language contamination or some degree of multilingual pretraining data, they are not o…

2025

Stress-testing Machine Generated Text Detection: Shifting Language Models Writing Style to Fool Detectors

ACL 2025finding

Recent advancements in Generative AI and Large Language Models (LLMs) have enabled the creation of highly realistic synthetic content, raising concerns about the potential for malicious use, such as misinformation and manipulation. Moreover, detecting Machine-Generated Text (MGT) remains challenging…

2025

TEXT-CAKE: Challenging Language Models on Local Text Coherence

COLING 2025main

We present a deep investigation of encoder-based Language Models (LMs) on their abilities to detect text coherence across four languages and four text genres using a new evaluation benchmark, TEXT-CAKE. We analyze both multilingual and monolingual LMs with varying architectures and parameters in dif…

Cited by 2SourcePDFScholar
2024

AI ‘News’ Content Farms Are Easy to Make and Hard to Detect: A Case Study in Italian

ACL 2024long

Large Language Models (LLMs) are increasingly used as ‘content farm’ models (CFMs), to generate synthetic text that could pass for real news articles. This is already happening even for languages that do not have high-quality monolingual LLMs. We show that fine-tuning Llama (v1), mostly trained on E…

2024

Evaluating Large Language Models via Linguistic Profiling

EMNLP 2024main

Large Language Models (LLMs) undergo extensive evaluation against various benchmarks collected in established leaderboards to assess their performance across multiple tasks. However, to the best of our knowledge, there is a lack of comprehensive studies evaluating these models’ linguistic abilities…

2024

Fine-tuning with HED-IT: The impact of human post-editing for dialogical language models

ACL 2024findings

Automatic methods for generating and gathering linguistic data have proven effective for fine-tuning Language Models (LMs) in languages less resourced than English. Still, while there has been emphasis on data quantity, less attention has been given to its quality. In this work, we investigate the i…

2024

Linguistic Knowledge Can Enhance Encoder-Decoder Models (If You Let It)

COLING 2024main

In this paper, we explore the impact of augmenting pre-trained Encoder-Decoder models, specifically T5, with linguistic knowledge for the prediction of a target task. In particular, we investigate whether fine-tuning a T5 model on an intermediate task that predicts structural linguistic properties o…

2023

Coherent or Not? Stressing a Neural Language Model for Discourse Coherence in Multiple Languages

ACL 2023findings

In this study, we investigate the capability of a Neural Language Model (NLM) to distinguish between coherent and incoherent text, where the latter has been artificially created to gradually undermine local coherence within text. While previous research on coherence assessment using NLMs has primari…

2022

How about Time? Probing a Multilingual Language Model for Temporal Relations

COLING 2022main

This paper presents a comprehensive set of probing experiments using a multilingual language model, XLM-R, for temporal relation classification between events in four languages. Results show an advantage of contextualized embeddings over static ones and a detrimen- tal role of sentence level embeddi…

2022

On the Nature of BERT: Correlating Fine-Tuning and Linguistic Competence

COLING 2022main

Several studies in the literature on the interpretation of Neural Language Models (NLM) focus on the linguistic generalization abilities of pre-trained models. However, little attention is paid to how the linguistic knowledge of the models changes during the fine-tuning steps. In this paper, we cont…

Cited by 3SourcePDFScholar
2022

Outlier Dimensions that Disrupt Transformers are Driven by Frequency

EMNLP 2022finding

While Transformer-based language models are generally very robust to pruning, there is the recently discovered outlier phenomenon: disabling only 48 out of 110M parameters in BERT-base drops its performance by nearly 30% on MNLI. We replicate the original evidence for the outlier phenomenon and we l…

2020

Linguistic Profiling of a Neural Language Model

COLING 2020main

In this paper we investigate the linguistic knowledge learned by a Neural Language Model (NLM) before and after a fine-tuning process and how this knowledge affects its predictions during several classification problems. We use a wide set of probing tasks, each of which corresponds to a distinct sen…

Cited by 70SourcePDFScholar