← Search

Els Lefever

8 accepted papers

2025

EnerGIZAr: Leveraging GIZA++ for Effective Tokenizer Initialization

ACL 2025finding

Continual pre-training has long been considered the default strategy for adapting models to non-English languages, but struggles with initializing new embeddings, particularly for non-Latin scripts. In this work, we propose EnerGIZAr, a novel methodology that improves continual pre-training by lever…

2025

Evaluating Transformers for OCR Post-Correction in Early Modern Dutch Theatre

COLING 2025main

This paper explores the effectiveness of two types of transformer models — large generative models and sequence-to-sequence models — for automatically post-correcting Optical Character Recognition (OCR) output in early modern Dutch plays. To address the need for optimally aligned data, we create a p…

Cited by 1SourcePDFScholar
2025

Lemmatisation & Morphological Analysis of Unedited Greek: Do Simple Tasks Need Complex Solutions?

ACL 2025finding

Fine-tuning transformer-based models for part-of-speech tagging of unedited Greek text has outperformed traditional systems. However, when applied to lemmatisation or morphological analysis, fine-tuning has not yet achieved competitive results. This paper explores various approaches to combine morph…

Cited by 0SourcePDFScholar
2024

At the Crossroad of Cuneiform and NLP: Challenges for Fine-grained Part-of-speech Tagging

COLING 2024main

The study of ancient Middle Eastern cultures is dominated by the vast number of cuneiform texts. Multiple languages and language families were expressed in cuneiform. The most dominant language written in cuneiform is the Semitic Akkadian, which is the focus of this paper. We are specifically focusi…

2024

Human and System Perspectives on the Expression of Irony: An Analysis of Likelihood Labels and Rationales

COLING 2024main

In this paper, we examine the recognition of irony by both humans and automatic systems. We achieve this by enhancing the annotations of an English benchmark data set for irony detection. This enhancement involves a layer of human-annotated irony likelihood using a 7-point Likert scale that combines…

2024

Lemmatisation of Medieval Greek: Against the Limits of Transformer’s Capabilities?

COLING 2024main

This paper presents preliminary experiments for the lemmatisation of unedited, Byzantine Greek epigrams. This type of Greek is quite different from its classical ancestor, mostly because of its orthographic inconsistencies. Existing lemmatisation algorithms display an accuracy drop of around 30pp wh…

Cited by 0SourcePDFScholar
2023

Misery Loves Complexity: Exploring Linguistic Complexity in the Context of Emotion Detection

EMNLP 2023long findings

Given the omnipresence of social media in our society, thoughts and opinions are being shared online in an unprecedented manner. This means that both positive and negative emotions can be equally and freely expressed. However, the negativity bias posits that human beings are inherently drawn to and…

Cited by 0SourceScholar
2022

When the Student Becomes the Master: Learning Better and Smaller Monolingual Models from mBERT

COLING 2022main

In this research, we present pilot experiments to distil monolingual models from a jointly trained model for 102 languages (mBERT). We demonstrate that it is possible for the target language to outperform the original model, even with a basic distillation setup. We evaluate our methodology for 6 lan…