← Search

Pranaydeep Singh

4 accepted papers

2025

EnerGIZAr: Leveraging GIZA++ for Effective Tokenizer Initialization

ACL 2025finding

Continual pre-training has long been considered the default strategy for adapting models to non-English languages, but struggles with initializing new embeddings, particularly for non-Latin scripts. In this work, we propose EnerGIZAr, a novel methodology that improves continual pre-training by lever…

2024

Lemmatisation of Medieval Greek: Against the Limits of Transformer’s Capabilities?

COLING 2024main

This paper presents preliminary experiments for the lemmatisation of unedited, Byzantine Greek epigrams. This type of Greek is quite different from its classical ancestor, mostly because of its orthographic inconsistencies. Existing lemmatisation algorithms display an accuracy drop of around 30pp wh…

Cited by 0SourcePDFScholar
2023

Misery Loves Complexity: Exploring Linguistic Complexity in the Context of Emotion Detection

EMNLP 2023long findings

Given the omnipresence of social media in our society, thoughts and opinions are being shared online in an unprecedented manner. This means that both positive and negative emotions can be equally and freely expressed. However, the negativity bias posits that human beings are inherently drawn to and…

Cited by 0SourceScholar
2022

When the Student Becomes the Master: Learning Better and Smaller Monolingual Models from mBERT

COLING 2022main

In this research, we present pilot experiments to distil monolingual models from a jointly trained model for 102 languages (mBERT). We demonstrate that it is possible for the target language to outperform the original model, even with a basic distillation setup. We evaluate our methodology for 6 lan…