← Search

Kumar Saunack

2 accepted papers

2021

How low is too low? A monolingual take on lemmatisation in Indian languages

NAACL 2021long

Lemmatization aims to reduce the sparse data problem by relating the inflected forms of a word to its dictionary form. Most prior work on ML based lemmatization has focused on high resource languages, where data sets (word forms) are readily available. For languages which have no linguistic work ava…

2020

Analysing cross-lingual transfer in lemmatisation for Indian languages

COLING 2020main

Lemmatization aims to reduce the sparse data problem by relating the inflected forms of a word to its dictionary form. However, most of the prior work on this topic has focused on high resource languages. In this paper, we evaluate cross-lingual approaches for low resource languages, especially in t…

Cited by 2SourcePDFScholar