← Search

Abhisek Chakrabarty

3 accepted papers

2024

NGLUEni: Benchmarking and Adapting Pretrained Language Models for Nguni Languages

COLING 2024main

The Nguni languages have over 20 million home language speakers in South Africa. There has been considerable growth in the datasets for Nguni languages, but so far no analysis of the performance of NLP models for these languages has been reported across languages and tasks. In this paper we study pr…

Cited by 1SourcePDFScholar
2022

FeatureBART: Feature Based Sequence-to-Sequence Pre-Training for Low-Resource NMT

COLING 2022main

In this paper we present FeatureBART, a linguistically motivated sequence-to-sequence monolingual pre-training strategy in which syntactic features such as lemma, part-of-speech and dependency labels are incorporated into the span prediction based pre-training framework (BART). These automatically e…

Cited by 5SourcePDFScholar
2020

Improving Low-Resource NMT through Relevance Based Linguistic Features Incorporation

COLING 2020main

In this study, linguistic knowledge at different levels are incorporated into the neural machine translation (NMT) framework to improve translation quality for language pairs with extremely limited data. Integrating manually designed or automatically extracted features into the NMT framework is know…