← Search

Sukumar Nandi

3 accepted papers

2024

Evaluating Performance of Pre-trained Word Embeddings on Assamese, a Low-resource Language

COLING 2024main

Word embeddings and Language models are the building blocks of modern Deep Neural Network-based Natural Language Processing. They are extensively explored in high-resource languages and provide state-of-the-art (SOTA) performance for a wide range of downstream tasks. Nevertheless, these word embeddi…

2024

IndiSentiment140: Sentiment Analysis Dataset for Indian Languages with Emphasis on Low-Resource Languages using Machine Translation

NAACL 2024long

Sentiment analysis, a fundamental aspect of Natural Language Processing (NLP), involves the classification of emotions, opinions, and attitudes in text data. In the context of India, with its vast linguistic diversity and low-resource languages, the challenge is to support sentiment analysis in nume…

Cited by 3SourcePDFScholar
2023

IndiSocialFT: Multilingual Word Representation for Indian languages in code-mixed environment

EMNLP 2023short findings

The increasing number of Indian language users on the internet necessitates the development of Indian language technologies. In response to this demand, our paper presents a generalized representation vector for diverse text characteristics, including native scripts, transliterated text, multilingua…

Cited by 0SourceScholar