EMNLP 2023short findings0 citations

IndiSocialFT: Multilingual Word Representation for Indian languages in code-mixed environment

Saurabh Kumar, Ranbir Singh Sanasam, Sukumar Nandi

Abstract

The increasing number of Indian language users on the internet necessitates the development of Indian language technologies. In response to this demand, our paper presents a generalized representation vector for diverse text characteristics, including native scripts, transliterated text, multilingual, code-mixed, and social media-related attributes. We gather text from both social media and well-formed sources and utilize the FastText model to create the "IndiSocialFT" embedding. Through intrinsic and extrinsic evaluation methods, we compare IndiSocialFT with three popular pretrained embeddings trained over Indian languages. Our findings show that the proposed embedding surpasses the baselines in most cases and languages, demonstrating its suitability for various NLP applications.

Indian LanguagesMultilingual Word EmbeddingCode-mixedSocial Media Text
BibTeX
@inproceedings{
kumar2023indisocialft,
title={IndiSocial{FT}: Multilingual Word Representation for Indian languages in code-mixed environment},
author={Saurabh Kumar and Ranbir Singh Sanasam and Sukumar Nandi},
booktitle={The 2023 Conference on Empirical Methods in Natural Language Processing},
year={2023},
url={https://openreview.net/forum?id=nMjktU5AiP}
}
IndiSocialFT: Multilingual Word Representation for Indian languages in code-mixed environment · EMNLP 2023