ACL 2022findings8 citations

SyMCoM - Syntactic Measure of Code Mixing A Study Of English-Hindi Code-Mixing

Prashant Kodali, Anmol Goel, Monojit Choudhury, Manish Shrivastava, Ponnurangam Kumaraguru

Abstract

Code mixing is the linguistic phenomenon where bilingual speakers tend to switch between two or more languages in conversations. Recent work on code-mixing in computational settings has leveraged social media code mixed texts to train NLP models. For capturing the variety of code mixing in, and across corpus, Language ID (LID) tags based measures (CMI) have been proposed. Syntactical variety/patterns of code-mixing and their relationship vis-a-vis computational model’s performance is under explored. In this work, we investigate a collection of English(en)-Hindi(hi) code-mixed datasets from a syntactic lens to propose, SyMCoM, an indicator of syntactic variety in code-mixed text, with intuitive theoretical bounds. We train SoTA en-hi PoS tagger, accuracy of 93.4%, to reliably compute PoS tags on a corpus, and demonstrate the utility of SyMCoM by applying it on various syntactical categories on a collection of datasets, and compare datasets using the measure.

BibTeX
@inproceedings{kodali-etal-2022-symcom,
    title = "{S}y{MC}o{M} - Syntactic Measure of Code Mixing A Study Of {E}nglish-{H}indi Code-Mixing",
    author = "Kodali, Prashant  and
      Goel, Anmol  and
      Choudhury, Monojit  and
      Shrivastava, Manish  and
      Kumaraguru, Ponnurangam",
    editor = "Muresan, Smaranda  and
      Nakov, Preslav  and
      Villavicencio, Aline",
    booktitle = "Findings of the Association for Computational Linguistics: ACL 2022",
    month = may,
    year = "2022",
    address = "Dublin, Ireland",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2022.findings-acl.40/",
    doi = "10.18653/v1/2022.findings-acl.40",
    pages = "472--480"
}
SyMCoM - Syntactic Measure of Code Mixing A Study Of English-Hindi Code-Mixing · ACL 2022