EMNLP 2024main1 citations

Understanding Slang with LLMs: Modelling Cross-Cultural Nuances through Paraphrasing

Ifeoluwa Wuraola, Nina Dethlefs, Daniel Marciniak

Abstract

In the realm of social media discourse, the integration of slang enriches communication, reflecting the sociocultural identities of users. This study investigates the capability of large language models (LLMs) to paraphrase slang within climate-related tweets from Nigeria and the UK, with a focus on identifying emotional nuances. Using DistilRoBERTa as the base-line model, we observe its limited comprehension of slang. To improve cross-cultural understanding, we gauge the effectiveness of leading LLMs ChatGPT 4, Gemini, and LLaMA3 in slang paraphrasing. While ChatGPT 4 and Gemini demonstrate comparable effectiveness in slang paraphrasing, LLaMA3 shows less coverage, with all LLMs exhibiting limitations in coverage, especially of Nigerian slang. Our findings underscore the necessity for culturally sensitive LLM development in emotion classification, particularly in non-anglocentric regions.

BibTeX
@inproceedings{wuraola-etal-2024-understanding,
    title = "Understanding Slang with {LLM}s: Modelling Cross-Cultural Nuances through Paraphrasing",
    author = "Wuraola, Ifeoluwa  and
      Dethlefs, Nina  and
      Marciniak, Daniel",
    editor = "Al-Onaizan, Yaser  and
      Bansal, Mohit  and
      Chen, Yun-Nung",
    booktitle = "Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing",
    month = nov,
    year = "2024",
    address = "Miami, Florida, USA",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2024.emnlp-main.869/",
    doi = "10.18653/v1/2024.emnlp-main.869",
    pages = "15525--15531"
}