← Search

MD Sakib Ul Rahman Sourove

2 accepted papers

2025

BanTH: A Multi-label Hate Speech Detection Dataset for Transliterated Bangla

NAACL 2025findings

The proliferation of transliterated texts in digital spaces has emphasized the need for detecting and classifying hate speech in languages beyond English, particularly in low-resource languages. As online discourse can perpetuate discrimination based on target groups, e.g. gender, religion, and orig…

Cited by 0SourcePDFScholar
2024

BanglaTLit: A Benchmark Dataset for Back-Transliteration of Romanized Bangla

EMNLP 2024finding

Low-resource languages like Bangla are severely limited by the lack of datasets. Romanized Bangla texts are ubiquitous on the internet, offering a rich source of data for Bangla NLP tasks and extending the available data sources. However, due to the informal nature of romanized text, they often lack…