← Search

Mohammad Mamun Or Rashid

3 accepted papers

2025

BanNERD: A Benchmark Dataset and Context-Driven Approach for Bangla Named Entity Recognition

NAACL 2025findings

In this study, we introduce BanNERD, the most extensive human-annotated and validated Bangla Named Entity Recognition Dataset to date, comprising over 85,000 sentences. BanNERD is curated from a diverse array of sources, spanning over 29 domains, thereby offering a comprehensive range of generalized…

2024

Unicode Normalization and Grapheme Parsing of Indic Languages

COLING 2024main

Writing systems of Indic languages have orthographic syllables, also known as complex graphemes, as unique horizontal units. A prominent feature of these languages is these complex grapheme units that comprise consonants/consonant conjuncts, vowel diacritics, and consonant diacritics, which, togethe…

2023

BanLemma: A Word Formation Dependent Rule and Dictionary Based Bangla Lemmatizer

EMNLP 2023long findings

Lemmatization holds significance in both natural language processing (NLP) and linguistics, as it effectively decreases data density and aids in comprehending contextual meaning. However, due to the highly inflected nature and morphological richness, lemmatization in Bangla text poses a complex chal…

Cited by 7SourcecodeScholar