← Search

Shammur Absar Chowdhury

11 accepted papers

2025

AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs

COLING 2025main

Arabic, with its rich diversity of dialects, remains significantly underrepresented in Large Language Models, particularly in dialectal variations. We address this gap by introducing seven synthetic datasets in dialects alongside Modern Standard Arabic (MSA), created using Machine Translation (MT) c…

Cited by 7SourcePDFScholar
2025

BnTTS: Few-Shot Speaker Adaptation in Low-Resource Setting

NAACL 2025findings

This paper introduces BnTTS (Bangla Text-To-Speech), the first framework for Bangla speaker adaptation-based TTS, designed to bridge the gap in Bangla speech synthesis using minimal training data. Building upon the XTTS architecture, our approach integrates Bangla into a multilingual TTS pipeline, w…

Cited by 0SourcePDFScholar
2025

NativQA: Multilingual Culturally-Aligned Natural Query for LLMs

ACL 2025finding

Natural Question Answering (QA) datasets play a crucial role in evaluating the capabilities of large language models (LLMs), ensuring their effectiveness in real-world applications. Despite the numerous QA datasets that have been developed and some work done in parallel, there is a notable lack of a…

Cited by 0SourcePDFScholar
2024

Speech Collage: Code-Switched Audio Generation by Collaging Monolingual Corpora

ICASSP 2024accepted

Designing effective automatic speech recognition (ASR) systems for Code-Switching (CS) often depends on the availability of the transcribed CS resources. To address data scarcity, this paper introduces Speech Collage, a method that synthesizes CS data from monolingual corpora by splicing audio segme…

Cited by 0SourceScholar
2021

QASR: QCRI Aljazeera Speech Resource A Large Scale Annotated Arabic Speech Corpus

ACL 2021long

We introduce the largest transcribed Arabic speech corpus, QASR, collected from the broadcast domain. This multi-dialect speech dataset contains 2,000 hours of speech sampled at 16kHz crawled from Aljazeera news channel. The dataset is released with lightly supervised transcriptions, aligned with th…

2017

A Deep Learning approach to modeling competitiveness in spoken conversations

ICASSP 2017accepted

The motivation behind the research on overlapping speech has always been dominated by the need to model human-machine interaction for dialog systems and conversation analysis. To have more complex insights of the interlocutors' intentions behind the interaction, we need to understand the type of ove…

Cited by 0SourceScholar
2016

Discourse connective detection in spoken conversations

ICASSP 2016accepted

Discourse parsing is an important task in Language Understanding with applications to human-human and human-machine communication modeling. However, most of the research has focused on written text, and parsers heavily rely on syntactic parsers that themselves have low performance on dialog data. In…

Cited by 0SourceScholar
2015

Annotating and categorizing competition in overlap speech

ICASSP 2015accepted

Overlapping speech is a common and relevant phenomenon in human conversations, reflecting many aspects of discourse dynamics. In this paper, we focus on the pragmatic role of overlaps in turn-in-progress, where it can be categorized as competitive or non-competitive. Previous studies on these two ca…

Cited by 0SourceScholar