← Search

Abdul Waheed

7 accepted papers

2026

VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding

ICLR 2026poster

Precisely evaluating video understanding models remains challenging: commonly used metrics such as BLEU, ROUGE, and BERTScore fail to capture the fineness of human judgment, while obtaining such judgments through manual evaluation is costly. Recent work has explored using large language models (LLMs…

Cited by 0SourceScholar
2025

Lost in Transcription, Found in Distribution Shift: Demystifying Hallucination in Speech Foundation Models

ACL 2025finding

Speech foundation models trained at a massive scale, both in terms of model and data size, result in robust systems capable of performing multiple speech tasks, including automatic speech recognition (ASR). These models transcend language and domain barriers, yet effectively measuring their performa…

2025

uDistil-Whisper: Label-Free Data Filtering for Knowledge Distillation in Low-Data Regimes

NAACL 2025long

Recent work on distilling Whisper’s knowledge into small models using pseudo-labels shows promising performance while reducing the size by up to 50%. This results in small, efficient, and dedicated models. However, a critical step of distillation using pseudo-labels involves filtering high-quality p…

2024

To Distill or Not to Distill? On the Robustness of Robust Knowledge Distillation

ACL 2024long

Arabic is known to present unique challengesfor Automatic Speech Recognition (ASR). Onone hand, its rich linguistic diversity andwide range of dialects complicate the de-velopment of robust, inclusive models. Onthe other, current multilingual ASR modelsare compute-intensive and lack proper com-prehe…

2023

GPTAraEval: A Comprehensive Evaluation of ChatGPT on Arabic NLP

EMNLP 2023long main

ChatGPT's emergence heralds a transformative phase in NLP, particularly demonstrated through its excellent performance on many English benchmarks. However, the model's efficacy across diverse linguistic contexts remains largely uncharted territory. This work aims to bridge this knowledge gap, with a…

Cited by 0SourceScholar