← Search

Hawau Olamide Toyin

8 accepted papers

2025

All Languages Matter: Evaluating LMMs on Culturally Diverse 100 Languages

CVPR 2025highlight

Existing Large Multimodal Models (LMMs) generally focus on only a few regions and languages. As LMMs continue to improve, it is increasingly important to ensure they understand cultural contexts, respect local sensitivities, and support low-resource languages, all while effectively integrating corr…

2025

Dialectal Coverage And Generalization in Arabic Speech Recognition

ACL 2025long

Developing robust automatic speech recognition (ASR) systems for Arabic requires effective strategies to manage its diversity. Existing ASR systems mainly cover the modern standard Arabic (MSA) variety and few high-resource dialects, but fall short in coverage and generalization across the multitude…

2025

Exploring the Limitations of Detecting Machine-Generated Text

COLING 2025main

Recent improvements in the quality of the generations by large language models have spurred research into identifying machine-generated text. Such work often presents high-performing detectors. However, humans and machines can produce text in different styles and domains, yet the the performance imp…

Cited by 2SourcePDFScholar
2025

Infant Cry Detection Using Causal Temporal Representation

ICASSP 2025accepted

This paper addresses a major challenge in acoustic event detection, in particular infant cry detection in the presence of other sounds and background noises: the lack of precise annotated data. We present two contributions for supervised and unsupervised infant cry detection. The first is an annotat…

Cited by 0SourceScholar
2025

Voice of a Continent: Mapping Africa’s Speech Technology Frontier

EMNLP 2025

Africa’s rich linguistic diversity remains significantly underrepresented in speech technologies, creating barriers to digital inclusion. To alleviate this challenge, we systematically map the continent’s speech space of datasets and technologies, leading to a new comprehensive benchmark SimbaBench

2025

Where Are We? Evaluating LLM Performance on African Languages

ACL 2025long

Africa’s rich linguistic heritage remains underrepresented in NLP, largely due to historical policies that favor foreign languages and create significant data inequities. In this paper, we integrate theoretical insights on Africa’s language landscape with an empirical evaluation using Sahara— a comp…

Cited by 0SourcePDFScholar
2024

PolyWER: A Holistic Evaluation Framework for Code-Switched Speech Recognition

EMNLP 2024finding

Code-switching in speech, particularly between languages that use different scripts, can potentially be correctly transcribed in various forms, including different ways of transliteration of the embedded language into the matrix language script. Traditional methods for measuring accuracy, such as Wo…