← Search

Hanan Aldarmaki

9 accepted papers

2026

RELUNET: RELATIVE CHANNEL FUSION U-NET FOR MULTICHANNEL SPEECH ENHANCEMENT

ICASSP 2026poster

Neural multi-channel speech enhancement models, in particular those based on the U-Net architecture, demonstrate promising performance and generalization potential. These models typically encode input channels independently, and integrate the channels during later stages of the network. In this pape…

Cited by 0SourcePDFScholar
2025

Dialectal Coverage And Generalization in Arabic Speech Recognition

ACL 2025long

Developing robust automatic speech recognition (ASR) systems for Arabic requires effective strategies to manage its diversity. Existing ASR systems mainly cover the modern standard Arabic (MSA) variety and few high-resource dialects, but fall short in coverage and generalization across the multitude…

2025

Infant Cry Detection Using Causal Temporal Representation

ICASSP 2025accepted

This paper addresses a major challenge in acoustic event detection, in particular infant cry detection in the presence of other sounds and background noises: the lack of precise annotated data. We present two contributions for supervised and unsupervised infant cry detection. The first is an annotat…

Cited by 0SourceScholar
2025

SPIRIT: Patching Speech Language Models against Jailbreak Attacks

EMNLP 2025

Speech Language Models (SLMs) enable natural interactions via spoken instructions, which more effectively capture user intent by detecting nuances in speech. The richer speech signal introduces new security risks compared to text-based models, as adversaries can better bypass safety mechanisms by in

Cited by 0SourcePDFScholar
2025

Voice of a Continent: Mapping Africa’s Speech Technology Frontier

EMNLP 2025

Africa’s rich linguistic diversity remains significantly underrepresented in speech technologies, creating barriers to digital inclusion. To alleviate this challenge, we systematically map the continent’s speech space of datasets and technologies, leading to a new comprehensive benchmark SimbaBench

2024

PALM: Few-Shot Prompt Learning for Audio Language Models

EMNLP 2024main

Audio-Language Models (ALMs) have recently achieved remarkable success in zero-shot audio recognition tasks, which match features of audio waveforms with class-specific text prompt features, inspired by advancements in Vision-Language Models (VLMs). Given the sensitivity of zero-shot performance to…

2024

PolyWER: A Holistic Evaluation Framework for Code-Switched Speech Recognition

EMNLP 2024finding

Code-switching in speech, particularly between languages that use different scripts, can potentially be correctly transcribed in various forms, including different ways of transliteration of the embedded language into the matrix language script. Traditional methods for measuring accuracy, such as Wo…