← Search

Zakaria Aldeneh

14 accepted papers

2026

Closing the Gap Between Text and Speech Understanding in LLMs

ICLR 2026poster

Large Language Models (LLMs) can be adapted to extend their text capabilities to speech inputs. However, these speech-adapted LLMs consistently underperform their text-based counterparts—and even cascaded pipelines—on language understanding tasks. We term this shortfall the text–speech understanding…

Cited by 0SourcecodeScholar
2026

LEVERAGING AUDIO-VISUAL DATA TO REDUCE THE MULTILINGUAL GAP IN SELF-SUPERVISED SPEECH MODELS

ICASSP 2026poster

Self-supervised learning (SSL) has made significant advances in speech representation learning. Models like wav2vec 2.0 and HuBERT have achieved state-of-the-art results in tasks such as speech recognition, particularly in monolingual settings. However, multilingual SSL models tend to underperform t…

Cited by 0SourcePDFScholar
2025

A Variational Framework for Improving Naturalness in Generative Spoken Language Models

ICML 2025poster

The success of large language models in text processing has inspired their adaptation to speech modeling. However, since speech is continuous and complex, it is often discretized for autoregressive modeling. Speech tokens derived from self-supervised models (known as semantic tokens) typically focus…

2025

Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models

ICASSP 2025accepted

Speech foundation models, such as HuBERT and its variants, are pre-trained on large amounts of unlabeled speech data and then used for a range of downstream tasks. These models use a masked prediction objective, where the model learns to predict information about masked input segments from the unmas…

Cited by 0SourceScholar
2025

Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels

ICASSP 2025accepted

Iterative self-training, or iterative pseudo-labeling (IPL)—using an improved model from the current iteration to provide pseudo-labels for the next iteration—has proven to be a powerful approach to enhance the quality of speaker representations. Recent applications of IPL in unsupervised speaker re…

Cited by 0SourceScholar
2025

Towards Automatic Assessment of Self-Supervised Speech Models using Rank

ICASSP 2025accepted

This study explores using embedding rank as an unsupervised evaluation metric for general-purpose speech encoders trained via self-supervised learning (SSL). Traditionally, assessing the performance of these encoders is resource-intensive and requires labeled data from the downstream tasks. Inspired…

Cited by 0SourceScholar
2024

Learning Spatially-Aware Language and Audio Embeddings

NeurIPS 2024poster

Humans can picture a sound scene given an imprecise natural language description. For example, it is easy to imagine an acoustic environment given a phrase like "the lion roar came from right behind me!". For a machine to have the same degree of comprehension, the machine must know what a lion is (…

Cited by 1SourcePDFScholar
2023

Naturalistic Head Motion Generation from Speech

ICASSP 2023accepted

Synthesizing natural head motion to accompany speech for an embodied conversational agent is necessary for pro-viding a rich interactive experience. Most prior works assess the quality of generated head motion by comparing them against a single ground-truth using an objective metric. Yet there are m…

Cited by 0SourceScholar
2023

On the Role of LIP Articulation in Visual Speech Perception

ICASSP 2023accepted

Generating realistic lip motion from audio to simulate speech production is critical for driving natural character animation. Previous research has shown that traditional metrics used to optimize and assess models for generating lip motion from speech are not a good indicator of subjective opinion o…

Cited by 0SourceScholar
2021

Learning Paralinguistic Features from Audiobooks through Style Voice Conversion

NAACL 2021long

Paralinguistics, the non-lexical components of speech, play a crucial role in human-human interaction. Models designed to recognize paralinguistic information, particularly speech emotion and style, are difficult to train because of the limited labeled datasets available. In this work, we present a…

Cited by 2SourcePDFScholar
2021

On The Role of Visual Cues in Audiovisual Speech Enhancement

ICASSP 2021accepted

We present an introspection of an audiovisual speech enhancement model. In particular, we focus on interpreting how a neural audiovisual speech enhancement model uses visual cues to improve the quality of the target speech signal. We show that visual cues provide not only high-level information abou…

Cited by 0SourceScholar
2019

Muse-ing on the Impact of Utterance Ordering on Crowdsourced Emotion Annotations

ICASSP 2019accepted

Emotion recognition algorithms rely on data annotated with high quality labels. However, emotion expression and perception are inherently subjective. There is generally not a single annotation that can be unambiguously declared "correct." As a result, annotations are colored by the manner in which t…

Cited by 0SourceScholar
2018

Improving End-of-Turn Detection in Spoken Dialogues by Detecting Speaker Intentions as a Secondary Task

ICASSP 2018accepted

This work focuses on the use of acoustic cues for modeling turn-taking in dyadic spoken dialogues. Previous work has shown that speaker intentions (e.g., asking a question, uttering a backchannel, etc.) can influence turn-taking behavior and are good predictors of turn-transitions in spoken dialogue…

Cited by 0SourceScholar