← Search

Mahmoud Al Ismail

4 accepted papers

2023

CLAP Learning Audio Concepts from Natural Language Supervision

ICASSP 2023accepted

Mainstream machine listening models are trained to learn audio concepts under the paradigm of one class label to many recordings focusing on one task. Learning under such restricted supervision limits the flexibility of models because they require labeled audio for training and can only predict the…

Cited by 0SourceScholar
2019

Disjoint Mapping Network for Cross-modal Matching of Voices and Faces

ICLR 2019poster

We propose a novel framework, called Disjoint Mapping Network (DIMNet), for cross-modal biometric matching, in particular of voices and faces. Different from the existing methods, DIMNet does not explicitly learn the joint relationship between the modalities. Instead, DIMNet learns a shared represen…

Cited by 92SourcePDFScholar