← Search

Yasser Abdelaziz Dahou Djilali

4 accepted papers

2025

Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment

CVPR 2025poster

Recent contrastive multimodal vision-language models like CLIP have demonstrated robust open-world semantic understanding, becoming the standard image backbones for vision-language applications. However, recent findings suggest high semantic similarity between well-trained unimodal encoders, which r…

2024

Do Vision and Language Encoders Represent the World Similarly?

CVPR 2024poster

Aligned text-image encoders such as CLIP have become the de-facto model for vision-language tasks. Furthermore modality-specific encoders achieve impressive performances in their respective domains. This raises a central question: does an alignment exist between uni-modal vision and language encoder…

2023

Lip2Vec: Efficient and Robust Visual Speech Recognition via Latent-to-Latent Visual to Audio Representation Mapping

ICCV 2023poster

Visual Speech Recognition (VSR) differs from the common perception tasks as it requires deeper reasoning over the video sequence, even by human experts. Despite the recent advances in VSR, current approaches rely on labeled data to fully train or finetune their models predicting the target speech. T…

Cited by 9PDFcodeScholar
2021

Rethinking 360deg Image Visual Attention Modelling With Unsupervised Learning.

ICCV 2021poster

Despite the success of self-supervised representation learning on planar data, to date it has not been studied on 360deg images. In this paper, we extend recent advances in contrastive learning to learn latent representations that are sufficiently invariant to be highly effective for spherical salie…

Cited by 16PDFcodeScholar