← Search

Raiymbek Akshulakov

2 accepted papers

2025

Harnessing Frozen Unimodal Encoders for Flexible Multimodal Alignment

CVPR 2025poster

Recent contrastive multimodal vision-language models like CLIP have demonstrated robust open-world semantic understanding, becoming the standard image backbones for vision-language applications. However, recent findings suggest high semantic similarity between well-trained unimodal encoders, which r…

2024

Do Vision and Language Encoders Represent the World Similarly?

CVPR 2024poster

Aligned text-image encoders such as CLIP have become the de-facto model for vision-language tasks. Furthermore modality-specific encoders achieve impressive performances in their respective domains. This raises a central question: does an alignment exist between uni-modal vision and language encoder…