← Search

Yassine Ouali

10 accepted papers

2026

VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions

CVPR 2026

Existing approaches for improving the efficiency of Large Vision-Language Models (LVLMs) are largely based on the concept of visual token reduction. This approach, however, creates an information bottleneck that impairs performance, especially on challenging tasks that require fine-grained understan

Cited by 0SourceScholar
2025

Compress & Cache: Vision token compression for efficient generation and retrieval

NeurIPS 2025poster

This work aims to compress the vision tokens of an LVLM into a representation that is simultaneously suitable for (a) generative and (b) discriminative tasks, (c) is nearly lossless, and (d) storage-efficient. To this end, we propose C&C, a novel compression method that leverages the LVLM itself fo…

Cited by 0SourceScholar
2025

Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions

EMNLP 2025

Contrastively-trained Vision-Language Models (VLMs), such as CLIP, have become the standard approach for learning discriminative vision-language representations. However, these models often exhibit shallow language understanding, manifesting bag-of-words behaviour. These limitations are reinforced b

Cited by 0SourcePDFScholar
2025

VladVA: Discriminative Fine-tuning of LVLMs

CVPR 2025poster

Contrastively-trained Vision-Language Models (VLMs) like CLIP have become the de facto approach for discriminative vision-language representation learning. However, these models have limited language understanding, often exhibiting a "bag of words" behavior. At the same time, Large Vision-Language M…

Cited by 0SourcePDFScholar
2024

Efficient Vision-Language pre-training via domain-specific learning for human activities

EMNLP 2024main

Current Vision-Language (VL) models owe their success to large-scale pre-training on web-collected data, which in turn requires high-capacity architectures and large compute resources for training. We posit that when the downstream tasks are known in advance, which is in practice common, the pretrai…

2024

FFF: Fixing Flawed Foundations in Contrastive Pre-Training Results in Very Strong Vision-Language Models

CVPR 2024poster

Despite noise and caption quality having been acknowledged as important factors impacting vision-language contrastive pre-training in this paper we show that the full potential of improving the training process by addressing such issues is yet to be realized. Specifically we firstly study and analyz…

Cited by 5SourcePDFScholar
2023

Black Box Few-Shot Adaptation for Vision-Language Models

ICCV 2023poster

Vision-Language (V-L) models trained with contrastive learning to align the visual and language modalities have been shown to be strong few-shot learners. Soft prompt learning is the method of choice for few-shot downstream adaption aiming to bridge the modality gap caused by the distribution shift…

Cited by 47PDFcodeScholar
2020

Semi-Supervised Semantic Segmentation With Cross-Consistency Training

CVPR 2020poster

In this paper, we present a novel cross-consistency based semi-supervised approach for semantic segmentation. Consistency training has proven to be a powerful semi-supervised learning framework for leveraging unlabeled data under the cluster assumption, in which the decision boundary should lie in l…

Cited by 1026PDFcodeScholar