← Search

Walid Bousselham

6 accepted papers

2026

MaskInversion: Localized Embeddings via Optimization of Explainability Maps

ICLR 2026poster

Vision-language foundation models such as CLIP have achieved tremendous results in global vision-language alignment, but still show some limitations in creating representations for specific image regions. To address this problem, we propose MaskInversion, a method that leverages the feature represe…

Cited by 0SourcecodeScholar
2026

VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation

CVPR 2026

Training vision-language models (VLMs) for complex reasoning remains a challenging task, i.a. due to the scarcity of high-quality image-text reasoning data. Conversely, text-based reasoning resources are abundant and scalable, but it is still an open question how to leveraging them for VLM reasoning

Cited by 0SourcecodeScholar
2025

LeGrad: An Explainability Method for Vision Transformers via Feature Formation Sensitivity

ICCV 2025poster

Vision Transformers (ViTs) have become a standard architecture in computer vision. However, because of their modeling of long-range dependencies through self-attention mechanisms, the explainability of these models remains a challenge. To address this, we propose LeGrad, an explainability method spe…

2025

VideoGEM: Training-free Action Grounding in Videos

CVPR 2025poster

Vision-language foundation models have shown impressive capabilities across various zero-shot tasks, including training-free localization and grounding, primarily focusing on localizing objects in images. However, leveraging those capabilities to localize actions and events in videos is challenging,…

2024

Grounding Everything: Emerging Localization Properties in Vision-Language Transformers

CVPR 2024poster

Vision-language foundation models have shown remarkable performance in various zero-shot settings such as image retrieval classification or captioning. But so far those models seem to fall behind when it comes to zero-shot localization of referential expressions and objects in images. As a result th…

2023

Learning Situation Hyper-Graphs for Video Question Answering

CVPR 2023poster

Answering questions about complex situations in videos requires not only capturing of the presence of actors, objects, and their relations, but also the evolution of these relationships over time. A situation hyper-graph is a representation that describes situations as scene sub-graphs for video fra…