← Search

Sofian Chaybouti

4 accepted papers

2026

MaskInversion: Localized Embeddings via Optimization of Explainability Maps

ICLR 2026poster

Vision-language foundation models such as CLIP have achieved tremendous results in global vision-language alignment, but still show some limitations in creating representations for specific image regions. To address this problem, we propose MaskInversion, a method that leverages the feature represe…

Cited by 0SourcecodeScholar
2026

SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models

CVPR 2026

Vision foundation models trained via multi-teacher distillation offer a promising path toward unified visual representations, yet the learning dynamics and data efficiency of such approaches remain underexplored. In this paper, we systematically study multi-teacher distillation for vision foundation

Cited by 0SourcecodeScholar
2026

VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs

CVPR 2026

Vision-Language Models (VLMs) have achieved remarkable progress across tasks such as visual question answering and image captioning. Yet, the extent to which these models perform visual reasoning as opposed to relying on linguistic priors remains unclear. To address this, we introduce VisRes Bench,

Cited by 0SourcecodeScholar
2025

LeGrad: An Explainability Method for Vision Transformers via Feature Formation Sensitivity

ICCV 2025poster

Vision Transformers (ViTs) have become a standard architecture in computer vision. However, because of their modeling of long-range dependencies through self-attention mechanisms, the explainability of these models remains a challenge. To address this, we propose LeGrad, an explainability method spe…