← Search

Alessio Tonioni*

1 accepted papers

2024

BRAVE: Broadening the visual encoding of vision-language models

ECCV 2024oral

"Vision-language models (VLMs) are typically composed of a vision encoder, e.g. CLIP, and a language model (LM) that interprets the encoded features to solve downstream tasks. Despite remarkable progress, VLMs are subject to several shortcomings due to the limited capabilities of vision encoders, e.…