Shared collection

Vision-Language Models

6 papers · shared by clip_notes

Vision-language and multimodal models that connect images and text.

Follow · See new papers the owner adds
Save a copy · Copy into your own editable folder
2022

BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

ICML 2022spotlight

Vision-Language Pre-training (VLP) has advanced the performance for many vision-language tasks. However, most existing pre-trained models only excel in either understanding-based tasks or generation-based tasks. Furthermore, performance improvement has been largely achieved by scaling up the dataset…

2021

Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

ICML 2021oral

Pre-trained representations are becoming crucial for many NLP and perception tasks. While representation learning in NLP has transitioned to training on raw text without human annotations, visual and vision-language representations still rely heavily on curated training datasets that are expensive o…

Cited by 4467SourcePDFScholar
2024

MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

ICLR 2024poster

The recent GPT-4 has demonstrated extraordinary multi-modal abilities, such as directly generating websites from handwritten text and identifying humorous elements within images. These features are rarely observed in previous vision-language models. However, the technical details behind GPT-4 contin…

2021

VinVL: Revisiting Visual Representations in Vision-Language Models

CVPR 2021poster

This paper presents a detailed study of improving vision features and develops an improved object detection model for vision language (VL) tasks. Compared to the most widely used bottom-up and top-down model [2], the new model is bigger, pre-trained on much larger training corpora that combine multi…

Cited by 1156PDFcodeScholar
2023

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

CoRL 2023poster

We study how vision-language models trained on Internet-scale data can be incorporated directly into end-to-end robotic control to boost generalization and enable emergent semantic reasoning. Our goal is to enable a single end-to-end trained model to both learn to map robot observations to actions a…

Cited by 1068SourceScholar
Vision-Language Models · AIConfPaper