← Search

Ildoo Kim

5 accepted papers

2021

ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision

ICML 2021oral

Vision-and-Language Pre-training (VLP) has improved performance on various joint vision-and-language downstream tasks. Current approaches to VLP heavily rely on image feature extraction processes, most of which involve region supervision (e.g., object detection) and the convolutional architecture (e…