← Search

Yankui Sun

1 accepted papers

2023

ViLTA: Enhancing Vision-Language Pre-training through Textual Augmentation

ICCV 2023poster

Vision-language pre-training (VLP) methods are blossoming recently, and its crucial goal is to jointly learn visual and textual features via a transformer-based architecture, demonstrating promising improvements on a variety of vision-language tasks. Prior arts usually focus on how to align visual a…

Cited by 13PDFScholar