BUS: Efficient and Effective Vision-Language Pre-Training with Bottom-Up Patch Summarization.
Vision Transformer (ViT) based Vision-Language Pretraining (VLP) models recently demonstrated impressive performance in various tasks. However, the lengthy visual token sequences used in these models can lead to inefficient and ineffective performance. Existing methods to address these issues lack t…