2024
MCAD: Multi-teacher Cross-modal Alignment Distillation for efficient image-text retrieval
NAACL 2024findings
Due to the success of large-scale visual-language pretraining (VLP) models and the widespread use of image-text retrieval in industry areas, it is now critically necessary to reduce the model size and streamline their mobile-device deployment. Single- and dual-stream model structures are commonly us…