Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge
CLIP (Contrastive Language-Image Pre-training) has attracted widespread attention for its multimodal generalizable knowledge, which is significant for downstream tasks. However, the computational overhead of a large number of parameters and large-scale pre-training poses challenges of pre-training a