2025
CovMatch: Cross-Covariance Guided Multimodal Dataset Distillation with Trainable Text Encoder
NeurIPS 2025poster
Multimodal dataset distillation aims to synthesize a small set of image-text pairs that enables efficient training of large-scale vision-language models. While dataset distillation has shown promise in unimodal tasks, extending it to multimodal contrastive learning presents key challenges: learning…