2024
Embracing Language Inclusivity and Diversity in CLIP through Continual Language Learning
AAAI 2024technical
While vision-language pre-trained models (VL-PTMs) have advanced multimodal research in recent years, their mastery in a few languages like English restricts their applicability in broader communities. To this end, there is an increasing interest in developing multilingual VL models via a joint-lear…