2025
ViCo: A Multitask Video-enhanced and Cognition-preserving Modality Alignment Training Framework
ICASSP 2025accepted
The rapid development of multimodal large language models (MLLMs) has brought significant breakthroughs to this field. However, current MLLMs typically rely on vision instruction tuning based on large language models (LLMs) to endow them with multimodal capabilities, which may lead to low video util…