EM-KD: Distilling Efficient Multimodal Large Language Model with Unbalanced Vision Tokens
Efficient Multimodal Large Language Models (MLLMs) compress vision tokens to reduce resource consumption, but the loss of visual information can degrade comprehension capabilities. Although some priors introduce Knowledge Distillation to enhance student models, they overlook the fundamental differen