2024
Instruction Tuning-free Visual Token Complement for Multimodal LLMs
ECCV 2024poster
"As the open community of large language models (LLMs) matures, multimodal LLMs (MLLMs) have promised an elegant bridge between vision and language. However, current research is inherently constrained by challenges such as the need for high-quality instruction pairs and the loss of visual informatio…