2024
Osprey: Pixel Understanding with Visual Instruction Tuning
CVPR 2024poster
Multimodal large language models (MLLMs) have recently achieved impressive general-purpose vision-language capabilities through visual instruction tuning. However current MLLMs primarily focus on image-level or box-level understanding falling short in achieving fine-grained vision-language alignment…