2024
Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models
CVPR 2024highlight
Large Multimodal Models (LMMs) have shown promise in vision-language tasks but struggle with high-resolution input and detailed scene understanding. Addressing these challenges we introduce Monkey to enhance LMM capabilities. Firstly Monkey processes input images by dividing them into uniform patche…