2024
LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
ECCV 2024poster
"This paper presents (), a general-purpose multimodal assistant trained using an end-to-end approach that systematically expands the capabilities of large multimodal models (LMMs). maintains a skill repository that contains a wide range of vision and vision-language pre-trained models (tools), and i…