2024
LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models
ECCV 2024poster
"With the recent significant advancements in large multimodal models (LMMs), the importance of their grounding capability in visual chat is increasingly recognized. Despite recent efforts to enable LMMs to support grounding, their capabilities for grounding and chat are usually separate, and their c…