← Search

Yi Jiang*

1 accepted papers

2024

Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models

ECCV 2024poster

"We introduce Groma, a Multimodal Large Language Model (MLLM) with grounded and fine-grained visual perception ability. Beyond holistic image understanding, Groma is adept at region-level tasks such as region captioning and visual grounding. Such capabilities are built upon a localized visual tokeni…