2024
LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
ECCV 2024poster
"In this work, we present a novel method to tackle the token generation challenge in Vision Language Models (VLMs) for video and image understanding, called LLaMA-VID. Current VLMs, while proficient in tasks like image captioning and visual question answering, face computational burdens when process…