2025
B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens
ICCV 2025poster
Recently, Vision Large Language Models (VLLMs) with integrated vision encoders have shown promising performance in vision understanding. They encode visual content into sequences of visual tokens, enabling joint processing of visual and textual data. However, understanding videos, especially long vi…