2024
VISA: Reasoning Video Object Segmentation via Large Language Model
ECCV 2024poster
"Existing Video Object Segmentation (VOS) relies on explicit user instructions, such as categories, masks, or short phrases, restricting their ability to perform complex video segmentation requiring reasoning with world knowledge. In this paper, we introduce a new task, Reasoning Video Object Segmen…