2024
One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos
NeurIPS 2024poster
We introduce VideoLISA, a video-based multimodal large language model designed to tackle the problem of language-instructed reasoning segmentation in videos. Leveraging the reasoning capabilities and world knowledge of large language models, and augmented by the Segment Anything Model, VideoLISA gen…