2022
End-to-End Referring Video Object Segmentation With Multimodal Transformers
CVPR 2022poster
The referring video object segmentation task (RVOS) involves segmentation of a text-referred object instance in the frames of a given video. Due to the complex nature of this multimodal task, which combines text reasoning, video understanding, instance segmentation and tracking, existing approaches…