2024
LoSh: Long-Short Text Joint Prediction Network for Referring Video Object Segmentation
CVPR 2024poster
Referring video object segmentation (RVOS) aims to segment the target instance referred by a given text expression in a video clip. The text expression normally contains sophisticated description of the instance's appearance action and relation with others. It is therefore rather difficult for a RVO…