2026
VIRST: Video-Instructed Reasoning Assistant for SpatioTemporal Segmentation
CVPR 2026
Referring Video Object Segmentation (RVOS) aims to segment target objects in videos based on natural language descriptions. However, fixed keyframe-based approaches that couple a vision language model with a separate propagation module often fail to capture rapidly changing spatiotemporal dynamics a