← Search

Sitong Gong

2 accepted papers

2026

Reinforcing Video Object Segmentation to Think before it Segments

CVPR 2026

Video reasoning segmentation (VRS) endeavors to delineate referred objects in videos guided by implicit instructions that encapsulate human intent and temporal logic. Previous approaches leverage large vision language models (LVLMs) to encode object semantics into \SEG tokens for mask prediction. Ho

Cited by 0SourceScholar
2025

The Devil is in Temporal Token: High Quality Video Reasoning Segmentation

CVPR 2025poster

Existing methods for Video Reasoning Segmentation rely heavily on a single special token to represent the object in the keyframe or the entire video, inadequately capturing spatial complexity and inter-frame motion. To overcome these challenges, we propose VRS-HQ, an end-to-end video reasoning segme…