2025
DTOS: Dynamic Time Object Sensing with Large Multimodal Model
CVPR 2025poster
Existing multimodal large language models (MLLMs) face significant challenges in Referring Video Object Segmentation(RVOS). We identify three critical challenges: (C1) insufficient quantitative representation of textual numerical data, (C2) repetitive and degraded response templates for spatiotempor…