← Search

Gabriel Fiastre

1 accepted papers

2026

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects

CVPR 2026

Dense Video Object Captioning (DVOC) is the task of jointly detecting, tracking, and captioning object trajectories in a video, requiring the ability to understand spatio-temporal details and describe them in natural language.Due to the complexity of the task and the high cost associated with manual

Cited by 1SourceScholar