Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
Yolo Yunlong Tang, Jing Bi, Chao Huang, Susan Liang, Daiki Shimada, Hang Hua, Yunzhong Xiao, Yizhi Song
Abstract
In this work, we introduce CAT-V (Caption Anything in Video), a training-free framework for fine-grained object-centric video captioning of user-selected instances. CAT-V combines (i) a SAMURAI-based Segmenter for precise object masks across frames, (ii) a TRACE-Uni Temporal Analyzer for event boundary detection and coarse event descriptions, and (iii) an InternVL-2.5 Captioner that, conditioned on spatiotemporal visual prompts and chain-of-thought (CoT) guidance, produces detailed, temporally coherent captions about object attributes, actions, states, interactions, and context. The system supports point, box, and region prompts and maintains temporal sensitivity by tracking object states across segments. In contrast to vanilla video captioning that is overly abstract and dense video captioning that is often terse, CAT-V enables object-level specificity with spatial accuracy and temporal coherence, without additional training data.
BibTeX
@inproceedings{aaai2026_captionanythingi,
title = {Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting},
author = {Yolo Yunlong Tang and Jing Bi and Chao Huang and Susan Liang and Daiki Shimada and Hang Hua and Yunzhong Xiao and Yizhi Song and Pinxin Liu and Mingqian Feng and Junjia Guo and Zhuo Liu and Luchuan Song and Ali Vosoughi and Jinxi He and Liu He and Zeliang Zhang and Jiebo Luo and Chenliang Xu},
booktitle = {AAAI 2026},
year = {2026}
}