2026
Asynchronous Temporal Modeling with Two-Agent Framework for Streaming Dense Video Captioning
CVPR 2026
Streaming dense video captioning requires real-time processing of continuous visual input while determining precisely when and what to caption. Current approaches primarily focus on designing complex external memory mechanisms, failing to leverage Large Multimodal Models' (LMMs) inherent long-contex