Asynchronous Temporal Modeling with Two-Agent Framework for Streaming Dense Video Captioning
Streaming dense video captioning requires real-time processing of continuous visual input while determining precisely when and what to caption. Current approaches primarily focus on designing complex external memory mechanisms, failing to leverage Large Multimodal Models' (LMMs) inherent long-contex