Self-Critical Distillation Network for Video-based Commonsense Captioning
Video-based commonsense captioning aims to generate captions for the video content while providing multiple commonsense about the underlying events. Existing approaches rely on constructing a "video -> content caption -> commonsense" reasoning chain, which generates visually ungrounded commonsense a