2020
Spatio-Temporal Graph for Video Captioning With Knowledge Distillation
CVPR 2020poster
Video captioning is a challenging task that requires a deep understanding of visual scenes. State-of-the-art methods generate captions using either scene-level or object-level information but without explicitly modeling object interactions. Thus, they often fail to make visually grounded predictions…