2022
SwinBERT: End-to-End Transformers With Sparse Attention for Video Captioning
CVPR 2022poster
The canonical approach to video captioning dictates a caption generation model to learn from offline-extracted dense video features. These feature extractors usually operate on video frames sampled at a fixed frame rate and are often trained on image/video understanding tasks, without adaption to vi…