2021
Learning from Inside: Self-driven Siamese Sampling and Reasoning for Video Question Answering
NeurIPS 2021poster
Recent advances in the video question answering (i.e., VideoQA) task have achieved strong success by following the paradigm of fine-tuning each clip-text pair independently on the pretrained transformer-based model via supervised learning. Intuitively, multiple samples (i.e., clips) should be interd…