2024
Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality Fusion
ICML 2024poster
While VideoQA Transformer models demonstrate competitive performance on standard benchmarks, the reasons behind their success are not fully understood. Do these models capture the rich multimodal structures and dynamics from video and text jointly? Or are they achieving high scores by exploiting bia…