2024
ViLA: Efficient Video-Language Alignment for Video Question Answering
ECCV 2024poster
"We propose an efficient Video-Language Alignment (ViLA) network. Our ViLA model addresses both efficient frame sampling and effective cross-modal alignment in a unified way. In our ViLA network, we design a new learnable text-guided Frame-Prompter together with a cross-modal distillation (QFormer-D…