2025
Assessing Modality Bias in Video Question Answering Benchmarks with Multimodal Large Language Models
AAAI 2025technical
Multimodal large language models (MLLMs) can simultaneously process visual, textual, and auditory data, capturing insights that complement human analysis. However, existing video question-answering (VidQA) benchmarks and datasets often exhibit a bias toward a single modality, despite the goal of re…