2026
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
ICLR 2026poster
Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities across audio and visual modalities, often neglecting either one of the modaliti…