2026
MMBench-Live: A Continuously Evolving Benchmark for Multimodal Models
ICML 2026poster
Evaluation benchmarks play a central role in assessing vision–language models (VLMs). However, most existing multimodal benchmarks are static, making them increasingly vulnerable to data contamination, temporal staleness, and high construction costs. In this work, we introduce MMBench-Live, a multi-…