Dynamic Multimodal Evaluation via Knowledge-Enhanced Benchmark Evolution
The rapid development of multimodal large language models (MLLMs) has created an urgent demand for more reliable and robust evaluation protocols. However, existing static benchmarks are prone to data contamination and performance saturation, which can result in inflated or misleading evaluation resu…