2025
MLLM-Bench: Evaluating Multimodal LLMs with Per-sample Criteria
NAACL 2025long
Multimodal large language models (MLLMs) have broadened the scope of AI applications. Existing automatic evaluation methodologies for MLLMs are mainly limited in evaluating objective queries without considering real-world user experiences, inadequately addressing the nuances of creative and associat…