2026
RPTS: Tree-Structured Reasoning Process Scoring for Faithful Multimodal Evaluation
AAAI 2026technical
Large Vision-Language Models (LVLMs) excel in multimodal reasoning and have shown impressive performance on various multimodal benchmarks. However, most of these benchmarks evaluate models primarily through multiple-choice or short-answer formats, which do not take the reasoning process into account