2026
Evaluating Cross-Modal Reasoning Ability and Problem Charactaristics with Multimodal Item Response Theory
ICLR 2026poster
Multimodal Large Language Models (MLLMs) have recently emerged as general architectures capable of reasoning over diverse modalities. Benchmarks for MLLMs should measure their ability for cross‑modal integration. However, current benchmarks are filled with shortcut questions, which can be solved usi…