2026
MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
ICLR 2026poster
Recent advances in large language models have laid the foundation for multimodal LLMs (MLLMs), which unify text, speech, and vision within a single framework. As these models are rapidly evolving toward general-purpose instruction following across diverse and complex tasks, a key frontier is evaluat…