2026
Composition-Grounded Instruction Synthesis for Visual Reasoning
ICLR 2026poster
Pretrained multi-modal large language models (MLLMs) demonstrate strong performance on diverse multimodal tasks, but remain limited in reasoning capabilities for domains where annotations are difficult to collect. In this work, we focus on artificial image domains such as charts, rendered documents,…