2026
Vision-aligned Latent Reasoning for Multi-Modal Large Language Model
ICML 2026poster
Despite recent advancements in Multi-modal Large Language Models (MLLMs) on diverse understanding tasks, these models struggle to solve problems which require extensive multi-step reasoning. This is primarily due to the progressive dilution of visual information during long-context generation, which…