2026
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
CVPR 2026
Despite recent advances in multimodal reasoning, Multimodal Large Language Models (MLLMs) still struggle on complex tasks where initial visual perceptions can be misleading. This performance gap stems from a critical reasoning flaw we term Visual Inertia: while MLLMs excel at iterative reflection in