ICML 2026poster0 citations

MentisOculi: Revealing the Limits of Reasoning with Mental Imagery

Jana Zeller, Thaddäus Wiedemer, Fanfei Li, Thomas Klein, Prasanna Mayilvahanan, Matthias Bethge, Felix Wichmann, Ryan Cotterell

Abstract

Frontier models are transitioning from _multimodal large language models_ (MLLMs) that merely ingest visual information to _unified multimodal models_ (UMMs) capable of native interleaved generation. This shift has sparked interest in using intermediate visualizations as a reasoning aid, akin to human _mental imagery_. Central to this idea is the ability to form, maintain, and manipulate visual representations in a goal-oriented manner. To evaluate and probe this capability, we develop MentisOculi, a procedural, stratified suite of multi-step reasoning problems amenable to visual solution, tuned to challenge frontier models. Evaluating visual strategies ranging from latent tokens to explicit generated imagery, we find they generally fail to improve performance. Analysis of UMMs specifically exposes a critical limitation: While they possess the textual reasoning capacity to solve a task and can sometimes generate correct visuals, they suffer from compounding generation errors and fail to leverage even ground-truth visualizations. Our findings suggest that despite their inherent appeal, _visual thoughts do not yet benefit model reasoning_. MentisOculi establishes the necessary foundation to analyze and close this gap across diverse model families.

LLMVisionMultimodalRetrieval
BibTeX
@inproceedings{
zeller2026mentisoculi,
title={MentisOculi: Revealing the Limits of Reasoning with Mental Imagery},
author={Jana Ricarda Zeller and Thadd{\"a}us Wiedemer and Fanfei Li and Thomas Klein and Prasanna Mayilvahanan and Matthias Bethge and Felix A. Wichmann and Ryan Cotterell and Wieland Brendel},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=sxvuK2x3eA}
}