2026
ReasonMap: Towards Fine-Grained Visual Reasoning from Transit Maps
CVPR 2026
Multimodal large language models (MLLMs) have demonstrated significant progress in semantic scene understanding and text-image alignment, with reasoning variants enhancing performance on more complex tasks involving mathematics and logic. However, their proficiency in tasks requiring both fine-grain