← Search

Xiantao Zhang

2 accepted papers

2025

Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical Reasoning

AAAI 2025technical

Multimodal Large Language Models (MLLMs) excel in solving text-based mathematical problems, but they struggle with mathematical diagrams since they are primarily trained on natural scene images. For humans, visual aids generally enhance problem-solving, but MLLMs perform worse as information shifts…