DeepPhy: Benchmarking Agentic VLMs on Physical Reasoning
Xinrun Xu, Pi Bu, Ye Wang, Börje F. Karlsson, Ziming Wang, Tengtao Song, Qi Zhu, Jun Song
Abstract
Although Vision Language Models (VLMs) exhibit strong perceptual abilities and impressive visual reasoning, they struggle with attention to detail and precise action planning in complex, dynamic environments, leading to subpar performance. Real-world tasks typically require complex interactions, advanced spatial reasoning, long-term planning, and continuous strategy refinement, usually necessitating understanding the physics rules of the target scenario. However, evaluating these capabilities in real-world scenarios is often prohibitively expensive. To bridge this gap, we introduce DeepPHY, a novel benchmark framework designed to systematically evaluate VLMs
BibTeX
@inproceedings{aaai2026_deepphybenchmark,
title = {DeepPhy: Benchmarking Agentic VLMs on Physical Reasoning},
author = {Xinrun Xu and Pi Bu and Ye Wang and Börje F. Karlsson and Ziming Wang and Tengtao Song and Qi Zhu and Jun Song and Zhiming Ding and Bo Zheng},
booktitle = {AAAI 2026},
year = {2026}
}