2026
GUI-CEval: A Hierarchical and Comprehensive Chinese Benchmark for Mobile GUI Agents
CVPR 2026
Recent progress in Multimodal Large Language Models (MLLMs) has enabled mobile GUI agents capable of visual perception, cross-modal reasoning, and interactive control. However, existing benchmarks are largely English-centric and fail to capture the linguistic and interaction characteristics of the C