Focus Like a Human: Efficient GUI Grounding via Coarse-to-Fine Visual Attention and Parallel Verification
Building upon powerful Large Visual Language Models, recent GUI agents have revolutionized autonomous GUI interaction. Given the high information density and structural complexity of GUI layouts, a critical challenge lies in accurately identifying where to focus, i.e., precise GUI grounding. To ensu