D-GARA: A Dynamic Benchmarking Framework for GUI Agent Robustness in Real-World Anomalies
Developing intelligent agents capable of operating a wide range of Graphical User Interfaces (GUIs) with human-level proficiency is a key milestone on the path toward Artificial General Intelligence. While most existing datasets and benchmarks for training and evaluating GUI agents are static and id