2026
GUI-Shift: Enhancing VLM-Based GUI Agents through Self-supervised Reinforcement Learning
ICLR 2026poster
Training effective Vision-Language Models (VLMs) for GUI agents typically depends on large-scale annotated datasets, whose collection is both labor-intensive and error-prone. We introduce K-step GUI Transition, a self-supervised inverse dynamics task in which VLMs learn GUI dynamics by predicting th…