2024
VGA: Vision GUI Assistant - Minimizing Hallucinations through Image-Centric Fine-Tuning
EMNLP 2024finding
Large Vision-Language Models (VLMs) have already been applied to the understanding of Graphical User Interfaces (GUIs) and have achieved notable results. However, existing VLMs often overly rely on internal text-based knowledge while neglecting visual inputs. This imbalance may lead models to produc…