← Search

Jingfan Chen

3 accepted papers

2025

Activation Steering Decoding: Mitigating Hallucination in Large Vision-Language Models through Bidirectional Hidden State Intervention

ACL 2025long

Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities in multimodal understanding, but they frequently suffer from hallucination - generating content inconsistent with visual inputs. In this work, we explore a novel perspective on hallucination mitigation by examining the in…

Cited by 0SourcePDFScholar
2025

AutoGUI: Scaling GUI Grounding with Automatic Functionality Annotations from LLMs

ACL 2025long

User interface understanding with vision-language models (VLMs) has received much attention due to its potential for enhancing software automation.However, existing datasets used to build UI-VLMs either only contain large-scale context-free element annotations or contextualized functional descriptio…

2025

UIPro: Unleashing Superior Interaction Capability For GUI Agents

ICCV 2025poster

Building autonomous agents that perceive and operate graphical user interfaces (GUIs) like humans has long been a vision in the field of artificial intelligence. Central to these agents is the capability for GUI interaction, which involves GUI understanding and planning capabilities. Existing method…