← Search

Fei Tang

6 accepted papers

2026

GUI-G²: Gaussian Reward Modeling for GUI Grounding

AAAI 2026technical

Graphical User Interface (GUI) grounding maps natural language instructions to precise interface locations for autonomous interaction. Current reinforcement learning approaches use binary rewards that treat elements as hit-or-miss targets, creating sparse signals that ignore the continuous nature of

Cited by 0SourcePDFScholar
2026

GUI-SAGE: Enhancing GUI Automation with Self-Explanatory Learning

CVPR 2026

Reinforcement learning with verifiable rewards (RLVR) has shown promise for GUI automation, enabling agents to learn from binary task completion signals. However, when task difficulty exceeds model capacity, on-policy exploration fails to discover correct actions, creating zero-advantage traps that

Cited by 0SourceScholar
2026

Test-Time Reinforcement Learning for GUI Grounding via Region Consistency

AAAI 2026technical

Graphical User Interface (GUI) grounding, the task of mapping natural language instructions to precise screen coordinates, is fundamental to autonomous GUI agents. While existing methods achieve strong performance through extensive supervised training or reinforcement learning with labeled rewards,

Cited by 0SourcePDFScholar
2025

Social Robot Haru Assisting Dynamic Group Discussion with Autonomous Eye Gaze Behavior

IROS 2025

Due to recent advances in large language models and robotics, social robots will potentially play an important role in people’s daily lives soon, and are expected to improve dynamic multi-party group discussions in social scenarios. In this paper, we developed a system to assist dynamic group discus

Cited by 0SourceScholar
2024

Assisting Group Discussions Using Desktop Robot Haru

ICRA 2024poster

Socially assistive robots are potentially to be integrated with human daily lives in the near future, and expected to be able to improve group dynamics when interacting with groups of people in social settings. In this paper, we developed a system with desktop robot Haru to assist group discussions.…

Cited by 1SourceScholar
2022

CLOOB: Modern Hopfield Networks with InfoLOOB Outperform CLIP

NeurIPS 2022accept

CLIP yielded impressive results on zero-shot transfer learning tasks and is considered as a foundation model like BERT or GPT3. CLIP vision models that have a rich representation are pre-trained using the InfoNCE objective and natural language supervision before they are fine-tuned on particular tas…