← Search

Teng Yihua

2 accepted papers

2026

MMSearch-Plus: Benchmarking Provenance-Aware Search for Multimodal Browsing Agents

ICLR 2026poster

Existing multimodal browsing benchmarks often fail to require genuine multimodal reasoning, as many tasks can be solved with text-only heuristics without vision-in-the-loop verification. We introduce MMSearch-Plus, a 311-task benchmark that enforces multimodal understanding by requiring extraction a…

Cited by 0SourcecodeScholar
2024

Android in the Zoo: Chain-of-Action-Thought for GUI Agents

EMNLP 2024finding

Large language model (LLM) leads to a surge of autonomous GUI agents for smartphone, which completes a task triggered by natural language through predicting a sequence of actions of API. Even though the task highly relies on past actions and visual observations, existing studies typically consider l…