← Search

Haohao Luo

6 accepted papers

2026

IntentRL: Training Proactive User-intent Agents for Open-ended Deep Research via Reinforcement Learning

ICML 2026poster

Deep Research (DR) agents extend Large Language Models (LLMs) beyond parametric knowledge by autonomously retrieving and synthesizing evidence from large web corpora into long-form reports, enabling a long-horizon agentic paradigm. However, unlike real-time conversational assistants, DR is computati…

Cited by 0SourceScholar
2025

Browsing Like Human: A Multimodal Web Agent with Experiential Fast-and-Slow Thinking

ACL 2025long

Automating web navigation which aims to build a web agent that follows user instructions to complete tasks like booking flights by interacting with websites, has received increasing attention due to its practical value. Although existing web agents are mostly equipped with visual perception, plannin…

2025

Express What You See: Can Multimodal LLMs Decode Visual Ciphers with Intuitive Semiosis Comprehension?

ACL 2025finding

Bridging the gap between visual and language remains a pivotal challenge for the multimodal community. Traditional VQA benchmarks encounter a modality gap and over-reliance on language priors, whereas human cognition excels at intuitive semiosis, associating abstract visual symbols to linguistic sem…

Cited by 0SourcePDFScholar
2025

INREACT: An Inspire-Then-Reinforce Training Framework For Multimodal GUI Agent

EMNLP 2025

Graphical User Interface (GUI) interaction, which aims to develop an intelligent GUI agent that executes user instructions to perform tasks such as installing applications by controlling digital devices, has gained significant attention due to its practical value. Although current advanced multimoda

2024

Chain-of-Exemplar: Enhancing Distractor Generation for Multimodal Educational Question Generation

ACL 2024long

Multiple-choice questions (MCQs) are important in enhancing concept learning and student engagement for educational purposes. Despite the multimodal nature of educational content, current methods focus mainly on text-based inputs and often neglect the integration of visual information. In this work,…

Cited by 8SourcePDFScholar