← Search

Dakuan Lu

3 accepted papers

2026

GUI-Eyes: Tool-Augmented Perception for Visual Grounding in GUI Agents

AAAI 2026technical

Recent advances in vision-language models (VLMs) and reinforcement learning (RL) have driven progress in GUI automation. However, most existing methods rely on static, one-shot visual inputs and passive perception, lacking the ability to adaptively determine when, whether, and how to observe the int

Cited by 0SourcePDFScholar
2025

Guess What I am Thinking: A Benchmark for Inner Thought Reasoning of Role-Playing Language Agents

EMNLP 2025

Recent advances in Large Language Model (LLM)-based Role-Playing Language Agents (RPLAs) have attracted broad attention in various applications. While chain-of-thought reasoning has shown importance in many tasks for LLMs, the internal thinking processes of RPLAs remain unexplored. Understanding cha

Cited by 0SourcePDFScholar
2025

ORIGAMISPACE: Benchmarking Multimodal LLMs in Multi-Step Spatial Reasoning with Mathematical Constraints

NeurIPS 2025spotlight

Spatial reasoning is a key capability in the field of artificial intelligence, especially crucial in areas such as robotics, computer vision, and natural language understanding. However, evaluating the ability of multimodal large language models (MLLMs) in complex spatial reasoning still faces chall…

Cited by 0SourceScholar