← Search

Enshen Zhou

10 accepted papers

2026

Prune4Web: DOM Tree Pruning Programming for Web Agent

AAAI 2026technical

Web automation uses intelligent agents to perform high-level tasks by mimicking human interactions with webpages. Despite recent advances in LLM-based web agents, efficiently navigating complex, real-world webpages remains challenging due to massive DOM structures (10,000 ~ 100,000 tokens). Current

Cited by 0SourcePDFScholar
2026

SaPaVe: Towards Active Perception and Manipulation in Vision-Language Action Models for Robotics

CVPR 2026

Active perception and manipulation are crucial for robots to interact with complex scenes. Existing methods struggle to unify semantic-driven perception actively with robust, viewpoint-invariant execution accordingly. To this end, we propose SaPaVe, an end-to-end framework that jointly learns these

Cited by 0SourceScholar
2026

TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics

ICRA 2026poster

Vision-Language Models (VLMs) have shown remarkable capabilities in spatial reasoning, yet they remain fundamentally limited to qualitative assessments and lack the computational precision required for real-world robotics. Current approaches fail to leverage metric information from depth sensors and…

2025

AGFSync: Leveraging AI-Generated Feedback for Preference Optimization in Text-to-Image Generation

AAAI 2025technical

Text-to-Image (T2I) diffusion models have achieved remarkable success in image generation. Despite their progress, challenges remain in both prompt-following ability, image quality and lack of high-quality datasets, which are essential for refining these models. As acquiring labeled data is costly,…

Cited by 2SourcePDFScholar
2025

Chain-of-Imagination for Reliable Instruction Following in Decision Making

IROS 2025

Enabling the embodied agent to imagine step-by-step the future states and sequentially approach these situation-aware states can enhance its capability to make reliable action decisions from textual instructions. In this work, we introduce a simple but effective mechanism called Chain-of-Imagination

Cited by 0SourceScholar
2025

Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection

CVPR 2025poster

Automatic detection and prevention of open-set failures are crucial in closed-loop robotic systems. Recent studies often struggle to simultaneously identify unexpected failures reactively after they occur and prevent foreseeable ones proactively. To this end, we propose Code-as-Monitor (CaM), a nove…

Cited by 7SourcePDFScholar
2025

RH20T-P: A Primitive-Level Robotic Manipulation Dataset towards Composable Generalization Agents in Real-world Scenarios

IROS 2025

Achieving generalizability in solving out-of-distribution tasks is one of the ultimate goals of learning robotic manipulation. Recent progress of Vision-Language Models (VLMs) has shown that VLM-based task planners can alleviate the difficulty of solving novel tasks, by decomposing the compounded ta

Cited by 1SourceScholar
2025

RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics

NeurIPS 2025poster

Spatial referring is a fundamental capability of embodied robots to interact with the 3D physical world. However, even with the powerful pretrained VLMs, recent approaches are still not qualified to accurately understand the complex 3D scenes and dynamically reason about the instruction-indicated lo…

Cited by 0SourceScholar
2025

WorldSimBench: Towards Video Generation Models as World Simulators

ICML 2025poster

Recent advancements in predictive models have demonstrated exceptional capabilities in predicting the future state of objects and scenes. However, the lack of categorization based on inherent characteristics continues to hinder the progress of predictive model development. Additionally, existing ben…

Cited by 18SourcePDFScholar
2024

MP5: A Multi-modal Open-ended Embodied System in Minecraft via Active Perception

CVPR 2024poster

It is a long-lasting goal to design an embodied system that can solve long-horizon open-world tasks in human-like ways. However existing approaches usually struggle with compound difficulties caused by the logic-aware decomposition and context-aware execution of these tasks. To this end we introduce…

Cited by 32SourcePDFScholar