← Search

Beichen Wang

1 accepted papers

2025

VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model

IROS 2025

Large Vision Language Models (VLMs) have been adopted in robotics for their strong common sense understanding and generalization capabilities. Existing works leverage VLMs for task and motion planning based on language instructions and robot observations. In this work, we explore using VLM to interp

Cited by 45SourcecodeScholar