2025
VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
IROS 2025
Large Vision Language Models (VLMs) have been adopted in robotics for their strong common sense understanding and generalization capabilities. Existing works leverage VLMs for task and motion planning based on language instructions and robot observations. In this work, we explore using VLM to interp