← Search

Dejie Yang

4 accepted papers

2025

AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical Reasoning

ICCV 2025poster

Visual Robot Manipulation (VRM) aims to enable a robot to follow natural language instructions based on robot states and visual observations, and therefore requires costly multi- modal data. To compensate for the deficiency of robot data, existing approaches have employed vision-language pre- traini…

2025

PlanLLM: Video Procedure Planning with Refinable Large Language Models

AAAI 2025technical

Video procedure planning, i.e., planning a sequence of action steps given the video frames of start and goal states, is an essential ability for embodied AI. Recent works utilize Large Language Models (LLMs) to generate enriched action step description texts to guide action step decoding. Although L…

2024

3D Vision and Language Pretraining with Large-Scale Synthetic Data

IJCAI 2024poster

3D Vision-Language Pre-training (3D-VLP) aims to provide a pre-train model which can bridge 3D scenes with natural language, which is an important technique for embodied intelligence. However, current 3D-VLP datasets are hindered by limited scene-level diversity and insufficient fine-grained annot…

2024

Active Object Detection with Knowledge Aggregation and Distillation from Large Models

CVPR 2024poster

Accurately detecting active objects undergoing state changes is essential for comprehending human interactions and facilitating decision-making. The existing methods for active object detection (AOD) primarily rely on visual appearance of the objects within input such as changes in size shape and re…