Rapid Robot Manipulation Policy Learning Via Hierarchical Foundation-Model Prior Distillation
Qingwei Dong, Jiyuan Zhang, Guangxi Wan, Ruikai Liu, Peng Zeng
Abstract
In robotic skill acquisition, rapid policy learning remains challenging due to high-dimensional state-action spaces and inefficient exploration in the early stage of training cite{p1}. Although the pre-trained OpenVLA model exhibits cross-task generalization and can generate goal-directed actions for unseen tasks under suitable prompts, its direct application to novel manipulation tasks remains limited, while full fine-tuning is computationally expensive. To address this issue, we propose a hierarchical framework that combines OpenVLA with reinforcement learning for efficient skill acquisition. Specifically, OpenVLA is used to generate diverse task-related prior trajectories through prompt engineering, and reinforcement learning leverages these priors to fit local dynamics and constrain policy exploration. In this way, the proposed method improves adaptation efficiency and accelerates policy learning on new tasks. We evaluate the framework on multiple manipulation tasks in the LIBERO environment.