DB-MPO: Demonstration Boosted Reactive Grasping For Two-Finger Gripper
Boya Zhang, Andreas Zell, Georg Martius
Abstract
Prior knowledge vastly exists in the automation industry, especially for tasks like pick-and-place, where simple programmatic demonstrations with online generation ability can be acquired easily. How to learn a policy faster with higher flexibility and generalization ability based on these demonstrations is a question to be answered. End-to-end target learning and imitation learning are widely discussed in previous works. Here, we focus on the online generation ability of the demonstration and propose a demo injection method based on actor-critic off-policy reinforcement learning (RL) for the interaction and policy optimization phase. We conduct experiments and an ablation study based on four research questions around a two-finger reactive grasping task with a Panda robot. The result shows our proposed injection method increases the training stability, strongly reduces the time to convergence and benefits sim-2-real transfer with smooth motion.
BibTeX
@inproceedings{iros2025_dbmpodemonstrati,
title = {DB-MPO: Demonstration Boosted Reactive Grasping For Two-Finger Gripper},
author = {Boya Zhang and Andreas Zell and Georg Martius},
booktitle = {IROS 2025},
year = {2025}
}