Multi-policy Grounding and Ensemble Policy Learning for Transfer Learning with Dynamics Mismatch
We propose a new transfer learning algorithm between tasks with different dynamics. The proposed algorithm solves an Imitation from Observation problem (IfO) to ground the source environment to the target task before learning an optimal policy in the grounded environment. The learned policy is deplo…