RA-L 20250 citations

Zero-Shot Adaptation at Task-Level via Coarse-to-Fine Policy Refinement and Holistic-Local Contrastive Representation

Zhengwei Li, Zhenyang Lin, Yurou Chen, Lu Zhang, Zhiyong Liu

Abstract

Meta-reinforcement learning offers a mechanism for zero-shot adaptation, enabling agents to handle new tasks with parametric variation in real-world environments. However, existing methods still struggle with task-level adaptation, which demands generalization beyond simple variations within tasks, thereby limiting their practical effectiveness. This limitation stems from a failure to leverage the hierarchical characteristics inherent in task-level adaptation, resulting in inadequate task representations and constrained policy adaptation mechanisms. To address these challenges, we propose a <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">C</b>oarse-to-<bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">F</b>ine p<bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">O</b>licy refinement process combined with a <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">H</b>olistic-<bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">L</b>ocal contrastive <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">R</b>epresentation (<bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">CFOHLR</b>) method to enable effective zero-shot policy adaptation. Specifically, we utilize task language instructions as prior knowledge to select different parameterized modules as a coarse policy. This coarse policy is subsequently refined by a fine policy generated through a hypernetwork, which produces a task-aware policy based on task representations. Additionally, since task representations can influence the effectiveness of task-aware policies, we employ contrastive learning from both holistic and local perspectives to enhance these representations for more effective policy generation. Experimental results demonstrate that our method significantly improves learning efficiency and zero-shot adaptation on new tasks, outperforming previous methods on the Meta-World ML-10 and ML-45 benchmarks.

BibTeX
@inproceedings{ral2025_zeroshotadaptati,
  title = {Zero-Shot Adaptation at Task-Level via Coarse-to-Fine Policy Refinement and Holistic-Local Contrastive Representation},
  author = {Zhengwei Li and Zhenyang Lin and Yurou Chen and Lu Zhang and Zhiyong Liu},
  booktitle = {RA-L 2025},
  year = {2025}
}
Zero-Shot Adaptation at Task-Level via Coarse-to-Fine Policy Refinement and Holistic-Local Contrastive Representation · RA-L 2025