RA-L 20260 citations

Collaborative Map-Based and Route-Based Policy Learning for Continuous Vision-and-Language Navigation

Jiewen Hou, Meina Kan, Lixuan Zhang, Hao Liang, Shiguang Shan, Xilin Chen

Abstract

Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires an agent to follow language instructions to reach a target in unseen, 3D environments. A powerful VLN-CE agent requires two crucial abilities during cross-modal planning: spatial reasoning to explore towards the target location in partially observable environments, and procedural alignment to match navigation routes with language instructions for sequential guidance. Existing cross-modal planning policies are divided into two types, each focusing on one of these capabilities. The map-based policy promotes spatial reasoning by grounding instructions on graph representations, while the route-based policy favors procedural alignment by aligning sequential observations with language instructions. However, these policies are typically studied independently, making agents that rely on a single crucial ability unable to plan effectively in complex scenes. Inspired by human navigation, we propose a collaborative policy learning framework that integrates the advantages of both policies for cross-modal planning. This framework includes three processes: Spatio-Procedural Topological Mapping, which constructs a multiplex graph to support the learning of map-based and route-based policies; Dual-Stream Encoding, which performs cross-modal encoding for both policies in parallel; and Hierarchical Policy Integration, which fuses both policies via feature-level collaboration and logit-level fusion. Extensive experiments on the VLN-CE datasets confirm the effectiveness of our framework.

BibTeX
@inproceedings{ral2026_collaborativemap,
  title = {Collaborative Map-Based and Route-Based Policy Learning for Continuous Vision-and-Language Navigation},
  author = {Jiewen Hou and Meina Kan and Lixuan Zhang and Hao Liang and Shiguang Shan and Xilin Chen},
  booktitle = {RA-L 2026},
  year = {2026}
}
Collaborative Map-Based and Route-Based Policy Learning for Continuous Vision-and-Language Navigation · RA-L 2026