Collaborative Map-Based and Route-Based Policy Learning for Continuous Vision-and-Language Navigation
Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires an agent to follow language instructions to reach a target in unseen, 3D environments. A powerful VLN-CE agent requires two crucial abilities during cross-modal planning: spatial reasoning to explore towards the target locat