FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-and-Language Navigation
Achieving human-level performance in Vision-and-Language Navigation (VLN) requires an embodied agent to understand textual instructions, perceive visual observations, and reason over long action sequences. Recent works, such as NavCoT and NavGPT-2, demonstrate the potential of Chain-of-Thought (CoT)