DecoVLN: Decoupling Observation, Reasoning, and Correction for Vision-and-Language Navigation
Vision-and-Language Navigation (VLN) requires agents to follow long-horizon instructions and navigate complex 3D environments. However, existing approaches face two major challenges: constructing an effective long-term memory bank and overcoming the compounding errors problem. To address these issue