NaVLA$^2$: A Vision-Language-Audio-Action Model for Multimodal Instruction Navigation
Embodied navigation is a fundamental capability for intelligent agents, yet remains challenging in partially observable environments where navigation instructions can be difficult to interpret. However, existing tasks only provide unimodal instructions, which are ambiguous in complex multimodal envi