AAAI 2026technical0 citations

NaVLA$^2$: A Vision-Language-Audio-Action Model for Multimodal Instruction Navigation

Jugang Fan, Peihao Chen, Changhao Li, Qing Du, Jian Chen, Mingkui Tan

Abstract

Embodied navigation is a fundamental capability for intelligent agents, yet remains challenging in partially observable environments where navigation instructions can be difficult to interpret. However, existing tasks only provide unimodal instructions, which are ambiguous in complex multimodal environments with multiple similar objects, and may result in misinterpretation and navigation failure. To overcome these limitations, we introduce MINav, a novel task where the navigation path is precisely described by a multimodal instruction. The instruction provides multimodal cues, including object categories, RGB images, language descriptions, and auditory descriptions, which help the agent to disambiguate and ground objects in the environment and navigate effectively. We further construct a large-scale dataset of 43.9K navigation episodes using a two-stage pipeline that first annotates multimodal references of objects and then synthesizes diverse multimodal instructions. We find that existing methods struggle on MINav task, indicating substantial room for improvement in agents

BibTeX
@inproceedings{aaai2026_navla2avisionlan,
  title = {NaVLA$^2$: A Vision-Language-Audio-Action Model for Multimodal Instruction Navigation},
  author = {Jugang Fan and Peihao Chen and Changhao Li and Qing Du and Jian Chen and Mingkui Tan},
  booktitle = {AAAI 2026},
  year = {2026}
}
NaVLA$^2$: A Vision-Language-Audio-Action Model for Multimodal Instruction Navigation · AAAI 2026