IJCAI 20260 citations

FILD-Nav:Vision-and-Language Navigation with Instruction Landmark Features in Continuous Environments

Chuangye Hu, Lulu Liu, Huaiwei Si, Yawen Zhao, Nan Ding

Abstract

Vision-and-language navigation (VLN) requires agents to follow natural language instructions to navigate autonomously in continuous environments. However, existing approaches often lack high-level semantic guidance in waypoint prediction and explicit language–landmark alignment in cross-modal planning. To address these limitations, we propose FILD-Nav, a vision-and-language navigation framework that integrates instruction landmark features. FILD-Nav extracts task-relevant landmarks from instructions and incorporates landmark semantics into both waypoint prediction and topological planning. Specifically, landmark-guided waypoint prediction improves waypoint relevance, while landmark-enhanced cross-modal planning enables more effective long-horizon navigation. Extensive experiments on the VLN-CE benchmark demonstrate that FILD-Nav consistently outperforms prior methods, achieving improvements of 2% in Success Rate (SR), 3% in Success weighted by Path Length (SPL), and 7% in Oracle Success Rate (OSR), particularly in unseen environments.

Robot control, planning, and execution with guarantees: Architectures connecting high-level intent and constraints to low-level trajectoriesLearning to understand, generalize, and explain actions: Robust generalization and transfer across tasks, objects, environments, embodiments, and long-horizon scenariosAIR: Robot control, planning, and execution with guaranteesFoundations of human–robot interaction and assistance: Learning and inference methods for aligning robot behavior with human instructions, demonstrations, and feedback
BibTeX
@inproceedings{ijcai2026_fildnavvisionand,
  title = {FILD-Nav:Vision-and-Language Navigation with Instruction Landmark Features in Continuous Environments},
  author = {Chuangye Hu and Lulu Liu and Huaiwei Si and Yawen Zhao and Nan Ding},
  booktitle = {IJCAI 2026},
  year = {2026}
}
FILD-Nav:Vision-and-Language Navigation with Instruction Landmark Features in Continuous Environments · IJCAI 2026