2024
Augmented Commonsense Knowledge for Remote Object Grounding
AAAI 2024technical
The vision-and-language navigation (VLN) task necessitates an agent to perceive the surroundings, follow natural language instructions, and act in photo-realistic unseen environments. Most of the existing methods employ the entire image or object features to represent navigable viewpoints. However,…