2026
MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming
AAAI 2026technical
Vision-Language Navigation (VLN) tasks often leverage panoramic RGB and depth inputs to provide rich spatial cues for action planning, but these sensors can be costly or less accessible in real-world deployments. Recent approaches based on Vision-Language Action (VLA) models achieve strong results w