2026
D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation
CVPR 2026
Embodied agents face a critical dilemma that end-to-end models lack interpretability and explicit 3D reasoning, while modular systems ignore cross-component interdependencies and synergies. To bridge this gap, we propose the Dynamic 3D Vision-Language-Planning Model (D3D-VLP). Our model introduces t