← Search

Dong An

6 accepted papers

2026

Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation

CVPR 2026

Vision-and-Language Navigation (VLN) requires the agent to navigate based on natural instructions. This task is challenging due to partial observability, which makes it difficult to align perception with language. Recent methods mitigate this by imagining future scenes, yet they rely on vision-based

Cited by 0SourcecodeScholar
2025

3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering

IROS 2025

With the growing need for diverse and scalable data in indoor scene tasks, such as question answering and dense captioning, we propose 3D-MoRe, a novel paradigm designed to generate large-scale 3D-language datasets by lever-aging the strengths of foundational models. The framework integrates key com

Cited by 12SourcecodeScholar
2025

Learning Fine-Grained Alignment for Aerial Vision-Dialog Navigation

AAAI 2025technical

Aerial Vision-Dialog Navigation (AVDN) is a new task that requires drones to navigate to a target location based on human-robot dialog history. This paper focuses on the critical fine-grained cross-modal alignment problem in AVDN, requiring the drone to align language entities with visual landmarks…

2025

NavBench: Probing Multimodal Large Language Models for Embodied Navigation

NeurIPS 2025poster

Multimodal Large Language Models (MLLMs) have demonstrated strong generalization in vision-language tasks, yet their ability to understand and act within embodied environments remains underexplored. We present NavBench, a benchmark to evaluate the embodied navigation capabilities of MLLMs under zero…

Cited by 0SourceScholar
2021

Landmark-RxR: Solving Vision-and-Language Navigation with Fine-Grained Alignment Supervision

NeurIPS 2021poster

In Vision-and-Language Navigation (VLN) task, an agent is asked to navigate inside 3D indoor environments following given instructions. Cross-modal alignment is one of the most critical challenges in VLN because the predicted trajectory needs to match the given instruction accurately. In this paper,…