← Search

Qunbo Wang

8 accepted papers

2026

SurveilNav: Collaborative Object Goal Navigation with Robot and Surveillance System

ICRA 2026poster

With the growing deployment of surveillance systems in factories, offices, and homes, integrating them with robots offers a promising direction for collaborative and efficient task execution. However, existing approaches largely focus on single-robot scenarios and struggle with multi-view collaborat…

2026

UrbanNav: Learning Language-Guided Embodied Urban Navigation from Web-Scale Human Trajectories

AAAI 2026technical

Navigating complex urban environments using natural language instructions poses significant challenges for embodied agents, including noisy language instructions, ambiguous spatial references, diverse landmarks, and dynamic street scenes. Current visual navigation methods are typically limited to si

Cited by 0SourcePDFScholar
2025

C-NAV: Towards Self-Evolving Continual Object Navigation in Open World

NeurIPS 2025poster

Embodied agents are expected to perform object navigation in dynamic, open-world environments. However, existing approaches typically rely on static trajectories and a fixed set of object categories during training, overlooking the real-world requirement for continual adaptation to evolving scenario…

Cited by 0SourcecodeScholar
2025

COSMO: Combination of Selective Memorization for Low-cost Vision-and-Language Navigation

ICCV 2025poster

Vision-and-Language Navigation (VLN) tasks have gained prominence within artificial intelligence research due to their potential application in fields like home assistants. Many contemporary VLN approaches, while based on transformer architectures, have increasingly incorporated additional component…

2025

TALKER: A Task-Activated Language Model Based Knowledge-Extension Reasoning System

RA-L 2025

Training drones to execute complex collective tasks via multi-agent reinforcement learning presents significant challenges. To address these challenges, this letter introduces the Task-Activated Language model-based Knowledge-Extension Reasoning system. Specifically, we trained drones in two fine-gr

Cited by 1SourceScholar
2024

Self-Bootstrapped Visual-Language Model for Knowledge Selection and Question Answering

EMNLP 2024main

While large pre-trained visual-language models have shown promising results on traditional visual question answering benchmarks, it is still challenging for them to answer complex VQA problems which requires diverse world knowledge. Motivated by the research of retrieval-augmented generation in the…

2024

Soft Knowledge Prompt: Help External Knowledge Become a Better Teacher to Instruct LLM in Knowledge-based VQA

ACL 2024long

LLM has achieved impressive performance on multi-modal tasks, which have received ever-increasing research attention. Recent research focuses on improving prediction performance and reliability (e.g., addressing the hallucination problem). They often prepend relevant external knowledge to the input…

2023

VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset

NeurIPS 2023poster

Vision and text have been fully explored in contemporary video-text foundational models, while other modalities such as audio and subtitles in videos have not received sufficient attention. In this paper, we resort to establish connections between multi-modality video tracks, including Vision, Audi…