← Search

Haitong Wang

6 accepted papers

2026

Mobile Robot Navigation Using Hand-Drawn Maps: A Vision Language Model Approach

ICRA 2026poster

Hand-drawn maps can be used to convey navigation instructions between humans and robots in a natural and efficient manner. However, these maps can often contain inaccuracies such as scale distortions and missing landmarks which present challenges for mobile robot navigation. This paper introduces a …

2026

X-Nav: Learning End-To-End Cross-Embodiment Navigation for Mobile Robots

ICRA 2026poster

Existing navigation methods are primarily designed for specific robot embodiments, limiting their generalizability across diverse robot platforms. In this paper, we introduce X-Nav, a novel framework for end-to-end cross-embodiment navigation where a single unified policy can be deployed across vari…

2025

Find Everything: A General Vision Language Model Approach to Multi-Object Search

IROS 2025

Efficient navigation and search in unknown environments for multiple objects is a fundamental challenge in robotics, particularly in applications such as warehouse management, domestic assistance, and search-and-rescue. The Multi-Object Search (MOS) problem involves navigating to a sequence of locat

Cited by 8SourcecodeScholar
2025

Mobile Robot Navigation Using Hand-Drawn Maps: A Vision Language Model Approach

RA-L 2025

Hand-drawn maps can be used to convey navigation instructions between humans and robots in a natural and efficient manner. However, these maps can often contain inaccuracies such as scale distortions and missing landmarks which present challenges for mobile robot navigation. This paper introduces a

Cited by 11SourceScholar
2024

NavFormer: A Transformer Architecture for Robot Target-Driven Navigation in Unknown and Dynamic Environments

RA-L 2024

In unknown cluttered and dynamic environments such as disaster scenes, mobile robots need to perform target-driven navigation in order to find people or objects of interest, where the only information provided about these targets are images of the individual targets. In this letter, we introduce Nav

Cited by 34SourceScholar