← Search

Zongtao He

7 accepted papers

2024

Enhanced Language-guided Robot Navigation with Panoramic Semantic Depth Perception and Cross-modal Fusion

IROS 2024poster

Integrating visual observation with linguistic instruction holds significant promise for enhancing robot navigation across unstructured environments and enriches the human-robot interaction experience. However, while panoramic RGB views furnish robots with extensive environmental visuals, current me…

Cited by 0SourcecodeScholar
2024

Multimodal Evolutionary Encoder for Continuous Vision-Language Navigation

IROS 2024poster

Can multimodal encoder evolve when facing increasingly tough circumstances? Our work investigates this possibility in the context of continuous vision-language navigation (continuous VLN), which aims to navigate robots under linguistic supervision and visual feedback. We propose a multimodal evoluti…

Cited by 0SourcecodeScholar
2024

Vision-and-Language Navigation via Causal Learning

CVPR 2024poster

In the pursuit of robust and generalizable environment perception and language understanding the ubiquitous challenge of dataset bias continues to plague vision-and-language navigation (VLN) agents hindering their performance in unseen environments. This paper introduces the generalized cross-modal…

2023

A Dual Semantic-Aware Recurrent Global-Adaptive Network for Vision-and-Language Navigation

IJCAI 2023poster

Vision-and-Language Navigation (VLN) is a realistic but challenging task that requires an agent to locate the target region using verbal and visual cues. While significant advancements have been achieved recently, there are still two broad limitations: (1) The explicit information mining for signifi…

2023

Multiple Thinking Achieving Meta-Ability Decoupling for Object Navigation

ICML 2023poster

We propose a meta-ability decoupling (MAD) paradigm, which brings together various object navigation methods in an architecture system, allowing them to mutually enhance each other and evolve together. Based on the MAD paradigm, we design a multiple thinking (MT) model that leverages distinct thinki…

Cited by 10SourcePDFScholar
2023

Search for or Navigate to? Dual Adaptive Thinking for Object Navigation

ICCV 2023poster

"Search for" or "Navigate to"? When we find a specific object in an unknown environment, the two choices always arise in our subconscious mind. Before we see the target, we search for the target based on prior experience. Once we have seen the target, we can navigate to it by remembering the target…

Cited by 21PDFScholar