NavTr: Object-Goal Navigation With Learnable Transformer Queries
Qiuyu Mao, Jikai Wang, Meng Xu, Zonghai Chen
Abstract
This letter introduces <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Nav</b>igation <bold xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">Tr</b>ansformer (NavTr), a novel framework for object-goal navigation using Transformer queries to enhance the learning and representation of environment states. By integrating semantic information, object positions, and neighborhood information, NavTr creates a unified, comprehensive, and extensible state representation for the object-goal navigating task. In the framework, the Transformer queries implicitly learn inter-object relationships, which facilitates high-level understanding of the environment. Additionally, NavTr implements target-oriented supervisory signals, such as rotation rewards and spatial loss, which improve exploration efficiency in the reinforcement learning framework. NavTr outperforms popular graph-based and Attention-based methods by a large margin in terms of success rate (SR) and success weighted by path length (SPL). Extensive experiments on the AI2-THOR dataset demonstrate the effectiveness of our approach.
BibTeX
@inproceedings{ral2024_navtrobjectgoaln,
title = {NavTr: Object-Goal Navigation With Learnable Transformer Queries},
author = {Qiuyu Mao and Jikai Wang and Meng Xu and Zonghai Chen},
booktitle = {RA-L 2024},
year = {2024}
}