STAF-Navi: Vision-Based Spatio-Temporal Attention Fusion Navigation Framework
Haowen Zhang, Fanghong Liu, Chaoyu Zhang, Qiuze Yu
Abstract
In cluttered, unknown, and partially observable environments, Unmanned Aerial Vehicle (UAV) navigation encounters formidable challenges. To address these challenges, we propose an innovative spatio-temporal attention fusion navigation framework called STAF-Navi. The framework integrates spatio-temporal attention mechanisms to model sequential dependencies. It captures spatial and temporal correlations from historical observations and actions to improve navigation and obstacle avoidance. STAF-Navi employs deep collision encoding to compress high-dimensional depth images into informative low-dimensional latent states, and a single-site Transformer to model historical sensor inputs and states, enhancing the utility of current observations. By exploiting temporal dependencies, this integration enables early braking and stable hovering. Extensive simulation experiments show that the framework increases the navigation success rate by 10% and improves path efficiency by 7%. Finally, the successful deployment of the proposed strategy in real-world scenarios validates its effectiveness.
BibTeX
@inproceedings{ral2025_stafnavivisionba,
title = {STAF-Navi: Vision-Based Spatio-Temporal Attention Fusion Navigation Framework},
author = {Haowen Zhang and Fanghong Liu and Chaoyu Zhang and Qiuze Yu},
booktitle = {RA-L 2025},
year = {2025}
}