ST-HNet: A CNN-LSM Hybrid Architecture for Spatio-Temporal Feature Learning in Event-Based Visual Place Recognition
Xun Xiao, Shasha Guo, Tie Junbo, Jingyue Zhao, Ziqi Wang, Yuan Li, Jingzhuo Yuan, Qiang Dou
Abstract
Visual Place Recognition (VPR) based on Dynamic Vision Sensors (DVSs) has gained attention due to their high temporal resolution and robustness under challenging lighting conditions. However, the sparse and asynchronous event stream output of DVS introduces unique challenges for effective VPR. In this paper, we propose ST-HNet, a novel framework for VPR that introduces improvements in event representation, spatio-temporal feature extraction, and loss design. Specifically, we introduce a compact and efficient event representation called Bipolar Binary Voxel Grid (BBVG). Then, we propose a hybrid feature extractor that combines a Convolutional Neural Network (CNN) for spatial encoding and a Liquid State Machine (LSM) for temporal aggregation. We refer to this combination as a CNN-LSM hybrid architecture. Moreover, we introduce a soft-margin triplet loss to better accommodate the gradual transitions between nearby locations in the event-based VPR task. Extensive experiments conducted on the Brisbane-Event-VPR and DDD20 datasets demonstrate that our method outperforms state-of-the-art approaches, achieving improvements of 11% and 23% in Recall@1 performance, respectively.