SS-VIO: A Novel Cross Modal State Space Visual Inertial Odometry
Wenyuan Qin, Zhi Lin, Lu Bai, Shaojun Yao, Hao Xu, Xiwang Dong
Abstract
Visual-inertial odometry (VIO) is fundamental to autonomous driving, embodied navigation, and robotics. While geometry based methods dominate the field, learning based approaches are still in their infancy despite promising progress. Current deep learning based VIO methods usually fuse visual images and inertial measurements directly, neglecting intra modal frequency heterogeneity, cross modal dependencies, and temporal continuity in state estimation. We propose a State Space Visual-Inertial Odometry (SS-VIO) framework that fuses modalities through a Mamba-based time-frequency modeling module. The first is a time frequency interaction module with Mamba that captures joint temporal and frequency representations and performs adaptive selection. The second is a cross modal fusion module that integrates inertial cues into visual features through learned weighting. Experiments on the KITTI dataset show that SS-VIO outperforms state of the art methods in both accuracy and efficiency. Additional tests on RealSense D435i quadruped robot data confirm the generalization ability of the framework across different platforms.
BibTeX
@inproceedings{ral2026_ssvioanovelcross,
title = {SS-VIO: A Novel Cross Modal State Space Visual Inertial Odometry},
author = {Wenyuan Qin and Zhi Lin and Lu Bai and Shaojun Yao and Hao Xu and Xiwang Dong},
booktitle = {RA-L 2026},
year = {2026}
}