Uni-Zipper: A Multi-modal Perception Framework of Deformable Objects with Unpaired Data
Qian Xie, Yanmin Zhou, Wei Wang, Yiyang Jin, Zhipeng Wang, Rong Jiang, Xin Li, Hongrui Sang
Abstract
Multi-modal perception plays a crucial role in preventing deformation and damage during the robotic manipulation of deformable objects. However, integrating new heterogeneous modalities into existing robotic perception frameworks remains a significant challenge, primarily due to the need for massive amounts of paired data. In this paper, we propose Uni-Zipper, a scalable multi-modal fusion framework designed to expand new modalities with the help of semantic enhancement without relying on paired data. Uni-Zipper consists of a tokenizer that projects various modalities into a shared embedding space, a summary word embedding layer with a feature dictionary, a modality alignment space, and dynamic reconfigurable task heads. To facilitate efficient integration and extension of new modalities, the Zipper alignment mechanism is employed, effectively bridging the modality gap between different input types. Our experimental results demonstrate that Uni-Zipper successfully fuses four modalities and enhances performance in downstream tasks. Despite a 12% decrease in parameter count, Uni-Zipper maintains comparable performance.
BibTeX
@inproceedings{iros2025_unizipperamultim,
title = {Uni-Zipper: A Multi-modal Perception Framework of Deformable Objects with Unpaired Data},
author = {Qian Xie and Yanmin Zhou and Wei Wang and Yiyang Jin and Zhipeng Wang and Rong Jiang and Xin Li and Hongrui Sang and Bin He},
booktitle = {IROS 2025},
year = {2025}
}