GIL-3D: U-Shaped Diffusion Transformers for Generalizable 3D Imitation Learning
Xiyue Wang, Aoran Mei, Linzhi Wu, Zhongxue Gan, Guo-Niu Zhu
Abstract
Imitation learning with 3D vision effectively alleviates the impact of variations in lighting, background, and texture. It exhibits superior robustness compared to 2D-based methods. However, existing 3D imitation learning methods often suffer from performance degradation as the task horizon increases, primarily due to insufficient modeling of temporal dependencies and misalignment between states and actions. To address these challenges, we propose GIL-3D, a novel framework that combines the stable training dynamics of diffusion models with the global temporal modeling capability of Transformers. To further enhance multi-scale temporal dependency modeling, we explore alternative U-shaped hierarchical architectures and introduce strided skip connections to reduce redundancy in dense feature fusion. Moreover, we present a full-sequence joint attention mechanism to strengthen cross-modal interactions and improve the consistency between visual perception and action generation. Extensive experiments demonstrate that our model consistently outperforms existing baselines, achieving a 16.1% improvement in success rate on simulated benchmarks and a 15.83% improvement on real-world manipulation tasks. In addition, comprehensive generalization studies show that GIL-3D maintains robust performance in previously unseen scenarios.
BibTeX
@inproceedings{ral2026_gil3dushapeddiff,
title = {GIL-3D: U-Shaped Diffusion Transformers for Generalizable 3D Imitation Learning},
author = {Xiyue Wang and Aoran Mei and Linzhi Wu and Zhongxue Gan and Guo-Niu Zhu},
booktitle = {RA-L 2026},
year = {2026}
}