GIL-3D: U-Shaped Diffusion Transformers for Generalizable 3D Imitation Learning
Imitation learning with 3D vision effectively alleviates the impact of variations in lighting, background, and texture. It exhibits superior robustness compared to 2D-based methods. However, existing 3D imitation learning methods often suffer from performance degradation as the task horizon increase