PIDiff: Integrating a High-Performance Transformer Into Diffusion Models for Robust and Efficient Imitation Learning
Yunzhi Huang, Peng Jin, Liang Han, Yingzhao Li, Xudong Hou
Abstract
Imitation learning is a critical approach for robots to acquire skills by mimicking human behavior. However, traditional imitation learning frameworks often exhibit poor action prediction accuracy and low robustness when handling complex tasks. To tackle these limitations, we propose the PIDiff policy that integrates an enhanced linear-complexity Transformer with a diffusion model and introduces a PointNet-based encoder to efficiently extract visual features, thereby enhancing the policy's understanding of the environment and improving learning efficiency and robustness. Specifically, we introduce the designed SubLink Block and CrossLink Block, enabling the improved Transformer to serve as the noise predictor in the diffusion strategy while retaining the diffusion model's ability to capture multi-modal distributions. To comprehensively evaluate the effectiveness of our method, we conducted systematic experiments across multiple tasks on three robotic manipulation benchmark platforms in simulation environments and performed evaluations in various real-world scenarios. The experimental results demonstrate that our approach significantly outperforms baseline methods, achieving higher success rates, faster inference speeds, reduced GPU memory usage, and enhanced robustness.
BibTeX
@inproceedings{ral2026_pidiffintegratin,
title = {PIDiff: Integrating a High-Performance Transformer Into Diffusion Models for Robust and Efficient Imitation Learning},
author = {Yunzhi Huang and Peng Jin and Liang Han and Yingzhao Li and Xudong Hou},
booktitle = {RA-L 2026},
year = {2026}
}