DiTEA: Mixture-of-Experts for Vision-Language-Action Model in Robotic Manipulation
The current diffusion-based Vision-Language-Action (VLA) models have faster inference speed and the ability to solve the action muti-modality problem in robot manipulation tasks compared to traditional autoregressive models after large-scale pre-training and post-training. However, the diffusion-bas