MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
Pre-trained Vision-Language-Action (VLA) models have achieved remarkable success in improving robustness and generalization for end-to-end robotic manipulation. However, these models struggle with long-horizon tasks due to their lack of memory and reliance solely on immediate sensory inputs. To addr…