2025
VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
NeurIPS 2025poster
Vision-Language-Action (VLA) models have demonstrated strong multi-modal reasoning capabilities, enabling direct action generation from visual perception and language instructions in an end-to-end manner. However, their substantial computational cost poses a challenge for real-time robotic control,…