2026
BFA++: Hierarchical Best-Feature-Aware Token Prune for Multi-View Vision Language Action Model
RA-L 2026
Vision-Language-Action (VLA) models have achieved significant breakthroughs by leveraging Large Vision Language Models (VLMs) to jointly interpret instructions and visual inputs. However, the substantial increase in visual tokens, particularly from multi-view inputs, poses serious challenges to real