2026
HiMe: Hierarchical Embodied Memory for Long-Horizon Vision-Language-Action Control
ICML 2026poster
Current Vision-Language-Action (VLA) models excel at robotic manipulation but often struggle with non-Markovian tasks requiring long-term memory and reasoning due to their reliance on immediate observations. Existing solutions face a frequency-competence paradox, where high-performance models are to…