HeStIa: Asynchronous Embodied Dynamic Locomotion Learning for Walking Robots through Multimodal Large Language Models
Xiaoyu Tan, Haoyu Wang, Sijia Li, Yinghui Xu, Xihe Qiu
Abstract
The control of locomotion in walking robots with various architectural designs presents significant challenges. While existing approaches primarily rely on low-level state information and isolated visual features, lacking the high-level semantic understanding that humans use to reason about movement and posture, we propose HeStIa, a novel framework that bridges visual perception, natural language understanding, and robotic control through multimodal learning. Our framework leverages multimodal large language models (MLLMs) to establish a semantic bridge between visual observations and motion control, enabling robots to understand and adjust their locomotion through both visual and linguistic modalities. By leveraging multimodal large language models (MLLMs), HeStIa establishes a semantic connection between visual observations and motion control, enabling robots to comprehend and adapt their locomotion through both visual and linguistic modalities. Our approach extracts spatiotemporal visual features from robot movements and transforms them into a cross-modal embedding space shared with textual descriptions. HeStIa incorporates an innovative vision-language-motion fusion mechanism to provide informed, context-aware feedback during the dynamic learning process. Through an asynchronous design, HeStIa effectively mitigates the inference delays typically associated with MLLMs while maintaining real-time performance in dynamic scenarios. The cross-modal representations learned by HeStIa facilitate more intuitive and efficient locomotion learning by grounding visual observations in natural language descriptions. Our comprehensive evaluation shows substantial improvements in motion naturalness, stability, and adaptability across diverse environmental conditions.
BibTeX
@inproceedings{iros2025_hestiaasynchrono,
title = {HeStIa: Asynchronous Embodied Dynamic Locomotion Learning for Walking Robots through Multimodal Large Language Models},
author = {Xiaoyu Tan and Haoyu Wang and Sijia Li and Yinghui Xu and Xihe Qiu},
booktitle = {IROS 2025},
year = {2025}
}