HiMRAG: Hierarchical Multimodal Retrieval-Augmented Generation for Robot Task Planning
Zhuoyi Zhang, Yixin Han, Renjun Li, Xiao Li
Abstract
Data-driven methods offer promising solutions for robotic manipulation in human-centric environments, but enabling robots to operate complex appliances from natural language remains a significant challenge. The ambiguity of human instructions and the visual diversity of real-world objects make it difficult to generate precise and reliable action sequences. In this paper, we propose a hierarchical multimodal Retrieval-Augmented Generation (RAG) framework that fuses visual perception with language understanding. Our framework uses a vision-based module to identify an appliance and its documentation from a snapshot, then leverages a task-oriented RAG pipeline to process user instructions, retrieve relevant manual sections, and generate executable action sequences. We train and validate this framework on a custom dataset of microwave oven operation tasks and demonstrate its effectiveness, robustness, and practical viability through extensive virtual and physical experiments on a robotic platform.
BibTeX
@inproceedings{ral2026_himraghierarchic,
title = {HiMRAG: Hierarchical Multimodal Retrieval-Augmented Generation for Robot Task Planning},
author = {Zhuoyi Zhang and Yixin Han and Renjun Li and Xiao Li},
booktitle = {RA-L 2026},
year = {2026}
}