ROD-VLM: A Framework of Real-time Robotic Perception, Reasoning and Manipulation
Yinkai Zhu, Xinbei Wang, Feilin Yu, Tianjiao Lei, Yizhuo Sun
Abstract
In recent years, Vision-Language Models (VLMs) have exhibited powerful capacity of reasoning, decomposing long-horizon tasks and motion planning in robotic manipulation tasks. However, the current operating speed of VLMs has limited the interaction frequency of users and the model to several seconds, which disables the real-time perception of environmental changes when executing tasks released by VLM. We propose Real-time Object Detection - VLM (ROD-VLM), a novel framework which combines classical Object Detection Algorithm YOLO-v5x with VLM to achieve the real-time robotic environmental perception, reasoning and manipulation. Specifically, we introduce the concept of key frame to VLM model, capturing the crucial information through object detection algorithm to assist VLM in perceiving varying environment. Our comprehensive real-world experiments show that ROD-VLM possess an excellent capability in real-time environmental understanding, decision-making and action executing.
BibTeX
@inproceedings{iros2025_rodvlmaframework,
title = {ROD-VLM: A Framework of Real-time Robotic Perception, Reasoning and Manipulation},
author = {Yinkai Zhu and Xinbei Wang and Feilin Yu and Tianjiao Lei and Yizhuo Sun},
booktitle = {IROS 2025},
year = {2025}
}