← Search

Yinkai Zhu

1 accepted papers

2025

ROD-VLM: A Framework of Real-time Robotic Perception, Reasoning and Manipulation

IROS 2025

In recent years, Vision-Language Models (VLMs) have exhibited powerful capacity of reasoning, decomposing long-horizon tasks and motion planning in robotic manipulation tasks. However, the current operating speed of VLMs has limited the interaction frequency of users and the model to several seconds

Cited by 0SourceScholar