IntelliRMS: A Robotic Manipulation System for Domain-Specific Tasks Using Vision and Language Foundational Models
Chandan Kumar Singh, Devesh Kumar, Vipul Sanap, Mayank Khandelwal, Rajesh Sinha
Abstract
Recent advancements in large language models (LLMs) have significantly enhanced machines' ability to understand and follow human instructions. In many tasks, LLMs have demonstrated performance that rivals human-level common sense. However, directly applying LLMs to domain-specific use cases, such as robotic pick-and-place, remains a challenge. Tasks that are intuitive for humans, who rely on prior knowledge and skills, become complex for robots. Industrial robotic applications like pick-and-place require a high degree of accuracy, often exceeding 90 %. In response to these challenges in domain-specific applications, we propose IntelliRMS, a novel system-oriented architecture for instruction-following robotic manipulation. The IntelliRMS synergizes the linguistic and open-vocabulary visual capabilities of foundational models to arrive at an accurate, robust and scalable system. Further, we demonstrate the effectiveness of IntelliRMS in a real-world industrial Bin-picking scenario within the retail sector, validating its performance with a comprehensive dataset.
BibTeX
@inproceedings{icra2025_intellirmsarobot,
title = {IntelliRMS: A Robotic Manipulation System for Domain-Specific Tasks Using Vision and Language Foundational Models},
author = {Chandan Kumar Singh and Devesh Kumar and Vipul Sanap and Mayank Khandelwal and Rajesh Sinha},
booktitle = {ICRA 2025},
year = {2025}
}