← Search

Iaroslav Ponomarenko

4 accepted papers

2025

ManipGPT: Is Affordance Segmentation by Large Vision Models Enough for Articulated Object Manipulation?

IROS 2025

Visual actionable affordance has emerged as a transformative approach in robotics, focusing on perceiving interaction areas prior to manipulation. Traditional methods rely on pixel sampling to identify successful interaction samples or processing pointclouds for affordance mapping. However, these ap

Cited by 0SourcecodeScholar
2025

Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation

CVPR 2025poster

In robotic manipulation, task goals can be conveyed through various modalities, such as language, goal images, and goal videos. However, natural language can be ambiguous, while images or videos may offer overly detailed specifications. To address these challenges, we propose a novel approach using…

Cited by 0SourcePDFScholar
2025

SpatialBot: Precise Spatial Understanding with Vision Language Models

ICRA 2025

Vision Language Models (VLMs) have achieved impressive performance in 2D image understanding; however, they still struggle with spatial understanding, which is fundamental to embodied AI. In this paper, we propose SpatialBot, a model designed to enhance spatial understanding by utilizing both RGB an

Cited by 167SourcecodeScholar
2024

ManipVQA: Injecting Robotic Affordance and Physically Grounded Information into Multi-Modal Large Language Models

IROS 2024poster

While the integration of Multi-modal Large Language Models (MLLMs) with robotic systems has significantly improved robots’ ability to understand and execute natural language instructions, their performance in manipulation tasks remains limited due to a lack of robotics-specific knowledge. Convention…

Cited by 27SourcecodeScholar