← Search

Sergio Povoli

2 accepted papers

2026

Obstruction Reasoning for Robotic Grasping

CVPR 2026

Successful robotic grasping in cluttered environments not only requires a model to visually ground a target object but also to reason about obstructions that must be cleared beforehand. While current vision-language embodied reasoning models show emergent spatial understanding, they remain limited i

Cited by 0SourceScholar
2025

Free-form language-based robotic reasoning and grasping

IROS 2025

Performing robotic grasping from a cluttered bin based on human instructions is a challenging task, as it requires understanding both the nuances of free-form language and the spatial relationships between objects. Vision-Language Models (VLMs) trained on web-scale data, such as GPT-4o, have demonst

Cited by 8SourcecodeScholar