← Search

Francesco Giuliari

7 accepted papers

2026

Obstruction Reasoning for Robotic Grasping

CVPR 2026

Successful robotic grasping in cluttered environments not only requires a model to visually ground a target object but also to reason about obstructions that must be cleared beforehand. While current vision-language embodied reasoning models show emergent spatial understanding, they remain limited i

Cited by 0SourceScholar
2025

Free-form language-based robotic reasoning and grasping

IROS 2025

Performing robotic grasping from a cluttered bin based on human instructions is a challenging task, as it requires understanding both the nuances of free-form language and the spatial relationships between objects. Vision-Language Models (VLMs) trained on web-scale data, such as GPT-4o, have demonst

Cited by 8SourcecodeScholar
2025

Functionality Understanding and Segmentation in 3D Scenes

CVPR 2025highlight

Understanding functionalities in 3D scenes involves interpreting natural language descriptions to locate functional interactive objects, such as handles and buttons, in a 3D environment. Functionality understanding is highly challenging, as it requires both world knowledge to interpret language and…

2025

Reasoning in Visual Navigation of End-to-end Trained Agents: A Dynamical Systems Approach

CVPR 2025highlight

Progress in Embodied AI has made it possible for end-to-end-trained agents to navigate in photo-realistic environments with high-level reasoning and zero-shot or language-conditioned behavior, but evaluations and benchmarks are still dominated by simulation. In this work, we focus on the fine-graine…

2024

DiffAssemble: A Unified Graph-Diffusion Model for 2D and 3D Reassembly

CVPR 2024poster

Reassembly tasks play a fundamental role in many fields and multiple approaches exist to solve specific reassembly problems. In this context we posit that a general unified model can effectively address them all irrespective of the input data type (image 3D etc.). We introduce DiffAssemble a Graph N…

2022

Spatial Commonsense Graph for Object Localisation in Partial Scenes

CVPR 2022poster

We solve object localisation in partial scenes, a new problem of estimating the unknown position of an object (e.g. where is the bag?) given a partial 3D scan of a scene. The proposed solution is based on a novel scene graph model, the Spatial Commonsense Graph (SCG), where objects are the nodes and…

Cited by 23PDFcodeScholar
2021

POMP++: Pomcp-based Active Visual Search in unknown indoor environments

IROS 2021poster

In this paper, we focus on the problem of learning online an optimal policy for Active Visual Search (AVS) of objects in unknown indoor environments. We propose POMP++, a planning strategy that introduces a novel formulation on top of the classic Partially Observable Monte Carlo Planning (POMCP) fra…

Cited by 17SourceScholar