GP3: A 3D Geometry-Aware Policy with Multi-View Images for Robotic Manipulation
Quanhao Qian, Guoyang Zhao, Gongjie Zhang, Jiuniu Wang, Junlong Gao, Deli Zhao, Ran Xu
Abstract
Effective robotic manipulation relies on a precise understanding of 3D scene geometry, and one of the most straightforward ways to acquire such geometry is through multi-view observations. Motivated by this, we present GP3—a 3D geometry-aware robotic manipulation policy that leverages multi-view input. GP3 employs a spatial encoder to infer dense spatial features from RGB observations, which enable the estimation of depth and camera parameters, leading to a compact yet expressive 3D scene representation tailored for manipulation. This representation is fused with language instructions and translated into continuous actions via a lightweight policy head. We further introduce G-FiLM, which applies language-conditioned FiLM only to cross-view global attention. Comprehensive experiments demonstrate that GP3 consistently outperforms state-of-the-art methods on simulated benchmarks. Furthermore, GP3 transfers effectively to realworld robots in depth-challenging scenes with only minimal fine-tuning. These results highlight GP3 as a practical, sensoragnostic solution for geometry-aware robotic manipulation.