LE-Object: Language Embedded Object-Level Neural Radiance Fields for Open-Vocabulary Scene
Mengting Wang, Yunzhou Zhang, Xingshuo Wang, Zhiyao Zhang, Zhiteng Li
Abstract
Recent advancements in Visual Language Models (VLMs) have significantly driven research in open-vocabulary 3D scene reconstruction, showcasing strong potential in open-set retrieval and semantic understanding. However, existing approaches face challenges in open-world environments: they either suffer from insufficient precision in semantic segmentation, leading to inadequate fine-grained scene understanding, or they are limited to object-level reconstruction, failing to capture intricate object details and lack applicability in open-world settings. To address these issues, we introduce LE-Object, an object-centric Neural Implicit Radiance Field (NeRF) method for open-world scenarios to achieve fine-grained scene understanding and high-fidelity object reconstruction. LE-Object integrates spatial features (SF) from object point clouds with visual features (VF) from VLMs to perform object association, ensuring spatiotemporal consistency in object mask segmentation, and extends VLM features from 2D images into 3D space, enabling precise open-world semantic inference and detailed object reconstruction. Experimental results demonstrate that LE-Object excels in zero-shot semantic segmentation and open-world object reconstruction, offering innovative solutions for global navigation and local object manipulation in open-world applications.
BibTeX
@inproceedings{icra2025_leobjectlanguage,
title = {LE-Object: Language Embedded Object-Level Neural Radiance Fields for Open-Vocabulary Scene},
author = {Mengting Wang and Yunzhou Zhang and Xingshuo Wang and Zhiyao Zhang and Zhiteng Li},
booktitle = {ICRA 2025},
year = {2025}
}