2024
SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding
ECCV 2024poster
"3D vision-language (3dvl) grounding, which aims to align language with 3D physical environments, stands as a cornerstone in developing embodied agents. In comparison to recent advancements in the 2D domain, grounding language in 3D scenes faces two significant challenges: (i) the scarcity of paired…