From Thousands to Billions: 3D Visual Language Grounding via Render-Supervised Distillation from 2D VLMs
3D vision-language grounding faces a fundamental data bottleneck: while 2D models train on billions of images, 3D models have access to only thousands of labeled scenes--a six-order-of-magnitude gap that severely limits performance. We introduce \textbf{\emph{LIFT-GS}}, a practical distillation tech…