2024
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
ICLR 2024poster
3D visual grounding is the ability to localize objects in 3D scenes conditioned by utterances. Most existing methods devote the referring head to localize the referred object directly, causing failure in complex scenarios. In addition, it does not illustrate how and why the network reaches the final…