EG-3DVG: Expression and Geometry Aware Grounding Decoder for 3D Visual Grounding
Despite recent progress in 3D visual grounding, existing methods still struggle with three core challenges: 1) cross-modal misalignment that prevents textual cues from being reliably delivered to visual representations, 2) intra-class confusion arising from insufficient understanding of fine-grained