← Search

Ronggang Huang

3 accepted papers

2025

ViewSRD: 3D Visual Grounding via Structured Multi-View Decomposition

ICCV 2025poster

3D visual grounding aims to identify and localize objects in a 3D space based on textual descriptions. However, existing methods struggle with disentangling targets from anchors in complex multi-anchor queries and resolving inconsistencies in spatial descriptions caused by perspective variations.To…

2024

Mask4Align: Aligned Entity Prompting with Color Masks for Multi-Entity Localization Problems

CVPR 2024poster

In Visual Question Answering (VQA) recognizing and localizing entities pose significant challenges. Pretrained vision-and-language models have addressed this problem by providing a text description as the answer. However in visual scenes with multiple entities textual descriptions struggle to distin…

Cited by 0SourcePDFScholar