← Search

Jin-Cheng Jhang

3 accepted papers

2026

Enhancing Part-Level Point Grounding for Any Open-Source MLLMs

CVPR 2026

Visual grounding aims to associate free-form textual queries with specific regions in an image. While recent Multimodal Large Language Models (MLLMs) have demonstrated promising capabilities in this domain, they primarily excel at object-level grounding and often struggle with part-level grounding--

Cited by 0SourceScholar
2026

Pointing at Parts: Training-Free Few-Shot Grounding in Multimodal LLMs

CVPR 2026

Part-level pointing is important for fine-grained interaction and reasoning, yet existing Multimodal Large Language Models (MLLMs) remain limited to instance-level pointing. Part-level pointing presents unique challenges: annotation is costly, parts are long-tail distributed, and many are difficult

Cited by 0SourceScholar
2024

No More Ambiguity in 360deg Room Layout via Bi-Layout Estimation

CVPR 2024poster

Inherent ambiguity in layout annotations poses significant challenges to developing accurate 360deg room layout estimation models. To address this issue we propose a novel Bi-Layout model capable of predicting two distinct layout types. One stops at ambiguous regions while the other extends to encom…

Cited by 5SourcePDFScholar