← Search

Yikang Zhou

4 accepted papers

2026

Grasp Any Region: Prompting MLLM to Understand the Dense World

ICLR 2026poster

While Multimodal Large Language Models (MLLMs) excel at holistic understanding, they struggle with the dense world, i.e., complex scenes requiring fine-grained analysis of intricate details and object inter-relationships. Region-level MLLMs have been a promising step. However, previous attempts are…

Cited by 0SourcecodeScholar
2026

SAMTok: Representing Any Mask with Two Words

CVPR 2026

Pixel-wise capabilities are essential for building interactive intelligent systems. However, pixel-wise multi-modal LLMs (MLLMs) remain difficult to scale due to complex region-level encoders, specialized segmentation decoders, and incompatible training objectives. To address these challenges, we pr

Cited by 0SourcecodeScholar
2025

Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs

ICCV 2025poster

Recent advancements in multimodal large language models (MLLM) have shown a strong ability in visual perception, reasoning abilities, and vision-language understanding. However, the visual matching ability of MLLMs is rarely studied, despite finding the visual correspondence of objects is essential…

2024

Improving Video Segmentation via Dynamic Anchor Queries

ECCV 2024poster

"Modern video segmentation methods adopt feature transitions between anchor and target queries to perform cross-frame object association. The smooth feature transitions between anchor and target queries enable these methods to achieve satisfactory performance when tracking continuously appearing obj…

Cited by 8SourcePDFScholar