← Search

Jiahe Zhao

3 accepted papers

2025

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs

ICCV 2025poster

In video Multimodal Large Language Models (video MLLMs), the visual encapsulation process plays a pivotal role in converting video contents into representative tokens for LLM input. While linear projectors are widely employed for encapsulation, they introduce semantic indistinctness and temporal inc…

2025

HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding

ICCV 2025poster

We propose a new task to benchmark human-in-scene understanding for embodied agents: Human-In-Scene Question Answering (HIS-QA). Given a human motion within a 3D scene, HIS-QA requires the agent to comprehend human states and behaviors, reason about its surrounding environment, and answer human-rela…

2025

un$^2$CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP

NeurIPS 2025poster

Contrastive Language-Image Pre-training (CLIP) has become a foundation model and has been applied to various vision and multimodal tasks. However, recent works indicate that CLIP falls short in distinguishing detailed differences in images and shows suboptimal performance on dense-prediction and vis…

Cited by 0SourcecodeScholar