← Search

Zaiquan Yang

3 accepted papers

2025

Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding

NeurIPS 2025poster

Spatio-temporal video grounding (STVG) aims at localizing the spatio-temporal tube of a video, as specified by the input text query. In this paper, we utilize multimodal large language models (MLLMs) to explore a zero-shot solution in STVG. We reveal two key insights about MLLMs: (1) MLLMs tend to…

Cited by 0SourcecodeScholar
2024

Boosting Weakly Supervised Referring Image Segmentation via Progressive Comprehension

NeurIPS 2024poster

This paper explores the weakly-supervised referring image segmentation (WRIS) problem, and focuses on a challenging setup where target localization is learned directly from image-text pairs. We note that the input text description typically already contains detailed information on how to localize t…

Cited by 2SourcePDFScholar
2022

Learning Prototype via Placeholder for Zero-shot Recognition

IJCAI 2022poster

Zero-shot learning (ZSL) aims to recognize unseen classes by exploiting semantic descriptions shared between seen classes and unseen classes. Current methods show that it is effective to learn visual-semantic alignment by projecting semantic embeddings into the visual space as class prototypes. Ho…