← Search

Yunyao Mao

7 accepted papers

2025

Leveraging Visual Captions for Enhanced Zero-Shot HOI Detection

ICASSP 2025accepted

Zero-shot Human-Object Interaction (HOI) detection aims to identify both seen and unseen HOI categories in an image. Most existing methods rely on semantic knowledge distilled from CLIP to find novel interactions but fail to fully exploit the powerful generalization ability of vision-language models…

Cited by 0SourceScholar
2023

CLIP4HOI: Towards Adapting CLIP for Practical Zero-Shot HOI Detection

NeurIPS 2023poster

Zero-shot Human-Object Interaction (HOI) detection aims to identify both seen and unseen HOI categories. A strong zero-shot HOI detector is supposed to be not only capable of discriminating novel interactions but also robust to positional distribution discrepancy between seen and unseen categories w…

Cited by 21SourcePDFScholar
2023

Masked Motion Predictors are Strong 3D Action Representation Learners

ICCV 2023poster

In 3D human action recognition, limited supervised data makes it challenging to fully tap into the modeling potential of powerful networks such as transformers. As a result, researchers have been actively investigating effective self-supervised pre-training strategies. In this work, we show that ins…

Cited by 46PDFcodeScholar
2022

CMD: Self-Supervised 3D Action Representation Learning with Cross-Modal Mutual Distillation

ECCV 2022poster

"In 3D action recognition, there exists rich complementary information between skeleton modalities. Nevertheless, how to model and utilize this information remains a challenging problem for self-supervised 3D action representation learning. In this work, we formulate the cross-modal interaction as a…

2022

CMT: Context-Matching-Guided Transformer for 3D Tracking in Point Clouds

ECCV 2022poster

"How to effectively match the target template features with the search area is the core problem in point-cloud-based 3D single object tracking. However, in the literature, most of the methods focus on devising sophisticated matching modules at point-level, while overlooking the rich spatial context…

Cited by 27SourcePDFScholar
2021

Joint Inductive and Transductive Learning for Video Object Segmentation

ICCV 2021poster

Semi-supervised video object segmentation is a task of segmenting the target object in a video sequence given only a mask annotation in the first frame. The limited information available makes it an extremely challenging task. Most previous best-performing methods adopt matching-based transductive r…

Cited by 122PDFcodeScholar