← Search

Yifei Yin

3 accepted papers

2026

Activating Visual Context and Commonsense Reasoning Through Masked Prediction in VLMs

AAAI 2026technical

Recent breakthroughs in reasoning models have markedly advanced the reasoning capabilities of large language models, particularly via training on tasks with verifiable rewards. Yet, a significant gap persists in their adaptation to real-world multimodal scenarios, most notably, vision-language tasks

Cited by 0SourcePDFScholar
2023

Hi4D: 4D Instance Segmentation of Close Human Interaction

CVPR 2023poster

We propose Hi4D, a method and dataset for the auto analysis of physically close human-human interaction under prolonged contact. Robustly disentangling several in-contact subjects is a challenging task due to occlusions and complex shapes. Hence, existing multi-view systems typically fuse 3D surface…

2021

Progressive Co-Teaching for Ambiguous Speech Emotion Recognition

ICASSP 2021accepted

Speech emotion recognition is a challenging task due to the ambiguity of emotion, which makes it difficult to learn the features of emotion data using machine learning algorithms. However, previous studies conventionally ignore the ambiguity of emotion and treat the emotion data as the same difficul…

Cited by 0SourceScholar