← Search

Pengfei Yan

6 accepted papers

2026

UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models

CVPR 2026

With the advancement of multi-modal Large Language Models (LLMs), Video LLMs have been further developed to perform on holistic and specialized video understanding. However, existing works are limited to specialized video understanding tasks, failing to achieve a comprehensive and multi-grained vide

Cited by 0SourceScholar
2025

ARIG: Autoregressive Interactive Head Generation for Real-time Conversations

ICCV 2025poster

Face-to-face communication, as a common human activity, motivates the research on interactive head generation. A virtual agent can generate motion responses with both listening and speaking capabilities based on the audio or motion signals of the other user and itself. However, previous clip-wise ge…

2024

Cooperation Does Matter: Exploring Multi-Order Bilateral Relations for Audio-Visual Segmentation

CVPR 2024highlight

Recently an audio-visual segmentation (AVS) task has been introduced aiming to group pixels with sounding objects within a given video. This task necessitates a first-ever audio-driven pixel-level understanding of the scene posing significant challenges. In this paper we propose an innovative audio-…

2024

CustomListener: Text-guided Responsive Interaction for User-friendly Listening Head Generation

CVPR 2024poster

Listening head generation aims to synthesize a non-verbal responsive listener head by modeling the correlation between the speaker and the listener in dynamic conversion. The applications of listener agent generation in virtual interaction have promoted many works achieving diverse and fine-grained…

Cited by 4SourcePDFScholar