← Search

Shunshun Yin

3 accepted papers

2026

RAP: Real-time Audio-driven Portrait Animation with Video Diffusion Transformer

ICASSP 2026oral

Audio-driven portrait animation aims to synthesize realistic and natural talking head videos from an input audio signal and a single reference image. While existing methods achieve high-quality results by leveraging high-dimensional intermediate representations and explicitly modeling motion dynamic…

Cited by 0SourcePDFScholar
2026

UniSS: Unified Expressive Speech-to-Speech Translation with Your Voice

ICLR 2026poster

The ultimate goal of expressive speech-to-speech translation (S2ST) is to accurately translate spoken content while preserving the speaker identity and emotional style. However, progress in this field is largely hindered by three key challenges: the scarcity of paired speech data that retains expres…

Cited by 0SourcecodeScholar
2025

Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion Generation

CVPR 2025poster

In this work, we introduce the first autoregressive framework for real-time, audio-driven portrait animation, a.k.a, talking head. Beyond the challenge of lengthy animation times, a critical challenge in realistic talking head generation lies in preserving the natural movement of diverse body parts.…

Cited by 0SourcePDFScholar