← Search

Gaojie Lin

6 accepted papers

2026

InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions

ICLR 2026poster

End-to-end human animation with rich multi-modal conditions, e.g., text, image and audio has achieved remarkable advancements in recent years. However, most existing methods could only animate a single subject and inject conditions in a global manner, ignoring scenarios that multiple concepts could…

Cited by 0SourceScholar
2025

CyberHost: A One-stage Diffusion Framework for Audio-driven Talking Body Generation

ICLR 2025oral

Diffusion-based video generation technology has advanced significantly, catalyzing a proliferation of research in human animation. While breakthroughs have been made in driving human animation through various modalities for portraits, most of current solutions for human body animation still focus on…

Cited by 0SourcePDFScholar
2025

FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG Distillation

CVPR 2025poster

Diffusion-based audio-driven talking avatar methods have recently gained attention for their high-fidelity, vivid, and expressive results. However, their slow inference speed limits practical applications. Despite the development of various distillation techniques for diffusion models, we found that…

2025

Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

ICLR 2025oral

With the introduction of video diffusion model, audio-conditioned human video generation has recently achieved significant breakthroughs in both the naturalness of motion and the synthesis of portrait details. Due to the limited control of audio signals in driving human motion, existing methods ofte…

2025

MobilePortrait: Real-Time One-Shot Neural Head Avatars on Mobile Devices

CVPR 2025poster

Existing neural head avatars methods have achieved significant progress in the image quality and motion range of portrait animation. However, these methods prioritize effectiveness over computational overhead. This paper presents MobilePortrait, a lightweight one-shot neural head avatars method that…

Cited by 8SourcePDFScholar
2025

OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

ICCV 2025poster

End-to-end human animation, such as audio-driven talking human generation, has undergone notable advancements in the recent few years. However, existing methods still struggle to scale up as large general video generation models, limiting their potential in real applications. In this paper, we propo…

Cited by 0SourcePDFScholar