← Search

Dongchan Min

6 accepted papers

2025

FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait

ICCV 2025poster

With the rapid advancement of diffusion-based generative models, portrait image animation has achieved remarkable results. However, it still faces challenges in temporally consistent video generation and fast sampling due to its iterative sampling nature. This paper presents FLOAT, an audio-driven t…

2024

Learning to Generate Conditional Tri-plane for 3D-aware Expression Controllable Portrait Animation

ECCV 2024poster

"In this paper, we present , a one-shot 3D-aware portrait animation method that is able to control the facial expression and camera view of a given portrait image. To achieve this, we introduce a tri-plane generator with an effective expression conditioning method, which directly generates a tri-pla…

Cited by 4SourcePDFScholar
2023

Grad-StyleSpeech: Any-Speaker Adaptive Text-to-Speech Synthesis with Diffusion Models

ICASSP 2023accepted

There has been a significant progress in Text-To-Speech (TTS) synthesis technology in recent years, thanks to the advancement in neural generative modeling. However, existing methods on any-speaker adaptive TTS have achieved unsatisfactory performance, due to their suboptimal accuracy in mimicking t…

Cited by 0SourceScholar
2021

Meta-GMVAE: Mixture of Gaussian VAE for Unsupervised Meta-Learning

ICLR 2021spotlight

Unsupervised learning aims to learn meaningful representations from unlabeled data which can captures its intrinsic structure, that can be transferred to downstream tasks. Meta-learning, whose objective is to learn to generalize across tasks such that the learned model can rapidly adapt to a novel t…

2021

Meta-StyleSpeech : Multi-Speaker Adaptive Text-to-Speech Generation

ICML 2021spotlight

With rapid progress in neural text-to-speech (TTS) models, personalized speech generation is now in high demand for many applications. For practical applicability, a TTS model should generate high-quality speech with only a few audio samples from the given speaker, that are also short in length. How…