← Search

Gyeongsu Chae

5 accepted papers

2025

FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait

ICCV 2025poster

With the rapid advancement of diffusion-based generative models, portrait image animation has achieved remarkable results. However, it still faces challenges in temporally consistent video generation and fast sampling due to its iterative sampling nature. This paper presents FLOAT, an audio-driven t…

2025

Improving Robustness of Diffusion-Based Zero-Shot Speech Synthesis via Stable Formant Generation

ICASSP 2025accepted

Diffusion models have achieved remarkable success in text-to-speech (TTS), even in zero-shot scenarios. Recent efforts aim to address the trade-off between inference speed and sound quality, often considered the primary drawback of diffusion models. However, we find a critical mispronunciation issue…

Cited by 0SourceScholar
2023

DisCoHead: Audio-and-Video-Driven Talking Head Generation by Disentangled Control of Head Pose and Facial Expressions

ICASSP 2023accepted

For realistic talking head generation, creating natural head motion while maintaining accurate lip synchronization is essential. To fulfill this challenging task, we propose DisCoHead, a novel method to disentangle and control head pose and facial expressions without supervision. DisCoHead uses a si…

Cited by 0SourceScholar
2023

Good Neighbors are All You Need for Chinese Grapheme-To-Phoneme Conversion

ICASSP 2023accepted

Most Chinese Grapheme-to-Phoneme (G2P) systems employ a three-stage framework that first transforms input sequences into character embeddings, obtains linguistic information using language models, and then predicts the phonemes based on global context about the entire input sequence. However, lingui…

Cited by 0SourceScholar
2021

KoDF: A Large-Scale Korean DeepFake Detection Dataset

ICCV 2021poster

A variety of effective face-swap and face-reenactment methods have been publicized in recent years, democratizing the face synthesis technology to a great extent. Videos generated as such have come to be called deepfakes with a negative connotation, for various social problems they have caused. Faci…

Cited by 147PDFcodeScholar