← Search

Junseok Ahn

3 accepted papers

2025

VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis

ICASSP 2025accepted

We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for intelligible speech, achieving this alignment in noisy conditions remains a significant and underexplored challenge in the…

Cited by 0SourceScholar
2024

Faces that Speak: Jointly Synthesising Talking Face and Speech from Text

CVPR 2024poster

The goal of this work is to simultaneously generate natural talking faces and speech outputs from text. We achieve this by integrating Talking Face Generation (TFG) and Text-to-Speech (TTS) systems into a unified framework. We address the main challenges of each task: (1) generating a range of head…

Cited by 11SourcePDFScholar