← Search

Seungjun Chung

2 accepted papers

2025

DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

ICLR 2025poster

Large-scale latent diffusion models (LDMs) excel in content generation across various modalities, but their reliance on phonemes and durations in text-to-speech (TTS) limits scalability and access from other fields. While recent studies show potential in removing these domain-specific factors, perfo…

2024

CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech

ICLR 2024poster

With the emergence of neural audio codecs, which encode multiple streams of discrete tokens from audio, large language models have recently gained attention as a promising approach for zero-shot Text-to-Speech (TTS) synthesis. Despite the ongoing rush towards scaling paradigms, audio tokenization ir…

Cited by 38SourcePDFScholar