← Search

Chunghsin Yeh

3 accepted papers

2025

Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation

ICASSP 2025accepted

Audio token modeling has become a powerful framework for speech synthesis, with two-stage approaches employing semantic tokens remaining prevalent. In this paper, we aim to simplify this process by introducing a semantic knowledge distillation method that enables high-quality speech generation in a…

Cited by 0SourceScholar
2023

Full-Band General Audio Synthesis with Score-Based Diffusion

ICASSP 2023accepted

Recent works have shown the capability of deep generative models to tackle general audio synthesis from a single label, producing a variety of impulsive, tonal, and environmental sounds. Such models operate on band-limited signals and, as a result of an autoregressive approach, they are typically co…

Cited by 0SourceScholar