← Search

Wonjune Kang

2 accepted papers

2024

Multi-Task Learning for Front-End Text Processing in TTS

ICASSP 2024accepted

We propose a multi-task learning (MTL) model for jointly performing three tasks that are commonly solved in a text-to-speech (TTS) front-end: text normalization (TN), part-of-speech (POS) tagging, and homograph disambiguation (HD). Our framework utilizes a tree-like structure with a trunk that learn…

Cited by 0SourceScholar
2020

Multimodal Speaker Diarization of Real-World Meetings Using D-Vectors With Spatial Features

ICASSP 2020accepted

Deep neural network based audio embeddings (d-vectors) have demonstrated superior performance in audio-only speaker diarization compared to traditional acoustic features such as mel-frequency cepstral coefficients (MFCCs) and i-vectors. However, there has been little work on multimodal diarization s…

Cited by 0SourceScholar