← Search

Changjin Han

3 accepted papers

2025

Improving Robustness of Diffusion-Based Zero-Shot Speech Synthesis via Stable Formant Generation

ICASSP 2025accepted

Diffusion models have achieved remarkable success in text-to-speech (TTS), even in zero-shot scenarios. Recent efforts aim to address the trade-off between inference speed and sound quality, often considered the primary drawback of diffusion models. However, we find a critical mispronunciation issue…

Cited by 0SourceScholar
2023

Good Neighbors are All You Need for Chinese Grapheme-To-Phoneme Conversion

ICASSP 2023accepted

Most Chinese Grapheme-to-Phoneme (G2P) systems employ a three-stage framework that first transforms input sequences into character embeddings, obtains linguistic information using language models, and then predicts the phonemes based on global context about the entire input sequence. However, lingui…

Cited by 0SourceScholar
2021

Image-to-Image Retrieval by Learning Similarity between Scene Graphs

AAAI 2021technical

As a scene graph compactly summarizes the high-level content of an image in a structured and symbolic manner, the similarity between scene graphs of two images reflects the relevance of their contents. Based on this idea, we propose a novel approach for image-to-image retrieval using scene graph sim…