← Search

Hiroki Kanagawa

3 accepted papers

2024

Noise-Robust Zero-Shot Text-to-Speech Synthesis Conditioned on Self-Supervised Speech-Representation Model with Adapters

ICASSP 2024accepted

The zero-shot text-to-speech (TTS) method, based on speaker embeddings extracted from reference speech using self-supervised learning (SSL) speech representations, can reproduce speaker characteristics very accurately. However, this approach suffers from degradation in speech synthesis quality when…

Cited by 0SourceScholar
2023

Enhancement of Text-Predicting Style Token With Generative Adversarial Network for Expressive Speech Synthesis

ICASSP 2023accepted

This work proposes an advanced text-predicting style embedding for expressive speech synthesis. Text-predicting global style token (TPGST) predicts style embedding from text instead of reference speech and uses it to condition a text-to-speech synthesis (TTS) model, resulting in style TTS without re…

Cited by 0SourceScholar