ICASSP 2023accepted0 citations

Synthesizing Speech from ECoG with a Combination of Transformer-Based Encoder and Neural Vocoder

Kai Shigemi, Shuji Komeiji, Takumi Mitsuhashi, Yasushi Iimura, Hiroharu Suzuki, Hidenori Sugano, Koichi Shinoda, Kohei Yatabe

Abstract

This paper reports on a novel invasive brain–computer interface (BCI) paradigm that has successfully reconstructed spoken sentences from invasive electrocorticogram (ECoG) signals using deep-neural-network-based encoders and a pre-trained neural vocoder. We recorded ECoG signals while 13 participants were speaking short sentences. Our BCI could map the ECoG recording to the log-mel spectrograms of the spoken sentences using a bidirectional long short-term memory (BLSTM) or a Transformer. The estimated log-mel spectrograms were used in Parallel WaveGAN to synthesize speech waveforms. An evaluation of the model performance revealed that the Transformer model significantly outperformed (Wilcoxon signed-rank test, p < 0.001) the BLSTM in terms of mean square error loss and Pearson correlation.

BibTeX
@inproceedings{icassp2023_synthesizingspee,
  title = {Synthesizing Speech from ECoG with a Combination of Transformer-Based Encoder and Neural Vocoder},
  author = {Kai Shigemi and Shuji Komeiji and Takumi Mitsuhashi and Yasushi Iimura and Hiroharu Suzuki and Hidenori Sugano and Koichi Shinoda and Kohei Yatabe and Toshihisa Tanaka},
  booktitle = {ICASSP 2023},
  year = {2023}
}
Synthesizing Speech from ECoG with a Combination of Transformer-Based Encoder and Neural Vocoder · ICASSP 2023