← Search

Shinnosuke Takamichi

20 accepted papers

2026

Low-Latency Real-Time Audio Game Commentary System via LLM-based Parallel Text Generation

IJCAI 2026

We present a low-latency real-time audio game commentary system that generates spoken commentary directly from live gameplay video. In this end-to-end setting, a key bottleneck is accumulated waiting time; conventional pipelines capture frames, generate text, and synthesize speech sequentially for e

Cited by 0Scholar
2024

Diversity-Based Core-Set Selection for Text-to-Speech with Linguistic and Acoustic Features

ICASSP 2024accepted

This paper proposes a method for extracting a lightweight subset from a text-to-speech (TTS) corpus ensuring synthetic speech quality. In recent years, methods have been proposed for constructing large-scale TTS corpora by collecting diverse data from massive sources such as audiobooks and YouTube.…

Cited by 0SourceScholar
2024

Do Learned Speech Symbols Follow Zipf's Law?

ICASSP 2024accepted

In this study, we investigate whether speech symbols, learned through deep learning, follow Zipf’s law, akin to natural language symbols. Zipf’s law is an empirical law that delineates the frequency distribution of words, forming fundamentals for statistical analysis in natural language processing.…

Cited by 0SourceScholar
2024

Environmental Sound Synthesis from Vocal Imitations and Sound Event Labels

ICASSP 2024accepted

One way of expressing an environmental sound is using vocal imitations, which involve the process of replicating or mimicking the rhythm and pitch of sounds by voice. We can effectively express the features of environmental sounds, such as rhythm and pitch, using vocal imitations, which cannot be ex…

Cited by 0SourceScholar
2023

Improving Speech Prosody of Audiobook Text-To-Speech Synthesis with Acoustic and Textual Contexts

ICASSP 2023accepted

We present a multi-speaker Japanese audiobook text-to-speech (TTS) system that leverages multimodal context information of preceding acoustic context and bilateral textual context to improve the prosody of synthetic speech. Previous work either uses unilateral or single-modality context, which does…

Cited by 0SourceScholar
2023

Learning to Speak from Text: Zero-Shot Multilingual Text-to-Speech with Unsupervised Text Pretraining

IJCAI 2023poster

While neural text-to-speech (TTS) has achieved human-like natural synthetic speech, multilingual TTS systems are limited to resource-rich languages due to the need for paired text and studio-quality audio data. This paper proposes a method for zero-shot multilingual TTS using text-only data for the…

2023

MID-Attribute Speaker Generation Using Optimal-Transport-Based Interpolation of Gaussian Mixture Models

ICASSP 2023accepted

In this paper, we propose a method for intermediating multiple speakers’ attributes and diversifying their voice characteristics in “speaker generation,” an emerging task that aims to synthesize a nonexistent speaker’s naturally sounding voice. The conventional TacoSpawn-based speaker generation met…

Cited by 0SourceScholar
2023

Visual Onoma-to-Wave: Environmental Sound Synthesis from Visual Onomatopoeias and Sound-Source Images

ICASSP 2023accepted

We propose a method for synthesizing environmental sounds from visually represented onomatopoeias and sound sources. An onomatopoeia is a word that imitates a sound structure, i.e., the text representation of sound. From this perspective, onoma-to-wave has been proposed to synthesize environmental s…

Cited by 0SourceScholar
2023

jaCappella Corpus: A Japanese a Cappella Vocal Ensemble Corpus

ICASSP 2023accepted

We construct a corpus of Japanese a cappella vocal ensembles (ja-Cappella corpus) for vocal ensemble separation and synthesis. It consists of 35 copyright-cleared vocal ensemble songs and their audio recordings of individual voice parts. These songs were arranged from out-of-copyright Japanese child…

Cited by 0SourceScholar
2021

Disentangled Speaker and Language Representations Using Mutual Information Minimization and Domain Adaptation for Cross-Lingual TTS

ICASSP 2021accepted

We propose a method for obtaining disentangled speaker and language representations via mutual information minimization and domain adaptation for cross-lingual text-to-speech (TTS) synthesis. The proposed method extracts speaker and language embeddings from acoustic features by a speaker encoder and…

Cited by 0SourceScholar
2021

Humanacgan: Conditional Generative Adversarial Network with Human-Based Auxiliary Classifier and its Evaluation in Phoneme Perception

ICASSP 2021accepted

We propose a conditional generative adversarial network (GAN) incorporating humans’ perceptual evaluations. A deep neural network (DNN)-based generator of a GAN can represent a real-data distribution accurately but can never represent a human-acceptable distribution, which are ranges of data in whic…

Cited by 0SourceScholar
2020

Humangan: Generative Adversarial Network With Human-Based Discriminator And Its Evaluation In Speech Perception Modeling

ICASSP 2020accepted

We propose the HumanGAN, a generative adversarial network (GAN) incorporating human perception as a discriminator. A basic GAN trains a generator to represent a real-data distribution by fooling the discriminator that distinguishes real and generated data. Therefore, the basic GAN cannot represent t…

Cited by 0SourceScholar
2020

Lifter Training and Sub-Band Modeling for Computationally Efficient and High-Quality Voice Conversion Using Spectral Differentials

ICASSP 2020accepted

In this paper, we propose computationally efficient and high-quality methods for statistical voice conversion (VC) with direct waveform modification based on spectral differentials. The conventional method with a minimum-phase filter achieves high-quality conversion but requires heavy computation in…

Cited by 0SourceScholar
2019

Generative Moment Matching Network-based Random Modulation Post-filter for DNN-based Singing Voice Synthesis and Neural Double-tracking

ICASSP 2019accepted

This paper proposes a generative moment matching network (GMMN)-based post-filter that provides inter-utterance pitch variation for deep neural network (DNN)-based singing voice synthesis. The natural pitch variation of a human singing voice leads to a richer musical experience and is used in double…

Cited by 0SourceScholar
2018

Non-Parallel Voice Conversion Using Variational Autoencoders Conditioned by Phonetic Posteriorgrams and D-Vectors

ICASSP 2018accepted

This paper proposes novel frameworks for non-parallel voice conversion (VC) using variational autoencoders (VAEs). Although conventional VAE-based VC models can be trained using non-parallel speech corpora with given speaker representations, phonetic contents of the converted speech tend to vanish b…

Cited by 0SourceScholar
2018

Text-to-Speech Synthesis Using STFT Spectra Based on Low-/Multi-Resolution Generative Adversarial Networks

ICASSP 2018accepted

This paper proposes novel training algorithms for vocoder-free statistical parametric speech synthesis (SPSS) using short-term Fourier transform (STFT) spectra. Recently, text-to-speech synthesis using STFT spectra has been investigated since it can avoid quality degradation caused by the vocoder-ba…

Cited by 0SourceScholar
2017

Blind source separation based on independent low-rank matrix analysis with sparse regularization for time-series activity

ICASSP 2017accepted

In this paper, we propose a new blind source separation (BSS) method based on independent low-rank matrix analysis (ILRMA) with novel sparse regularization. ILRMA is a recently proposed BSS algorithm that simultaneously estimates a demixing matrix and source spectrogram models based on nonnegative m…

Cited by 0SourceScholar
2017

Training algorithm to deceive Anti-Spoofing Verification for DNN-based speech synthesis

ICASSP 2017accepted

This paper proposes a novel training algorithm for high-quality Deep Neural Network (DNN)-based speech synthesis. The parameters of synthetic speech tend to be over-smoothed, and this causes significant quality degradation in synthetic speech. The proposed algorithm takes into account an Anti-Spoofi…

Cited by 0SourceScholar
2015

Modulation spectrum-constrained trajectory training algorithm for GMM-based Voice Conversion

ICASSP 2015accepted

This paper presents a novel training algorithm for Gaussian Mixture Model (GMM)-based Voice Conversion (VC). One of the advantages of GMM-based VC is computationally efficient conversion processing enabling to achieve real-time VC applications. On the other hand, the quality of the converted speech…

Cited by 0SourceScholar
2015

Parameter generation algorithm considering Modulation Spectrum for HMM-based speech synthesis

ICASSP 2015accepted

This paper proposes a novel parameter generation algorithm for high-quality speech generation in Hidden Markov Model (HMM)-based speech synthesis. One of the biggest issues causing significant quality degradation is the over-smoothing effect often observed in generated parameter trajectories. Global…

Cited by 0SourceScholar