← Search

Junichi Yamagishi

51 accepted papers

2026

Training Dynamics-Aware Multi-Factor Curriculum Learning for Target Speaker Extraction

ICASSP 2026oral

Target speaker extraction (TSE) aims to isolate a specific speaker's voice from multi-speaker mixtures. Despite strong benchmark results, real-world performance often degrades due to different interacting factors. Previous curriculum learning approaches for TSE typically address these factors separa…

Cited by 0SourcePDFScholar
2025

QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions

ACL 2025long

This paper explores a novel perspective to speech quality assessment by leveraging natural language descriptions, offering richer, more nuanced insights than traditional numerical scoring methods. Natural language feedback provides instructive recommendations and detailed evaluations, yet existing d…

2025

Towards An Integrated Approach for Expressive Piano Performance Synthesis from Music Scores

ICASSP 2025accepted

This paper presents an integrated system that transforms symbolic music scores into expressive piano performance audio. By combining a Transformer-based Expressive Performance Rendering (EPR) model with a fine-tuned neural MIDI synthesiser, our approach directly generates expressive audio performanc…

Cited by 0SourceScholar
2024

Bridging Textual and Tabular Worlds for Fact Verification: A Lightweight, Attention-Based Model

COLING 2024main

FEVEROUS is a benchmark and research initiative focused on fact extraction and verification tasks involving unstructured text and structured tabular data. In FEVEROUS, existing works often rely on extensive preprocessing and utilize rule-based transformations of data, leading to potential context lo…

2024

Can Large-Scale Vocoded Spoofed Data Improve Speech Spoofing Countermeasure with a Self-Supervised Front End?

ICASSP 2024accepted

A speech spoofing countermeasure (CM) that discriminates between unseen spoofed and bona fide data requires diverse training data. While many datasets use spoofed data generated by speech synthesis systems, it was recently found that data vocoded by neural vocoders were also effective as the spoofed…

Cited by 0SourceScholar
2024

Spoofing Attack Augmentation: Can Differently-Trained Attack Models Improve Generalisation?

ICASSP 2024accepted

A reliable deepfake detector or spoofing countermeasure (CM) should be robust in the face of unpredictable spoofing attacks. To encourage the learning of more generaliseable artefacts, rather than those specific only to known attacks, CMs are usually exposed to a broad variety of different attacks d…

Cited by 0SourceScholar
2024

Synvox2: Towards A Privacy-Friendly Voxceleb2 Dataset

ICASSP 2024accepted

The success of deep learning in speaker recognition relies heavily on the use of large datasets. However, the data-hungry nature of deep learning methods has already being questioned on account the ethical, privacy, and legal concerns that arise when using large-scale datasets of natural speech coll…

Cited by 0SourceScholar
2023

Can Knowledge of End-to-End Text-to-Speech Models Improve Neural Midi-to-Audio Synthesis Systems?

ICASSP 2023accepted

With the similarity between music and speech synthesis from symbolic input and the rapid development of text-to-speech (TTS) techniques, it is worthwhile to explore ways to improve the MIDI-to-audio performance by borrowing from TTS techniques. In this study, we analyze the shortcomings of a TTS-bas…

Cited by 0SourceScholar
2023

Hiding Speaker's Sex in Speech Using Zero-Evidence Speaker Representation in an Analysis/Synthesis Pipeline

ICASSP 2023accepted

The use of modern vocoders in an analysis/synthesis pipeline allows us to investigate high-quality voice conversion that can be used for privacy purposes. Here, we propose to transform the speaker embedding and the pitch in order to hide the sex of the speaker. ECAPA-TDNN-based speaker representatio…

Cited by 0SourceScholar
2023

Joint Noise Reduction and Listening Enhancement for Full-End Speech Enhancement

ICASSP 2023accepted

Speech enhancement (SE) methods mainly focus on recovering clean speech from noisy input. In real-world speech communication, however, noises often exist in not only speaker but also listener environments. Although SE methods can suppress the noise contained in the speaker’s voice, they cannot deal…

Cited by 0SourceScholar
2023

Spoofed Training Data for Speech Spoofing Countermeasure Can Be Efficiently Created Using Neural Vocoders

ICASSP 2023accepted

A good training set for speech spoofing countermeasures requires diverse TTS and VC spoofing attacks, but generating TTS and VC spoofed trials for a target speaker may be technically demanding. Instead of using full-fledged TTS and VC systems, this study uses neural-network-based vocoders to do copy…

Cited by 0SourceScholar
2022

Attention Back-End for Automatic Speaker Verification with Multiple Enrollment Utterances

ICASSP 2022accepted

Probabilistic linear discriminant analysis (PLDA) or cosine similarity have been widely used in traditional speaker verification systems as back-end techniques to measure pairwise similarities. To make better use of multiple enrollment utterances, we propose a novel attention back-end model that can…

Cited by 0SourceScholar
2022

LDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech

ICASSP 2022accepted

An effective approach to automatically predict the subjective rating for synthetic speech is to train on a listening test dataset with human-annotated scores. Although each speech sample in the dataset is rated by several listeners, most previous works only used the mean score as the training target…

Cited by 0SourceScholar
2022

On the Interplay between Sparsity, Naturalness, Intelligibility, and Prosody in Speech Synthesis

ICASSP 2022accepted

Are end-to-end text-to-speech (TTS) models over-parametrized? To what extent can these models be pruned, and what happens to their synthesis capabilities? This work serves as a starting point to explore pruning both spectrogram prediction networks and vocoders. We thoroughly investigate the tradeoff…

Cited by 0SourceScholar
2021

How Similar or Different is Rakugo Speech Synthesizer to Professional Performers?

ICASSP 2021accepted

We have been working on speech synthesis for rakugo (a traditional Japanese form of verbal entertainment similar to one-person stand-up comedy) toward speech synthesis that authentically entertains audiences. In this paper, we propose a novel evaluation methodology using synthesized rakugo speech an…

Cited by 0SourceScholar
2021

Learning Disentangled Phone and Speaker Representations in a Semi-Supervised VQ-VAE Paradigm

ICASSP 2021accepted

We present a new approach to disentangle speaker voice and phone content by introducing new components to the VQ-VAE architecture for speech synthesis. The original VQ-VAE does not generalize well to unseen speakers or content. To alleviate this problem, we have incorporated a speaker encoder and sp…

Cited by 0SourceScholar
2021

OpenForensics: Large-Scale Challenging Dataset for Multi-Face Forgery Detection and Segmentation In-the-Wild

ICCV 2021poster

The proliferation of deepfake media is raising concerns among the public and relevant authorities. It has become essential to develop countermeasures against forged faces in social media. This paper presents a comprehensive study on two new countermeasure tasks: multi-face forgery detection and segm…

Cited by 99PDFScholar
2020

Effect of Choice of Probability Distribution, Randomness, and Search Methods for Alignment Modeling in Sequence-to-Sequence Text-to-Speech Synthesis Using Hard Alignment

ICASSP 2020accepted

Sequence-to-sequence text-to-speech (TTS) is dominated by soft-attention-based methods. Recently, hard-attention-based methods have been proposed to prevent fatal alignment errors, but their sampling method of discrete alignment is poorly investigated. This research investigates various combinations…

Cited by 0SourceScholar
2020

Transferring Neural Speech Waveform Synthesizers to Musical Instrument Sounds Generation

ICASSP 2020accepted

Recent neural waveform synthesizers such as WaveNet, WaveG-low, and the neural-source-filter (NSF) model have shown good performance in speech synthesis despite their different methods of waveform generation. The similarity between speech and music audio synthesis techniques suggests interesting ave…

Cited by 0SourceScholar
2020

Zero-Shot Multi-Speaker Text-To-Speech with State-Of-The-Art Neural Speaker Embeddings

ICASSP 2020accepted

While speaker adaptation for end-to-end speech synthesis using speaker embeddings can produce good speaker similarity for speakers seen during training, there remains a gap for zero-shot adaptation to unseen speakers. We investigate multi-speaker modeling for end-to-end text-to-speech synthesis and…

Cited by 0SourceScholar
2019

Attentive Filtering Networks for Audio Replay Attack Detection

ICASSP 2019accepted

An attacker may use a variety of techniques to fool an automatic speaker verification system into accepting them as a genuine user. Anti-spoofing methods meanwhile aim to make the system robust against such attacks. The ASVspoof 2017 Challenge focused specifically on replay attacks, with the intenti…

Cited by 0SourceScholar
2019

Audiovisual Speaker Conversion: Jointly and Simultaneously Transforming Facial Expression and Acoustic Characteristics

ICASSP 2019accepted

An audiovisual speaker conversion method is presented for simultaneously transforming the facial expressions and voice of a source speaker into those of a target speaker. Transforming the facial and acoustic features together makes it possible for the converted voice and facial expressions to be hig…

Cited by 0SourceScholar
2019

Capsule-forensics: Using Capsule Networks to Detect Forged Images and Videos

ICASSP 2019accepted

Recent advances in media generation techniques have made it easier for attackers to create forged images and videos. State-of-the-art methods enable the real-time creation of a forged version of a single video obtained from a social network. Although numerous methods have been developed for detectin…

Cited by 0SourceScholar
2019

Cycle-consistent Adversarial Networks for Non-parallel Vocal Effort Based Speaking Style Conversion

ICASSP 2019accepted

Speaking style conversion (SSC) is the technology of converting natural speech signals from one style to another. In this study, we propose the use of cycle-consistent adversarial networks (CycleGANs) for converting styles with varying vocal effort, and focus on conversion between normal and Lombard…

Cited by 0SourceScholar
2019

Investigation of Enhanced Tacotron Text-to-speech Synthesis Systems with Self-attention for Pitch Accent Language

ICASSP 2019accepted

End-to-end speech synthesis is a promising approach that directly converts raw text to speech. Although it was shown that Tacotron2 outperforms classical pipeline systems with regards to naturalness in English, its applicability to other languages is still unknown. Japanese could be one of the most…

Cited by 0SourceScholar
2019

Neural Source-filter-based Waveform Model for Statistical Parametric Speech Synthesis

ICASSP 2019accepted

Neural waveform models such as the WaveNet are used in many recent text-to-speech systems, but the original WaveNet is quite slow in waveform generation because of its autoregressive (AR) structure. Although faster non-AR models were recently reported, they may be prohibitively complicated due to th…

Cited by 0SourceScholar
2019

STFT Spectral Loss for Training a Neural Speech Waveform Model

ICASSP 2019accepted

This paper proposes a new loss using short-time Fourier transform (STFT) spectra for the aim of training a high-performance neural speech waveform model that predicts raw continuous speech waveform samples directly. Not only amplitude spectra but also phase spectra obtained from generated speech wav…

Cited by 0SourceScholar
2019

Waveform Generation for Text-to-speech Synthesis Using Pitch-synchronous Multi-scale Generative Adversarial Networks

ICASSP 2019accepted

The state-of-the-art in text-to-speech (TTS) synthesis has recently improved considerably due to novel neural waveform generation methods, such as WaveNet. However, these methods suffer from their slow sequential inference process, while their parallel versions are difficult to train and even more c…

Cited by 0SourceScholar
2018

A Comparison of Recent Waveform Generation and Acoustic Modeling Methods for Neural-Network-Based Speech Synthesis

ICASSP 2018accepted

Recent advances in speech synthesis suggest that limitations such as the lossy nature of the amplitude spectrum with minimum phase approximation and the over-smoothing effect in acoustic modeling can be overcome by using advanced machine learning approaches. In this paper, we build a framework in wh…

Cited by 0SourceScholar
2018

Cyborg Speech: Deep Multilingual Speech Synthesis for Generating Segmental Foreign Accent with Natural Prosody

ICASSP 2018accepted

We describe a new application of deep-learning-based speech synthesis, namely multilingual speech synthesis for generating controllable foreign accent. Specifically, we train a DBLSTM-based acoustic model on non-accented multilingual speech recordings from a speaker native in several languages. By c…

Cited by 0SourceScholar
2018

High-Quality Nonparallel Voice Conversion Based on Cycle-Consistent Adversarial Network

ICASSP 2018accepted

Although voice conversion (VC) algorithms have achieved remarkable success along with the development of machine learning, superior performance is still difficult to achieve when using nonparallel data. In this paper, we propose using a cycle-consistent adversarial network (CycleGAN) for nonparallel…

Cited by 0SourceScholar
2018

Speech Waveform Synthesis from MFCC Sequences with Generative Adversarial Networks

ICASSP 2018accepted

This paper proposes a method for generating speech from filterbank mel frequency cepstral coefficients (MFCC), which are widely used in speech applications, such as ASR, but are generally considered unusable for speech synthesis. First, we predict fundamental frequency and voicing information from M…

Cited by 0SourceScholar
2017

Adapting and controlling DNN-based speech synthesis using input codes

ICASSP 2017accepted

Methods for adapting and controlling the characteristics of output speech are important topics in speech synthesis. In this work, we investigated the performance of DNN-based text-to-speech systems that in parallel to conventional text input also take speaker, gender, and age codes as inputs, in ord…

Cited by 0SourceScholar
2017

An autoregressive recurrent mixture density network for parametric speech synthesis

ICASSP 2017accepted

Neural-network-based generative models, such as mixture density networks, are potential solutions for speech synthesis. In this paper we follow this path and propose a recurrent mixture density network that incorporates a trainable autoregressive model. An advantage of incorporating an autoregressiv…

Cited by 0SourceScholar
2017

Non-parallel voice conversion using i-vector PLDA: towards unifying speaker verification and transformation

ICASSP 2017accepted

Text-independent speaker verification (recognizing speakers regardless of content) and non-parallel voice conversion (transforming voice identities without requiring content-matched training utterances) are related problems. We adopt i-vector method to voice conversion. An i-vector is a fixed-dimens…

Cited by 0SourceScholar
2016

A deep auto-encoder based low-dimensional feature extraction from FFT spectral envelopes for statistical parametric speech synthesis

ICASSP 2016accepted

In the state-of-the-art statistical parametric speech synthesis system, a speech analysis module, e.g. STRAIGHT spectral analysis, is generally used for obtaining accurate and stable spectral envelopes, and then low-dimensional acoustic features extracted from obtained spectral envelopes are used fo…

Cited by 0SourceScholar
2016

Deep neural network-guided unit selection synthesis

ICASSP 2016accepted

Vocoding of speech is a standard part of statistical parametric speech synthesis systems. It imposes an upper bound of the naturalness that can possibly be achieved. Hybrid systems using parametric models to guide the selection of natural speech units can combine the benefits of robust statistical m…

Cited by 0SourceScholar
2016

Initial investigation of speech synthesis based on complex-valued neural networks

ICASSP 2016accepted

Although frequency analysis often leads us to a speech signal in the complex domain, the acoustic models we frequently use are designed for real-valued data. Phase is usually ignored or modelled separately from spectral amplitude. Here, we propose a complex-valued neural network (CVNN) for directly…

Cited by 0SourceScholar
2016

Privacy-preserving sound to degrade automatic speaker verification performance

ICASSP 2016accepted

In this paper, a privacy protection method to prevent speaker identification from recorded speech is proposed and evaluated. Although many techniques for preserving various private information included in speech have been proposed, their impacts on human speech communication in physical space are no…

Cited by 0SourceScholar
2016

Testing the consistency assumption: Pronunciation variant forced alignment in read and spontaneous speech synthesis

ICASSP 2016accepted

Forced alignment for speech synthesis traditionally aligns a phoneme sequence predetermined by the front-end text processing system. This sequence is not altered during alignment, i.e., it is forced, despite possibly being faulty. The consistency assumption is the assumption that these mistakes do n…

Cited by 14SourceScholar
2016

Wavelet-based decomposition of F0 as a secondary task for DNN-based speech synthesis with multi-task learning

ICASSP 2016accepted

We investigate two wavelet-based decomposition strategies of the f0 signal and their usefulness as a secondary task for speech synthesis using multi-task deep neural networks (MTL-DNN). The first decomposition strategy uses a static set of scales for all utterances in the training data. We propose a…

Cited by 0SourceScholar
2015

Methods for applying dynamic sinusoidal models to statistical parametric speech synthesis

ICASSP 2015accepted

Sinusoidal vocoders can generate high quality speech, but they have not been extensively applied to statistical parametric speech synthesis. This paper presents two ways for using dynamic sinusoidal models for statistical speech synthesis, enabling the sinusoid parameters to be modelled in HMM-based…

Cited by 0SourceScholar
2015

SAS: A speaker verification spoofing database containing diverse attacks

ICASSP 2015accepted

This paper presents the first version of a speaker verification spoofing and anti-spoofing database, named SAS corpus. The corpus includes nine spoofing techniques, two of which are speech synthesis, and seven are voice conversion. We design two protocols, one for standard speaker verification evalu…

Cited by 0SourceScholar