← Search

Hiroshi Saruwatari

45 accepted papers

2026

DISSECTING PERFORMANCE DEGRADATION IN AUDIO SOURCE SEPARATION UNDER SAMPLING FREQUENCY MISMATCH

ICASSP 2026poster

Audio processing methods based on deep neural networks are typically trained at a single sampling frequency (SF). To handle untrained SFs, signal resampling is commonly employed, but it can degrade performance, particularly when the input SF is lower than the trained SF. This paper investigates the…

Cited by 0SourcePDFScholar
2025

Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features

ICASSP 2025accepted

Real-time speech enhancement (SE) is essential to online speech communication. Causal SE models use only the previous context while predicting future information, such as phoneme continuation, may help performing causal SE. The phonetic information is often represented by quantizing latent features…

Cited by 0SourceScholar
2025

Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis

NeurIPS 2025poster

We propose Shallow Flow Matching (SFM), a novel mechanism that enhances flow matching (FM)-based text-to-speech (TTS) models within a coarse-to-fine generation paradigm. Unlike conventional FM modules, which use the coarse representations from the weak generator as conditions, SFM constructs interme…

Cited by 0SourcecodeScholar
2024

Diversity-Based Core-Set Selection for Text-to-Speech with Linguistic and Acoustic Features

ICASSP 2024accepted

This paper proposes a method for extracting a lightweight subset from a text-to-speech (TTS) corpus ensuring synthetic speech quality. In recent years, methods have been proposed for constructing large-scale TTS corpora by collecting diverse data from massive sources such as audiobooks and YouTube.…

Cited by 0SourceScholar
2024

Do Learned Speech Symbols Follow Zipf's Law?

ICASSP 2024accepted

In this study, we investigate whether speech symbols, learned through deep learning, follow Zipf’s law, akin to natural language symbols. Zipf’s law is an empirical law that delineates the frequency distribution of words, forming fundamentals for statistical analysis in natural language processing.…

Cited by 0SourceScholar
2024

Localizing Acoustic Energy in Sound Field Synthesis by Directionally Weighted Exterior Radiation Suppression

ICASSP 2024accepted

A method for synthesizing the desired sound field while suppressing the exterior radiation power with directional weighting is proposed. The exterior radiation from the loudspeakers in sound field synthesis systems can be problematic in practical situations. Although several methods to suppress the…

Cited by 0SourceScholar
2023

Duration-Aware Pause Insertion Using Pre-Trained Language Model for Multi-Speaker Text-To-Speech

ICASSP 2023accepted

Pause insertion, also known as phrase break prediction and phrasing, is an essential part of TTS systems because proper pauses with natural duration significantly enhance the rhythm and intelligibility of synthetic speech. However, conventional phrasing models ignore various speakers’ different styl…

Cited by 0SourceScholar
2023

Improving Speech Prosody of Audiobook Text-To-Speech Synthesis with Acoustic and Textual Contexts

ICASSP 2023accepted

We present a multi-speaker Japanese audiobook text-to-speech (TTS) system that leverages multimodal context information of preceding acoustic context and bilateral textual context to improve the prosody of synthetic speech. Previous work either uses unilateral or single-modality context, which does…

Cited by 0SourceScholar
2023

Kernel Interpolation of Acoustic Transfer Functions with Adaptive Kernel for Directed and Residual Reverberations

ICASSP 2023accepted

An interpolation method for region-to-region acoustic transfer functions (ATFs) based on kernel ridge regression with an adaptive kernel is proposed. Most current ATF interpolation methods do not incorporate the acoustic properties for which measurements are performed. Our proposed method is based o…

Cited by 0SourceScholar
2023

Learning to Speak from Text: Zero-Shot Multilingual Text-to-Speech with Unsupervised Text Pretraining

IJCAI 2023poster

While neural text-to-speech (TTS) has achieved human-like natural synthetic speech, multilingual TTS systems are limited to resource-rich languages due to the need for paired text and studio-quality audio data. This paper proposes a method for zero-shot multilingual TTS using text-only data for the…

2023

MID-Attribute Speaker Generation Using Optimal-Transport-Based Interpolation of Gaussian Mixture Models

ICASSP 2023accepted

In this paper, we propose a method for intermediating multiple speakers’ attributes and diversifying their voice characteristics in “speaker generation,” an emerging task that aims to synthesize a nonexistent speaker’s naturally sounding voice. The conventional TacoSpawn-based speaker generation met…

Cited by 0SourceScholar
2023

Spatial Active Noise Control Method Based on Sound Field Interpolation from Reference Microphone Signals

ICASSP 2023accepted

A spatial active noise control (ANC) method based on the interpolation of a sound field from reference microphone signals is proposed. In most current spatial ANC methods, a sufficient number of error microphones are required to reduce noise over the target region because the sound field is estimate…

Cited by 0SourceScholar
2023

Visual Onoma-to-Wave: Environmental Sound Synthesis from Visual Onomatopoeias and Sound-Source Images

ICASSP 2023accepted

We propose a method for synthesizing environmental sounds from visually represented onomatopoeias and sound sources. An onomatopoeia is a word that imitates a sound structure, i.e., the text representation of sound. From this perspective, onoma-to-wave has been proposed to synthesize environmental s…

Cited by 0SourceScholar
2023

jaCappella Corpus: A Japanese a Cappella Vocal Ensemble Corpus

ICASSP 2023accepted

We construct a corpus of Japanese a cappella vocal ensembles (ja-Cappella corpus) for vocal ensemble separation and synthesis. It consists of 35 copyright-cleared vocal ensemble songs and their audio recordings of individual voice parts. These songs were arranged from out-of-copyright Japanese child…

Cited by 0SourceScholar
2022

Differentiable Digital Signal Processing Mixture Model for Synthesis Parameter Extraction from Mixture of Harmonic Sounds

ICASSP 2022accepted

A differentiable digital signal processing (DDSP) autoencoder is a musical sound synthesizer that combines a deep neural network (DNN) and spectral modeling synthesis. It allows us to flexibly edit sounds by changing the fundamental frequency, timbre feature, and loudness (synthesis parameters) extr…

Cited by 0SourceScholar
2022

Region-to-Region Kernel Interpolation of Acoustic Transfer Function with Directional Weighting

ICASSP 2022accepted

A method of interpolating the acoustic transfer function (ATF) between regions that takes into account both the physical properties of the ATF and the directionality of region configurations is proposed. Most spatial ATF interpolation methods are limited to estimation in the region of receivers. A k…

Cited by 4SourceScholar
2022

Spatial Active Noise Control Based on Individual Kernel Interpolation of Primary and Secondary Sound Fields

ICASSP 2022accepted

A spatial active noise control (ANC) method based on the individual kernel interpolation of primary and secondary sound fields is proposed. Spatial ANC is aimed at cancelling unwanted primary noise within a continuous region by using multiple secondary sources and microphones. A method based on the…

Cited by 0SourceScholar
2021

Amplitude Matching: Majorization-Minimization Algorithm for Sound Field Control Only with Amplitude Constraint

ICASSP 2021accepted

A sound field control method for synthesizing a desired amplitude distribution inside a target region, amplitude matching, is proposed. In the conventional pressure matching, a desired sound field is set as a pressure distribution including amplitude and phase. In personal audio applications, it is…

Cited by 0SourceScholar
2021

Deficient Basis Estimation of Noise Spatial Covariance Matrix for Rank-Constrained Spatial Covariance Matrix Estimation Method in Blind Speech Extraction

ICASSP 2021accepted

Rank-constrained spatial covariance matrix estimation (RCSCME) is a state-of-the-art blind speech extraction method applied to cases where one directional target speech and diffuse noise are mixed. In this paper, we proposed a new algorithmic extension of RCSCME. RCSCME complements a deficient one r…

Cited by 0SourceScholar
2021

Disentangled Speaker and Language Representations Using Mutual Information Minimization and Domain Adaptation for Cross-Lingual TTS

ICASSP 2021accepted

We propose a method for obtaining disentangled speaker and language representations via mutual information minimization and domain adaptation for cross-lingual text-to-speech (TTS) synthesis. The proposed method extracts speaker and language embeddings from acoustic features by a speaker encoder and…

Cited by 0SourceScholar
2021

Humanacgan: Conditional Generative Adversarial Network with Human-Based Auxiliary Classifier and its Evaluation in Phoneme Perception

ICASSP 2021accepted

We propose a conditional generative adversarial network (GAN) incorporating humans’ perceptual evaluations. A deep neural network (DNN)-based generator of a GAN can represent a real-data distribution accurately but can never represent a human-acceptable distribution, which are ranges of data in whic…

Cited by 0SourceScholar
2020

Convergence-Guaranteed Independent Positive Semidefinite Tensor Analysis Based on Student's T Distribution

ICASSP 2020accepted

In this paper, we address a blind source separation (BSS) problem and propose a new extended framework of independent positive semidefinite tensor analysis (IPSDTA). IPSDTA is a state-of-the-art BSS method that enables us to take interfrequency correlations into account, but the generative model is…

Cited by 0SourceScholar
2020

Humangan: Generative Adversarial Network With Human-Based Discriminator And Its Evaluation In Speech Perception Modeling

ICASSP 2020accepted

We propose the HumanGAN, a generative adversarial network (GAN) incorporating human perception as a discriminator. A basic GAN trains a generator to represent a real-data distribution by fooling the discriminator that distinguishes real and generated data. Therefore, the basic GAN cannot represent t…

Cited by 0SourceScholar
2020

Lifter Training and Sub-Band Modeling for Computationally Efficient and High-Quality Voice Conversion Using Spectral Differentials

ICASSP 2020accepted

In this paper, we propose computationally efficient and high-quality methods for statistical voice conversion (VC) with direct waveform modification based on spectral differentials. The conventional method with a minimum-phase filter achieves high-quality conversion but requires heavy computation in…

Cited by 0SourceScholar
2020

Mutual-Information-Based Sensor Placement for Spatial Sound Field Recording

ICASSP 2020accepted

A sensor (microphone) placement method based on mutual information for spatial sound field recording is proposed. The sound field recording methods using distributed sensors enable the estimation of the sound field inside a target region of arbitrary shape; however, it is a difficult task to find th…

Cited by 0SourceScholar
2020

Regularized Fast Multichannel Nonnegative Matrix Factorization with ILRMA-Based Prior Distribution of Joint-Diagonalization Process

ICASSP 2020accepted

In this paper, we address a convolutive blind source separation (BSS) problem and propose a new extended framework of FastMNMF by introducing prior information for joint diagonalization of the spatial covariance matrix model. Recently, FastMNMF has been proposed as a fast version of multichannel non…

Cited by 0SourceScholar
2020

Spatial Active Noise Control Based on Kernel Interpolation with Directional Weighting

ICASSP 2020accepted

A spatial active noise control (ANC) method taking prior information on the approximate direction of primary noise sources into consideration is proposed. ANC aims to cancel incoming primary noise using secondary loudspeakers. Conventional multipoint ANC does not guarantee the reduction of noise bet…

Cited by 0SourceScholar
2020

Time-Domain Audio Source Separation Based on Wave-U-Net Combined with Discrete Wavelet Transform

ICASSP 2020accepted

We propose a time-domain audio source separation method using down-sampling (DS) and up-sampling (US) layers based on a discrete wavelet transform (DWT). The proposed method is based on one of the state-of-the-art deep neural networks, Wave-U-Net, which successively down-samples and up-samples featu…

Cited by 0SourceScholar
2020

Utterance-Level Sequential Modeling for Deep Gaussian Process Based Speech Synthesis Using Simple Recurrent Unit

ICASSP 2020accepted

This paper presents a deep Gaussian process (DGP) model with a recurrent architecture for speech sequence modeling. DGP is a Bayesian deep model that can be trained effectively with the consideration of model complexity and is a kernel regression model that can have high expressibility. In the previ…

Cited by 0SourceScholar
2019

Feedforward Spatial Active Noise Control Based on Kernel Interpolation of Sound Field

ICASSP 2019accepted

A method for feedforward active noise control (ANC) over a spatial region is proposed. Conventional multipoint ANC aims to reduce the noise at multiple discrete positions; therefore, the noise reduction in the region between these points cannot be guaranteed. Recent studies revealed the possibility…

Cited by 0SourceScholar
2019

Generative Moment Matching Network-based Random Modulation Post-filter for DNN-based Singing Voice Synthesis and Neural Double-tracking

ICASSP 2019accepted

This paper proposes a generative moment matching network (GMMN)-based post-filter that provides inter-utterance pitch variation for deep neural network (DNN)-based singing voice synthesis. The natural pitch variation of a human singing voice leads to a richer musical experience and is used in double…

Cited by 0SourceScholar
2019

Robust Gridless Sound Field Decomposition Based on Structured Reciprocity Gap Functional in Spherical Harmonic Domain

ICASSP 2019accepted

A sound field reconstruction method for a region including sources is proposed. Under the assumption of spatial sparsity of the sources, this reconstruction problem has been solved by using sparse decomposition algorithms with the discretization of the target region. Since this discretization leads…

Cited by 0SourceScholar
2018

Sound Field Reproduction with Exterior Cancellation Using Analytical Weighting of Harmonic Coefficients

ICASSP 2018accepted

A method for sound field reproduction with the suppression of exterior radiation is proposed, which makes it possible to synthesize a desired sound field in a reverberant environment without prior knowledge of the transfer functions of the multiple loudspeakers. The objective function used to achiev…

Cited by 0SourceScholar
2018

Text-to-Speech Synthesis Using STFT Spectra Based on Low-/Multi-Resolution Generative Adversarial Networks

ICASSP 2018accepted

This paper proposes novel training algorithms for vocoder-free statistical parametric speech synthesis (SPSS) using short-term Fourier transform (STFT) spectra. Recently, text-to-speech synthesis using STFT spectra has been investigated since it can avoid quality degradation caused by the vocoder-ba…

Cited by 0SourceScholar
2018

Vectorwise Coordinate Descent Algorithm for Spatially Regularized Independent Low-Rank Matrix Analysis

ICASSP 2018accepted

Audio source separation is an important problem for many audio applications. Independent low-rank matrix analysis (ILRMA) is a recently proposed algorithm that employs the statistical independence between sources and the low-rankness of the time-frequency structure in each source. As reported in thi…

Cited by 0SourceScholar
2017

Blind source separation based on independent low-rank matrix analysis with sparse regularization for time-series activity

ICASSP 2017accepted

In this paper, we propose a new blind source separation (BSS) method based on independent low-rank matrix analysis (ILRMA) with novel sparse regularization. ILRMA is a recently proposed BSS algorithm that simultaneously estimates a demixing matrix and source spectrogram models based on nonnegative m…

Cited by 0SourceScholar
2017

Listening-area-informed sound field reproduction based on circular harmonic expansion

ICASSP 2017accepted

A sound field reproduction method that exploits prior information on listening areas is proposed. Most current methods are aimed at reproducing the sound field over the entire space or around listener locations. We formulate the objective function for this problem as the expectation minimization of…

Cited by 0SourceScholar
2017

Spatio-temporal sparse sound field decomposition considering acoustic source signal characteristics

ICASSP 2017accepted

We propose a sound field decomposition method that takes into consideration spatio-temporal sparsity. It has been proved that sparse representation of a sound field is effective in reducing errors originating from spatial aliasing artifacts compared with conventional plane wave decomposition. In mos…

Cited by 0SourceScholar
2017

Training algorithm to deceive Anti-Spoofing Verification for DNN-based speech synthesis

ICASSP 2017accepted

This paper proposes a novel training algorithm for high-quality Deep Neural Network (DNN)-based speech synthesis. The parameters of synthetic speech tend to be over-smoothed, and this causes significant quality degradation in synthetic speech. The proposed algorithm takes into account an Anti-Spoofi…

Cited by 0SourceScholar
2016

Multichannel blind source separation based on non-negative tensor factorization in wavenumber domain

ICASSP 2016accepted

Multichannel non-negative matrix factorization based on a spatial covariance model is one of the most promising techniques for blind source separation. However, this approach is not tractable for a large number of microphones, M, because the computational cost is of order O(M <sup xmlns:mml="http://…

Cited by 0SourceScholar
2016

Sound field decomposition in reverberant environment using sparse and low-rank signal models

ICASSP 2016accepted

A sound field decomposition method for a reverberant environment is proposed. Sound field decomposition is the foundation of various acoustic signal processing applications and enables the estimation of the entire sound field from pressure measurements. Although spatial Fourier analysis of the sound…

Cited by 0SourceScholar
2016

Sparse sound field decomposition with multichannel extension of complex NMF

ICASSP 2016accepted

A sparse sound field decomposition method using prior information on source signals in the time-frequency domain is proposed. Sparse sound field decomposition has been proved to be effective for various acoustic signal processing applications. Current methods for sparse decomposition are based only…

Cited by 0SourceScholar
2015

Efficient multichannel nonnegative matrix factorization exploiting rank-1 spatial model

ICASSP 2015accepted

This paper proposes a new efficient multichannel nonnegative matrix factorization (NMF) method. Recently, multichannel NMF (MNMF) has been proposed as a means of solving the blind source separation problem. This method estimates a mixing system of sources and attempts to separate them in a blind fas…

Cited by 0SourceScholar
2015

Statistical modeling of binaural signal and its application to binaural source separation

ICASSP 2015accepted

This paper addresses a new statistical model of binaural signals and its application to efficient binaural source separation. Binaural source separation is always required to retain a spatial cue of the separated sound, such as a head-related transfer function (HRTF). However, the direct use of an H…

Cited by 0SourceScholar
2015

Structured sparse signal models and decomposition algorithm for super-resolution in sound field recording and reproduction

ICASSP 2015accepted

A method for achieving super-resolution of sound field recording and reproduction is proposed. To obtain driving signals of loudspeakers for reproduction from received signals of microphones, sparse signal decomposition makes it possible to reduce spatial aliasing artifacts when the number of microp…

Cited by 0SourceScholar