← Search

Tomohiko Nakamura

7 accepted papers

2026

DISSECTING PERFORMANCE DEGRADATION IN AUDIO SOURCE SEPARATION UNDER SAMPLING FREQUENCY MISMATCH

ICASSP 2026poster

Audio processing methods based on deep neural networks are typically trained at a single sampling frequency (SF). To handle untrained SFs, signal resampling is commonly employed, but it can degrade performance, particularly when the input SF is lower than the trained SF. This paper investigates the…

Cited by 0SourcePDFScholar
2026

PHASE-RETRIEVAL-BASED PHYSICS-INFORMED NEURAL NETWORKS FOR ACOUSTIC MAGNITUDE FIELD RECONSTRUCTION

ICASSP 2026poster

We propose a method for estimating the magnitude distribution of an acoustic field from spatially sparse magnitude measurements. Such a method is useful when phase measurements are unreliable or inaccessible. Physics-informed neural networks (PINNs) have shown promise for sound field estimation by i…

Cited by 0SourcePDFScholar
2023

jaCappella Corpus: A Japanese a Cappella Vocal Ensemble Corpus

ICASSP 2023accepted

We construct a corpus of Japanese a cappella vocal ensembles (ja-Cappella corpus) for vocal ensemble separation and synthesis. It consists of 35 copyright-cleared vocal ensemble songs and their audio recordings of individual voice parts. These songs were arranged from out-of-copyright Japanese child…

Cited by 0SourceScholar
2022

Differentiable Digital Signal Processing Mixture Model for Synthesis Parameter Extraction from Mixture of Harmonic Sounds

ICASSP 2022accepted

A differentiable digital signal processing (DDSP) autoencoder is a musical sound synthesizer that combines a deep neural network (DNN) and spectral modeling synthesis. It allows us to flexibly edit sounds by changing the fundamental frequency, timbre feature, and loudness (synthesis parameters) extr…

Cited by 0SourceScholar
2020

Time-Domain Audio Source Separation Based on Wave-U-Net Combined with Discrete Wavelet Transform

ICASSP 2020accepted

We propose a time-domain audio source separation method using down-sampling (DS) and up-sampling (US) layers based on a discrete wavelet transform (DWT). The proposed method is based on one of the state-of-the-art deep neural networks, Wave-U-Net, which successively down-samples and up-samples featu…

Cited by 0SourceScholar
2016

Shifted and convolutive source-filter non-negative matrix factorization for monaural audio source separation

ICASSP 2016accepted

This paper proposes an extension of non-negative matrix factorization (NMF), which combines the shifted NMF model with the source-filter model. Shifted NMF was proposed as a powerful approach for monaural source separation and multiple fundamental frequency (F0) estimation, which is particularly uni…

Cited by 0SourceScholar
2015

Lp-norm non-negative matrix factorization and its application to singing voice enhancement

ICASSP 2015accepted

Measures of sparsity are useful in many aspects of audio signal processing including speech enhancement, audio coding and singing voice enhancement, and the well-known method for these applications is non-negative matrix factorization (NMF), which decomposes a non-negative data matrix into two non-n…

Cited by 0SourceScholar