← Search

Nobutaka Ito

10 accepted papers

2025

30+ Years of Source Separation Research: Achievements and Future Challenges

ICASSP 2025accepted

Source separation (SS) of acoustic signals is a research field that emerged in the mid-1990s and has flourished ever since. On the occasion of ICASSP’s 50<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">th</sup> anniversary, we review the major contribut…

Cited by 23SourceScholar
2019

FastMNMF: Joint Diagonalization Based Accelerated Algorithms for Multichannel Nonnegative Matrix Factorization

ICASSP 2019accepted

A multichannel extension of nonnegative matrix factorization (NMF) for audio/music data, called multichannel NMF (MNMF), has been proposed by Sawada et al ["Multichannel extensions of non-negative matrix factorization with complex-valued data IEEE Trans. ASLP, vol. 21, no. 5, pp. 971-982, May 2013].…

Cited by 0SourceScholar
2018

Frame-by-Frame Closed-Form Update for Mask-Based Adaptive MVDR Beamforming

ICASSP 2018accepted

Beamforming approaches using time-frequency masks have recently been investigated and have shown promising results for noise robust automatic speech recognition (ASR) in many tasks. The time-frequency masks are estimated to compute the spatial statistics of target speech and noise signals, and then…

Cited by 0SourceScholar
2018

Maximum-Likelihood Online Speaker Diarization in Noisy Meetings Based on Categorical Mixture Model and Probabilistic Spatial Dictionary

ICASSP 2018accepted

In this paper, we propose a maximum-likelihood online diarization method based on a probabilistic spatial dictionary. This dictionary consists of the given probability distribution of spatial features for each possible direction of arrival (DOA) of source signals. Recently, we have developed an onli…

Cited by 0SourceScholar
2018

Permutation-Free Cgmm: Complex Gaussian Mixture Model with Inverse Wishart Mixture Model Based Spatial Prior for Permutation-Free Source Separation and Source Counting

ICASSP 2018accepted

Here we propose a permutation-free cGMM (PF-cGMM), a new probabilistic model of observed mixtures, which can resolve permutation ambiguity between frequency bins, and is applicable even when the number of sources is unknown. A recently proposed complex Gaussian mixture model (cGMM) is highly effecti…

Cited by 0SourceScholar
2017

Integrating DNN-based and spatial clustering-based mask estimation for robust MVDR beamforming

ICASSP 2017accepted

Recently, time-frequency mask-based beamforming has been extensively studied as the frontend of deep neural network (DNN) based automatic speech recognition (ASR) in noisy environments. Two mask estimation approaches have been separately developed for this beamforming method, namely the the DNN-base…

Cited by 0SourceScholar
2017

Probabilistic spatial dictionary based online adaptive beamforming for meeting recognition in noisy and reverberant environments

ICASSP 2017accepted

Here we propose online adaptive beamforming for automatic speech recognition (ASR) in meetings in noisy, reverberant environments. The proposed method is based on recently developed mask-based beamforming, in which accurate mask estimation and diarization are paramount. Real-world experiments have s…

Cited by 0SourceScholar
2016

Modeling audio directional statistics using a complex bingham mixture model for blind source extraction from diffuse noise

ICASSP 2016accepted

Mask estimation is a central task in blind signal processing including source separation, denoising, and multi-source localization. In this paper, we define a complex Bingham mixture model (cBMM), and propose it as a model of directional statistics for mask estimation. The complex Bingham distributi…

Cited by 14SourceScholar
2016

Robust MVDR beamforming using time-frequency masks for online/offline ASR in noise

ICASSP 2016accepted

This paper considers acoustic beamforming for noise robust automatic speech recognition (ASR). A beamformer attenuates background noise by enhancing sound components coming from a direction specified by a steering vector. Hence, accurate steering vector estimation is paramount for successful noise r…

Cited by 0SourceScholar