← Search

Tomohiro Nakatani

63 accepted papers

2026

Joint Enhancement and Classification using Coupled Diffusion Models of Signals and Logits

ICML 2026poster

Robust classification in noisy environments remains a fundamental challenge in machine learning. Standard approaches typically treat signal enhancement and classification as separate, sequential stages: first enhancing the signal and then applying a classifier. This approach fails to leverage the se…

Cited by 0SourceScholar
2026

REFERENCE MICROPHONE SELECTION FOR GUIDED SOURCE SEPARATION BASED ON THE NORMALIZED L-P NORM

ICASSP 2026poster

Guided Source Separation (GSS) is a popular front-end for distant automatic speech recognition (ASR) systems using spatially distributed microphones. When considering spatially distributed microphones, the choice of reference microphone may have a large influence on the quality of the output signal…

Cited by 0SourcePDFScholar
2025

A Hybrid Probabilistic-Deterministic Model Recursively Enhancing Speech

ICASSP 2025accepted

This paper introduces Probabilistic-Deterministic Recursive Enhancement (PDRE), an innovative iterative Speech Enhancement (SE) approach that integrates probabilistic and deterministic methodologies. Recent advancements in diffusion models have demonstrated the exceptional effectiveness of probabili…

Cited by 0SourceScholar
2025

SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model

ICASSP 2025accepted

Target sound extraction (TSE) consists of isolating a desired sound from a mixture of arbitrary sounds using clues to identify it. A TSE system requires solving two problems at once, identifying the target source and extracting the target signal from the mixture. For increased practicability, the sa…

Cited by 0SourceScholar
2024

Neural Network-Based Virtual Microphone Estimation with Virtual Microphone and Beamformer-Level Multi-Task Loss

ICASSP 2024accepted

Array processing performance depends on the number of microphones available. Virtual microphone estimation (VME) has been proposed to increase the number of microphone signals artificially. Neural network-based VME (NN-VME) trains an NN with a VM-level loss to predict a signal at a microphone locati…

Cited by 0SourceScholar
2023

Fast Online Source Steering Algorithm for Tracking Single Moving Source Using Online Independent Vector Analysis

ICASSP 2023accepted

We address the problem of separating moving sources using online independent vector analysis (IVA). To solve this problem, researchers have extended the iterative projection (IP) and iterative source steering (ISS) algorithms developed for batch auxiliary-function-based IVA (AuxIVA) to online scenar…

Cited by 0SourceScholar
2022

Importance of Switch Optimization Criterion in Switching WPE Dereverberation

ICASSP 2022accepted

Weighted prediction error (WPE) is a fundamental dereverberation method to predict the late reverberation component of an observed signal based on linear prediction (LP). Recently, WPE was extended to Switching WPE (SwWPE), which optimizes (i) multiple LP filters and (ii) switching parameters to det…

Cited by 0SourceScholar
2022

Multi-Frame Full-Rank Spatial Covariance Analysis for Underdetermined BSS in Reverberant Environments

ICASSP 2022accepted

Full-rank spatial covariance analysis (FCA) is a blind source separation (BSS) method, and can be applied to underdetermined cases where the sources outnumber the microphones. This paper proposes a new extension of FCA, aiming to improve BSS performance for mixtures in which the length of reverberat…

Cited by 0SourceScholar
2021

Blind and Neural Network-Guided Convolutional Beamformer for Joint Denoising, Dereverberation, and Source Separation

ICASSP 2021accepted

This paper proposes an approach for optimizing a Convolutional BeamFormer (CBF) that can jointly perform denoising (DN), dereverberation (DR), and source separation (SS). First, we develop a blind CBF optimization algorithm that requires no prior information on the sources or the room acoustics, by…

Cited by 0SourceScholar
2021

Convolutive Transfer Function Invariant SDR Training Criteria for Multi-Channel Reverberant Speech Separation

ICASSP 2021accepted

Time-domain training criteria have proven to be very effective for the separation of single-channel non-reverberant speech mixtures. Likewise, mask-based beamforming has shown impressive performance in multi-channel reverberant speech enhancement and source separation. Here, we propose to combine ne…

Cited by 0SourceScholar
2021

Data Fusion for Audiovisual Speaker Localization: Extending Dynamic Stream Weights to the Spatial Domain

ICASSP 2021accepted

Estimating the positions of multiple speakers can be helpful for tasks like automatic speech recognition or speaker diarization. Both applications benefit from a known speaker position when, for instance, applying beamforming or assigning unique speaker identities. Recently, several approaches utili…

Cited by 0SourceScholar
2021

End-to-End Dereverberation, Beamforming, and Speech Recognition with Improved Numerical Stability and Advanced Frontend

ICASSP 2021accepted

Recently, the end-to-end approach has been successfully applied to multi-speaker speech separation and recognition in both singlechannel and multichannel conditions. However, severe performance degradation is still observed in the reverberant and noisy scenarios, and there is still a large performan…

Cited by 0SourceScholar
2021

Low Latency Online Blind Source Separation Based on Joint Optimization with Blind Dereverberation

ICASSP 2021accepted

This paper presents a new low-latency online blind source separation (BSS) algorithm. Although algorithmic delay of a frequency domain online BSS can be reduced simply by shortening the short-time Fourier transform (STFT) frame length, it degrades the source separation performance in the presence of…

Cited by 0SourceScholar
2021

Neural Network-Based Virtual Microphone Estimator

ICASSP 2021accepted

Developing microphone array technologies for a small number of microphones is important due to the constraints of many devices. One direction to address this situation consists of virtually augmenting the number of microphone signals, e.g., based on several physical model assumptions. However, such…

Cited by 0SourceScholar
2021

Speaker Activity Driven Neural Speech Extraction

ICASSP 2021accepted

Target speech extraction, which extracts the speech of a target speaker in a mixture given auxiliary speaker clues, has recently received increased interest. Various clues have been investigated such as pre-recorded enrollment utterances, direction information, or video of the target speaker. In thi…

Cited by 0SourceScholar
2020

A Dynamic Stream Weight Backprop Kalman Filter for Audiovisual Speaker Tracking

ICASSP 2020accepted

Audiovisual speaker tracking is an application that has been tackled by a wide range of classical approaches based on Gaussian filters, most notably the well-known Kalman filter. Recently, a specific Kalman filter implementation was proposed for this task, which incorporated dynamic stream weights t…

Cited by 0SourceScholar
2020

Beam-TasNet: Time-domain Audio Separation Network Meets Frequency-domain Beamformer

ICASSP 2020accepted

Recent studies have shown that acoustic beamforming using a microphone array plays an important role in the construction of high-performance automatic speech recognition (ASR) systems, especially for noisy and overlapping speech conditions. In parallel with the success of multichannel beamforming fo…

Cited by 0SourceScholar
2020

Convergence-Guaranteed Independent Positive Semidefinite Tensor Analysis Based on Student's T Distribution

ICASSP 2020accepted

In this paper, we address a blind source separation (BSS) problem and propose a new extended framework of independent positive semidefinite tensor analysis (IPSDTA). IPSDTA is a state-of-the-art BSS method that enables us to take interfrequency correlations into account, but the generative model is…

Cited by 0SourceScholar
2020

DNN-supported Mask-based Convolutional Beamforming for Simultaneous Denoising, Dereverberation, and Source Separation

ICASSP 2020accepted

In this article, we investigate an integrated mask-based convolutional beamforming method for performing simultaneous denoising, dereverberation, and source separation. Conventionally, it is difficult for neural network-supported mask-based source separation to perform denoising and dereverberation…

Cited by 0SourceScholar
2020

End-to-End Training of Time Domain Audio Separation and Recognition

ICASSP 2020accepted

The rising interest in single-channel multi-speaker speech separation sparked development of End-to-End (E2E) approaches to multi-speaker speech recognition. However, up until now, state-of-the-art neural network-based time domain source separation has not yet been combined with E2E speech recogniti…

Cited by 0SourceScholar
2020

Improving Noise Robust Automatic Speech Recognition with Single-Channel Time-Domain Enhancement Network

ICASSP 2020accepted

With the advent of deep learning, research on noise-robust automatic speech recognition (ASR) has progressed rapidly. However, ASR performance in noisy conditions of single-channel systems remains unsatisfactory. Indeed, most single-channel speech enhancement (SE) methods (denoising) have brought on…

Cited by 0SourceScholar
2020

Improving Speaker Discrimination of Target Speech Extraction With Time-Domain Speakerbeam

ICASSP 2020accepted

Target speech extraction, which extracts a single target source in a mixture given clues about the target speaker, has attracted increasing attention. We have recently proposed SpeakerBeam, which exploits an adaptation utterance of the target speaker to extract his/her voice characteristics that are…

Cited by 0SourceScholar
2020

Jointly Optimal Dereverberation and Beamforming

ICASSP 2020accepted

We previously proposed an optimal (in the maximum likelihood sense) convolutional beamformer that can perform simultaneous denoising and dereverberation, and showed its superiority over the widely used cascade of a Weighted Prediction Error (WPE) dereverberation filter and a conventional Minimum-Pow…

Cited by 0SourceScholar
2020

Tackling Real Noisy Reverberant Meetings with All-Neural Source Separation, Counting, and Diarization System

ICASSP 2020accepted

Automatic meeting analysis is an essential fundamental technology required to let, e.g. smart devices follow and respond to our conversations. To achieve an optimal automatic meeting analysis, we previously proposed an all-neural approach that jointly solves source separation, speaker diarization an…

Cited by 0SourceScholar
2019

A Unified Framework for Feature-based Domain Adaptation of Neural Network Language Models

ICASSP 2019accepted

An important task for language models is the adaptation of general-domain models to specific target domains. For neural network-based language models, feature-based domain adaptation has been a popular method in previous research. Conventional methods use an adaptation feature providing context info…

Cited by 0SourceScholar
2019

A Unified Framework for Neural Speech Separation and Extraction

ICASSP 2019accepted

The development of deep learning techniques has triggered the active investigation of neural network-based speech enhancement approaches. In particular, single-channel blind (uninformed) speech separation and speaker-aware (informed) speech extraction have received increased interest. Blind speech s…

Cited by 0SourceScholar
2019

All-neural Online Source Separation, Counting, and Diarization for Meeting Analysis

ICASSP 2019accepted

Automatic meeting analysis comprises the tasks of speaker counting, speaker diarization, and the separation of overlapped speech, followed by automatic speech recognition. This all has to be carried out on arbitrarily long sessions and, ideally, in an online or block-online manner. While significant…

Cited by 0SourceScholar
2019

Compact Network for Speakerbeam Target Speaker Extraction

ICASSP 2019accepted

Speech separation that separates a mixture of speech signals into each of its sources has been an active research topic for a long time and has seen recent progress with the advent of deep learning. A related problem is target speaker extraction, i.e. extraction of only speech of a target speaker ou…

Cited by 0SourceScholar
2019

FastMNMF: Joint Diagonalization Based Accelerated Algorithms for Multichannel Nonnegative Matrix Factorization

ICASSP 2019accepted

A multichannel extension of nonnegative matrix factorization (NMF) for audio/music data, called multichannel NMF (MNMF), has been proposed by Sawada et al ["Multichannel extensions of non-negative matrix factorization with complex-valued data IEEE Trans. ASLP, vol. 21, no. 5, pp. 971-982, May 2013].…

Cited by 0SourceScholar
2019

ILP-based Compressive Speech Summarization with Content Word Coverage Maximization and Its Oracle Performance Analysis

ICASSP 2019accepted

We propose an integer linear programming (ILP)-based compressive speech summarization method that maximizes the coverage of content words in a resultant summary. It is an unsupervised method and, under the designed constraints, it performs a single-step globally optimal summarization of a given long…

Cited by 0SourceScholar
2019

Joint Optimization of Neural Network-based WPE Dereverberation and Acoustic Model for Robust Online ASR

ICASSP 2019accepted

Signal dereverberation using the Weighted Prediction Error (WPE) method has been proven to be an effective means to raise the accuracy of far-field speech recognition. First proposed as an iterative algorithm, follow-up works have reformulated it as a recursive least squares algorithm and therefore…

Cited by 0SourceScholar
2019

Mask-based MVDR Beamformer for Noisy Multisource Environments: Introduction of Time-varying Spatial Covariance Model

ICASSP 2019accepted

This paper proposes a method for designing a time-varying minimum variance distortionless response (MVDR) beamformer using time-frequency masks, with the aim of improving speech enhancement in noisy multi-speaker environments. A key to successful beamforming is to estimate accurately a time-varying…

Cited by 0SourceScholar
2019

Semi-supervised End-to-end Speech Recognition Using Text-to-speech and Autoencoders

ICASSP 2019accepted

We introduce speech and text autoencoders that share encoders and decoders with an automatic speech recognition (ASR) model to improve ASR performance with large speech only and text only training datasets. To build the speech and text autoencoders, we leverage state-of-the-art ASR and text-to-speec…

Cited by 0SourceScholar
2018

Dual Frequency- and Block-Permutation Alignment for Deep Learning Based Block-Online Blind Source Separation

ICASSP 2018accepted

Deep attractor networks (DANs) are a recently introduced method to blindly separate sources from spectral features of a monaural recording using bidirectional long short-term memory networks (BLSTMs). Due to the nature of BLSTMs, this is inherently not online-ready and resorting to operating on bloc…

Cited by 0SourceScholar
2018

Frame-by-Frame Closed-Form Update for Mask-Based Adaptive MVDR Beamforming

ICASSP 2018accepted

Beamforming approaches using time-frequency masks have recently been investigated and have shown promising results for noise robust automatic speech recognition (ASR) in many tasks. The time-frequency masks are estimated to compute the spatial statistics of target speech and noise signals, and then…

Cited by 0SourceScholar
2018

Listening to Each Speaker One by One with Recurrent Selective Hearing Networks

ICASSP 2018accepted

Deep learning-based single-channel source separation algorithms are currently being actively investigated. Among them, Deep Clustering (DC) and Deep Attractor Networks (DANs) have made it possible to separate an arbitrary number of speakers. In particular, they cleverly combine a neural network and…

Cited by 0SourceScholar
2018

Maximum-Likelihood Online Speaker Diarization in Noisy Meetings Based on Categorical Mixture Model and Probabilistic Spatial Dictionary

ICASSP 2018accepted

In this paper, we propose a maximum-likelihood online diarization method based on a probabilistic spatial dictionary. This dictionary consists of the given probability distribution of spatial features for each possible direction of arrival (DOA) of source signals. Recently, we have developed an onli…

Cited by 0SourceScholar
2018

Optimization of Speaker-Aware Multichannel Speech Extraction with ASR Criterion

ICASSP 2018accepted

This paper addresses the problem of recognizing speech corrupted by overlapping speakers in a multichannel setting. To extract a target speaker from the mixture, we use a neural network based beamformer which uses masks estimated by a neural network to compute statistically optimal spatial filters.…

Cited by 0SourceScholar
2018

Permutation-Free Cgmm: Complex Gaussian Mixture Model with Inverse Wishart Mixture Model Based Spatial Prior for Permutation-Free Source Separation and Source Counting

ICASSP 2018accepted

Here we propose a permutation-free cGMM (PF-cGMM), a new probabilistic model of observed mixtures, which can resolve permutation ambiguity between frequency bins, and is applicable even when the number of sources is unknown. A recently proposed complex Gaussian mixture model (cGMM) is highly effecti…

Cited by 0SourceScholar
2018

Rescoring N-Best Speech Recognition List Based on One-on-One Hypothesis Comparison Using Encoder-Classifier Model

ICASSP 2018accepted

This paper proposes a new model for accurately rescoring (reranking) N-best speech recognition hypothesis lists. The model is based on state-of-the-art neural networks (NNs) and provides the minimum necessary functionality to perform N-best rescoring, i.e. one-on-one hypothesis comparison on a given…

Cited by 0SourceScholar
2018

Sequence Training of Encoder-Decoder Model Using Policy Gradient for End-to-End Speech Recognition

ICASSP 2018accepted

The standard evaluation metric of automatic speech recognition (ASR) is the word error rate (WER), which measures the dissimilarity between recognized word sequences and their ground truth. Many training algorithms designed to reduce sequence-level errors such as WER have been proposed for hidden Ma…

Cited by 0SourceScholar
2018

Single Channel Target Speaker Extraction and Recognition with Speaker Beam

ICASSP 2018accepted

This paper addresses the problem of single channel speech recognition of a target speaker in a mixture of speech signals. We propose to exploit auxiliary speaker information provided by an adaptation utterance from the target speaker to extract and recognize only that speaker. Using such auxiliary i…

Cited by 0SourceScholar
2017

Cumulative moving averaged bottleneck speaker vectors for online speaker adaptation of CNN-based acoustic models

ICASSP 2017accepted

Adapting acoustic models to speakers have shown to greatly improve performance for many tasks. Among the adaptation approaches, exploiting auxiliary features characterizing speakers or environments has received great attention because they allow rapid adaptation, i.e. adaptation with limited amount…

Cited by 0SourceScholar
2017

Deep mixture density network for statistical model-based feature enhancement

ICASSP 2017accepted

We propose a novel framework designed to extend conventional deep neural network (DNN)-based feature enhancement approaches. In general, the conventional DNN-based feature enhancement framework aims to map input noisy observation to clean speech or a binary/ soft mask in a deterministic way, assumin…

Cited by 0SourceScholar
2017

Feedback connection for deep neural network-based acoustic modeling

ICASSP 2017accepted

The use of auxiliary features is an effective way to improve the performance of deep neural network (DNN)-based acoustic models. Most approaches use auxiliary features that represent the speaker or the environment. These auxiliary features are usually computed independently of the acoustic model. Th…

Cited by 0SourceScholar
2017

Integrating DNN-based and spatial clustering-based mask estimation for robust MVDR beamforming

ICASSP 2017accepted

Recently, time-frequency mask-based beamforming has been extensively studied as the frontend of deep neural network (DNN) based automatic speech recognition (ASR) in noisy environments. Two mask estimation approaches have been separately developed for this beamforming method, namely the the DNN-base…

Cited by 0SourceScholar
2017

Online environmental adaptation of CNN-based acoustic models using spatial diffuseness features

ICASSP 2017accepted

We propose a new concept for adapting CNN-based acoustic models using spatial diffuseness features as auxiliary information about the acoustic environment: the spatial diffuseness features are simultaneously employed as acoustic-model input features and to estimate environmental cues for context ada…

Cited by 0SourceScholar
2017

Probabilistic spatial dictionary based online adaptive beamforming for meeting recognition in noisy and reverberant environments

ICASSP 2017accepted

Here we propose online adaptive beamforming for automatic speech recognition (ASR) in meetings in noisy, reverberant environments. The proposed method is based on recently developed mask-based beamforming, in which accurate mask estimation and diarization are paramount. Real-world experiments have s…

Cited by 0SourceScholar
2017

Unsupervised utterance-wise beamformer estimation with speech recognition-level criterion

ICASSP 2017accepted

In this paper, we perform beamforming with a speech recognition-level criterion. A beamformer is usually designed by optimizing signal-level criteria, e.g., by minimizing the beamformer output covariance or by maximizing the signal-to-noise ratio (SNR). Such signal-level criteria do not always guara…

Cited by 0SourceScholar
2016

A generative-discriminative hybrid approach to multi-channel noise reduction for robust automatic speech recognition

ICASSP 2016accepted

In the recent years, discriminative models have become a very attractive utility and gained a lot of attention in the speech research community, encompassing both front and back-end methods, thanks to their prominent discriminative power and the availability of improved training strategies. When it…

Cited by 0SourceScholar
2016

Context adaptive deep neural networks for fast acoustic model adaptation in noisy conditions

ICASSP 2016accepted

Deep neural network (DNN) based acoustic models have greatly improved the performance of automatic speech recognition (ASR) for various tasks. Further performance improvements have been reported when making DNNs aware of the acoustic context (e.g. speaker or environment) for example by adding auxili…

Cited by 0SourceScholar
2016

Modeling audio directional statistics using a complex bingham mixture model for blind source extraction from diffuse noise

ICASSP 2016accepted

Mask estimation is a central task in blind signal processing including source separation, denoising, and multi-source localization. In this paper, we define a complex Bingham mixture model (cBMM), and propose it as a model of directional statistics for mask estimation. The complex Bingham distributi…

Cited by 0SourceScholar
2016

Multi-pass feature enhancement based on generative-discriminative hybrid approach for noise robust speech recognition

ICASSP 2016accepted

This paper presents multi-pass feature enhancement technique that consists of three processing passes. In the proposed method, the first pass was described in our previous work, and consists of model-based feature enhancement realized by employing a generative-discriminative hybrid approach with Gau…

Cited by 0SourceScholar
2016

Noise robust speech recognition using recent developments in neural networks for computer vision

ICASSP 2016accepted

Convolutional Neural Networks (CNNs) are superior to fully connected neural networks in various speech recognition tasks and the advantage is pronounced in noisy environments. In recent years, many techniques have been proposed in the computer vision community to improve CNN's classification perform…

Cited by 0SourceScholar
2016

Real-time integration of statistical model-based speech enhancement with unsupervised noise PSD estimation using microphone array

ICASSP 2016accepted

We propose a technique of multi-channel speech enhancement based on integration of beamforming and statistical model-based speech enhancement to clearly extract the target speech, even in very noisy environments. Conventional microphone array-based techniques estimate speech and noise power spectral…

Cited by 0SourceScholar
2016

Robust MVDR beamforming using time-frequency masks for online/offline ASR in noise

ICASSP 2016accepted

This paper considers acoustic beamforming for noise robust automatic speech recognition (ASR). A beamformer attenuates background noise by enhancing sound components coming from a direction specified by a steering vector. Hence, accurate steering vector estimation is paramount for successful noise r…

Cited by 0SourceScholar
2016

Spatial correlation model based observation vector clustering and MVDR beamforming for meeting recognition

ICASSP 2016accepted

This paper addresses a minimum variance distortionless response (MVDR) beamforming based speech enhancement approach for meeting speech recognition. In a meeting situation, speaker overlaps and noise signals are not negligible. To handle these issues, we employ MVDR beamforming, where accurate estim…

Cited by 0SourceScholar
2015

Context adaptive deep neural networks for fast acoustic model adaptation

ICASSP 2015accepted

Deep neural networks (DNNs) are widely used for acoustic modeling in automatic speech recognition (ASR), since they greatly outperform legacy Gaussian mixture model-based systems. However, the levels of performance achieved by current DNN-based systems remain far too low in many tasks, e.g. when the…

Cited by 0SourceScholar
2015

Exploring multi-channel features for denoising-autoencoder-based speech enhancement

ICASSP 2015accepted

This paper investigates a multi-channel denoising autoencoder (DAE)-based speech enhancement approach. In recent years, deep neural network (DNN)-based monaural speech enhancement and robust automatic speech recognition (ASR) approaches have attracted much attention due to their high performance. Al…

Cited by 0SourceScholar
2015

Far-field speech recognition using CNN-DNN-HMM with convolution in time

ICASSP 2015accepted

Recent studies in speech recognition have shown that the performance of convolutional neural networks (CNNs) is superior to that of fully connected deep neural networks (DNNs). In this paper, we explore the use of CNNs in far-field speech recognition for dealing with reverberation, which blurs spect…

Cited by 0SourceScholar
2015

Feature enhancement based on generative-discriminative hybrid approach with gmms and DNNS for noise robust speech recognition

ICASSP 2015accepted

This paper presents a technique that combines generative and discriminative approaches with Gaussian mixture models (GMMs) and deep neural networks (DNNs) for model-based feature enhancement. Typical model-based feature enhancement employs a generative model approach. The enhanced features are obtai…

Cited by 0SourceScholar
2015

Modeling inter-node acoustic dependencies with Restricted Boltzmann Machine for distributed microphone array based BSS

ICASSP 2015accepted

An accurate estimation of a source activity information is essential for many speech enhancement algorithms including blind source separation (BSS). In this paper, we propose a novel BSS method that accurately models and estimates the source activity in distributed microphone array (DMA) scenarios.…

Cited by 0SourceScholar