← Search

Masahito Togami

16 accepted papers

2024

Real-Time Stereo Speech Enhancement with Spatial-Cue Preservation Based on Dual-Path Structure

ICASSP 2024accepted

We introduce a real-time, multichannel speech enhancement algorithm which maintains the spatial cues of stereo recordings including two speech sources. Recognizing that each source has unique spatial information, our method utilizes a dual-path structure, ensuring the spatial cues remain unaffected…

Cited by 0SourceScholar
2021

Joint Dereverberation and Separation With Iterative Source Steering

ICASSP 2021accepted

We propose a new algorithm for joint dereverberation and blind source separation (DR-BSS). Our work builds upon the IRLMA-T framework that applies a unified filter combining dereverberation and separation. One drawback of this framework is that it requires several matrix inversions, an operation inh…

Cited by 0SourceScholar
2021

Refinement of Direction of Arrival Estimators by Majorization-Minimization Optimization on the Array Manifold

ICASSP 2021accepted

We propose a generalized formulation of direction of arrival estimation that includes many existing methods such as steered response power, subspace, coherent and incoherent, as well as speech sparsity-based methods. Unlike most conventional methods that rely exclusively on grid search, we introduce…

Cited by 0SourceScholar
2020

Consistency-Aware Multi-Channel Speech Enhancement Using Deep Neural Networks

ICASSP 2020accepted

This paper proposes a deep neural network (DNN)–based multichannel speech enhancement system in which a DNN is trained to maximize the quality of the enhanced time-domain signal. DNN-based multi-channel speech enhancement is often conducted in the time-frequency (T-F) domain because spatial filterin…

Cited by 0SourceScholar
2020

Deep Speech Extraction with Time-Varying Spatial Filtering Guided By Desired Direction Attractor

ICASSP 2020accepted

In this investigation, a deep neural network (DNN) based speech extraction method is proposed to enhance a speech signal propagating from the desired direction. The proposed method integrates knowledge based on a sound propagation model and the time-varying characteristics of a speech source, into a…

Cited by 0SourceScholar
2020

Joint Training of Deep Neural Networks for Multi-Channel Dereverberation and Speech Source Separation

ICASSP 2020accepted

In this paper, we propose a joint training of two deep neural networks (DNNs) for dereverberation and speech source separation. The proposed method connects the first DNN, the dereverberation part, the second DNN, and the speech source separation part in a cascade manner. The proposed method does no…

Cited by 9SourceScholar
2020

Multi-Channel Speech Source Separation and Dereverberation With Sequential Integration of Determined and Underdetermined Models

ICASSP 2020accepted

In this paper, we propose a joint multi-channel speech source separation and dereverberation method in which multiple speech sources and late reverberation are separated in an unsupervised manner. The proposed method jointly optimizes an auto-regressive (AR) model based speech dereverberation and a…

Cited by 0SourceScholar
2020

Scene-Dependent Acoustic Event Detection with Scene Conditioning and Fake-Scene-Conditioned Loss

ICASSP 2020accepted

In this paper, we propose scene-dependent acoustic event detection (AED) with scene conditioning and fake-scene-conditioned loss. The proposed method employs a multitask network, that has not only AED part but also acoustic scene classification (ASC). The scenes predicted by ASC are employed as an a…

Cited by 0SourceScholar
2020

Unsupervised Training for Deep Speech Source Separation with Kullback-Leibler Divergence Based Probabilistic Loss Function

ICASSP 2020accepted

In this paper, we propose a multi-channel speech source separation method with a deep neural network (DNN) which is trained under the condition that no clean signal is available. As an alternative to a clean signal, the proposed method adopts an estimated speech signal by an unsupervised speech sour…

Cited by 0SourceScholar
2019

Simultaneous Optimization of Forgetting Factor and Time-frequency Mask for Block Online Multi-channel Speech Enhancement

ICASSP 2019accepted

In this paper, we propose a block-online multi-channel speech enhancement technique which simultaneously optimizes time-frequency masks and forgetting factors for estimation of multichannel covariance matrices of the desired speech signal and the noise signal so as to maximize speech enhancement per…

Cited by 0SourceScholar