← Search

Rintaro Ikeshita

14 accepted papers

2024

How Does End-To-End Speech Recognition Training Impact Speech Enhancement Artifacts?

ICASSP 2024accepted

Jointly training a speech enhancement (SE) front-end and an automatic speech recognition (ASR) back-end has been investigated as a way to mitigate the influence of processing distortion generated by single-channel SE on ASR. In this paper, we investigate the effect of such joint training on the sign…

Cited by 0SourceScholar
2024

Neural Network-Based Virtual Microphone Estimation with Virtual Microphone and Beamformer-Level Multi-Task Loss

ICASSP 2024accepted

Array processing performance depends on the number of microphones available. Virtual microphone estimation (VME) has been proposed to increase the number of microphone signals artificially. Neural network-based VME (NN-VME) trains an NN with a VM-level loss to predict a signal at a microphone locati…

Cited by 0SourceScholar
2023

Fast Online Source Steering Algorithm for Tracking Single Moving Source Using Online Independent Vector Analysis

ICASSP 2023accepted

We address the problem of separating moving sources using online independent vector analysis (IVA). To solve this problem, researchers have extended the iterative projection (IP) and iterative source steering (ISS) algorithms developed for batch auxiliary-function-based IVA (AuxIVA) to online scenar…

Cited by 0SourceScholar
2022

Importance of Switch Optimization Criterion in Switching WPE Dereverberation

ICASSP 2022accepted

Weighted prediction error (WPE) is a fundamental dereverberation method to predict the late reverberation component of an observed signal based on linear prediction (LP). Recently, WPE was extended to Switching WPE (SwWPE), which optimizes (i) multiple LP filters and (ii) switching parameters to det…

Cited by 0SourceScholar
2022

Multi-Frame Full-Rank Spatial Covariance Analysis for Underdetermined BSS in Reverberant Environments

ICASSP 2022accepted

Full-rank spatial covariance analysis (FCA) is a blind source separation (BSS) method, and can be applied to underdetermined cases where the sources outnumber the microphones. This paper proposes a new extension of FCA, aiming to improve BSS performance for mixtures in which the length of reverberat…

Cited by 0SourceScholar
2021

Blind and Neural Network-Guided Convolutional Beamformer for Joint Denoising, Dereverberation, and Source Separation

ICASSP 2021accepted

This paper proposes an approach for optimizing a Convolutional BeamFormer (CBF) that can jointly perform denoising (DN), dereverberation (DR), and source separation (SS). First, we develop a blind CBF optimization algorithm that requires no prior information on the sources or the room acoustics, by…

Cited by 0SourceScholar
2021

Low Latency Online Blind Source Separation Based on Joint Optimization with Blind Dereverberation

ICASSP 2021accepted

This paper presents a new low-latency online blind source separation (BSS) algorithm. Although algorithmic delay of a frequency domain online BSS can be reduced simply by shortening the short-time Fourier transform (STFT) frame length, it degrades the source separation performance in the presence of…

Cited by 0SourceScholar
2021

Neural Network-Based Virtual Microphone Estimator

ICASSP 2021accepted

Developing microphone array technologies for a small number of microphones is important due to the constraints of many devices. One direction to address this situation consists of virtually augmenting the number of microphone signals, e.g., based on several physical model assumptions. However, such…

Cited by 0SourceScholar
2020

Beam-TasNet: Time-domain Audio Separation Network Meets Frequency-domain Beamformer

ICASSP 2020accepted

Recent studies have shown that acoustic beamforming using a microphone array plays an important role in the construction of high-performance automatic speech recognition (ASR) systems, especially for noisy and overlapping speech conditions. In parallel with the success of multichannel beamforming fo…

Cited by 0SourceScholar
2020

Convergence-Guaranteed Independent Positive Semidefinite Tensor Analysis Based on Student's T Distribution

ICASSP 2020accepted

In this paper, we address a blind source separation (BSS) problem and propose a new extended framework of independent positive semidefinite tensor analysis (IPSDTA). IPSDTA is a state-of-the-art BSS method that enables us to take interfrequency correlations into account, but the generative model is…

Cited by 0SourceScholar
2020

DNN-supported Mask-based Convolutional Beamforming for Simultaneous Denoising, Dereverberation, and Source Separation

ICASSP 2020accepted

In this article, we investigate an integrated mask-based convolutional beamforming method for performing simultaneous denoising, dereverberation, and source separation. Conventionally, it is difficult for neural network-supported mask-based source separation to perform denoising and dereverberation…

Cited by 25SourceScholar
2019

Acoustic Modeling for Distant Multi-talker Speech Recognition with Single- and Multi-channel Branches

ICASSP 2019accepted

This paper presents a novel heterogeneous-input multi-channel acoustic model (AM) that has both single-channel and multi-channel input branches. In our proposed training pipeline, a single-channel AM is trained first, then a multi-channel AM is trained starting from the single-channel AM with a rand…

Cited by 0SourceScholar
2018

Independent Low-Rank Matrix Analysis Based on Multivariate Complex Exponential Power Distribution

ICASSP 2018accepted

Independent low-rank matrix analysis (ILRMA), a unified method of independent vector analysis (IVA) and nonnegative matrix factorization (NMF), is a state-of-the-art blind source separation method for convolutive mixtures. Although ILRMA provides high separation performance for music signals whose s…

Cited by 0SourceScholar