← Search

Keisuke Kinoshita

44 accepted papers

2025

Long-Form Speech Generation with Spoken Language Models

ICML 2025oral

We consider the generative modeling of speech over multiple minutes, a requirement for long-form multimedia generation and audio-native voice assistants. However, textless spoken language models struggle to generate plausible speech past tens of seconds, due to high temporal resolution of speech tok…

2023

On Word Error Rate Definitions and Their Efficient Computation for Multi-Speaker Speech Recognition Systems

ICASSP 2023accepted

We propose a general framework to compute the word error rate (WER) of ASR systems that process recordings containing multiple speakers at their input and that produce multiple output word sequences (MIMO). Such ASR systems are typically required, e.g., for meeting transcription. We provide an effic…

Cited by 0SourceScholar
2022

Importance of Switch Optimization Criterion in Switching WPE Dereverberation

ICASSP 2022accepted

Weighted prediction error (WPE) is a fundamental dereverberation method to predict the late reverberation component of an observed signal based on linear prediction (LP). Recently, WPE was extended to Switching WPE (SwWPE), which optimizes (i) multiple LP filters and (ii) switching parameters to det…

Cited by 0SourceScholar
2022

Learning to Enhance or Not: Neural Network-Based Switching of Enhanced and Observed Signals for Overlapping Speech Recognition

ICASSP 2022accepted

The combination of a deep neural network (DNN) -based speech enhancement (SE) front-end and an automatic speech recognition (ASR) back-end is a widely used approach to implement overlapping speech recognition. However, the SE front-end generates processing artifacts that can degrade the ASR performa…

Cited by 0SourceScholar
2022

Multi-Frame Full-Rank Spatial Covariance Analysis for Underdetermined BSS in Reverberant Environments

ICASSP 2022accepted

Full-rank spatial covariance analysis (FCA) is a blind source separation (BSS) method, and can be applied to underdetermined cases where the sources outnumber the microphones. This paper proposes a new extension of FCA, aiming to improve BSS performance for mixtures in which the length of reverberat…

Cited by 0SourceScholar
2022

SA-SDR: A Novel Loss Function for Separation of Meeting Style Data

ICASSP 2022accepted

Many state-of-the-art neural network-based source separation systems use the averaged Signal-to-Distortion Ratio (SDR) as a training objective function. The basic SDR is, however, undefined if the network reconstructs the reference signal perfectly or if the reference signal contains silence, e.g.,…

Cited by 0SourceScholar
2022

Tight Integration Of Neural- And Clustering-Based Diarization Through Deep Unfolding Of Infinite Gaussian Mixture Model

ICASSP 2022accepted

Speaker diarization has been investigated extensively as an important central task for meeting analysis. Recent trend shows that integration of end-to-end neural (EEND)- and clustering-based diarization is a promising approach to handle realistic conversational data containing overlapped speech with…

Cited by 0SourceScholar
2021

Blind and Neural Network-Guided Convolutional Beamformer for Joint Denoising, Dereverberation, and Source Separation

ICASSP 2021accepted

This paper proposes an approach for optimizing a Convolutional BeamFormer (CBF) that can jointly perform denoising (DN), dereverberation (DR), and source separation (SS). First, we develop a blind CBF optimization algorithm that requires no prior information on the sources or the room acoustics, by…

Cited by 0SourceScholar
2021

Convolutive Transfer Function Invariant SDR Training Criteria for Multi-Channel Reverberant Speech Separation

ICASSP 2021accepted

Time-domain training criteria have proven to be very effective for the separation of single-channel non-reverberant speech mixtures. Likewise, mask-based beamforming has shown impressive performance in multi-channel reverberant speech enhancement and source separation. Here, we propose to combine ne…

Cited by 0SourceScholar
2021

Data Fusion for Audiovisual Speaker Localization: Extending Dynamic Stream Weights to the Spatial Domain

ICASSP 2021accepted

Estimating the positions of multiple speakers can be helpful for tasks like automatic speech recognition or speaker diarization. Both applications benefit from a known speaker position when, for instance, applying beamforming or assigning unique speaker identities. Recently, several approaches utili…

Cited by 0SourceScholar
2021

Dual-Path Modeling for Long Recording Speech Separation in Meetings

ICASSP 2021accepted

The continuous speech separation (CSS) is a task to separate the speech sources from a long, partially overlapped recording, which involves a varying number of speakers. A straightforward extension of conventional utterance-level speech separation to the CSS task is to segment the long recording wit…

Cited by 0SourceScholar
2021

End-to-End Dereverberation, Beamforming, and Speech Recognition with Improved Numerical Stability and Advanced Frontend

ICASSP 2021accepted

Recently, the end-to-end approach has been successfully applied to multi-speaker speech separation and recognition in both singlechannel and multichannel conditions. However, severe performance degradation is still observed in the reverberant and noisy scenarios, and there is still a large performan…

Cited by 0SourceScholar
2021

Integrating End-to-End Neural and Clustering-Based Diarization: Getting the Best of Both Worlds

ICASSP 2021accepted

Recent diarization technologies can be categorized into two approaches, i.e., clustering and end-to-end neural approaches, which have different pros and cons. The clustering-based approaches assign speaker labels to speech regions by clustering speaker embeddings such as x-vectors. While it can be s…

Cited by 0SourceScholar
2021

Low Latency Online Blind Source Separation Based on Joint Optimization with Blind Dereverberation

ICASSP 2021accepted

This paper presents a new low-latency online blind source separation (BSS) algorithm. Although algorithmic delay of a frequency domain online BSS can be reduced simply by shortening the short-time Fourier transform (STFT) frame length, it degrades the source separation performance in the presence of…

Cited by 0SourceScholar
2021

Neural Network-Based Virtual Microphone Estimator

ICASSP 2021accepted

Developing microphone array technologies for a small number of microphones is important due to the constraints of many devices. One direction to address this situation consists of virtually augmenting the number of microphone signals, e.g., based on several physical model assumptions. However, such…

Cited by 0SourceScholar
2021

Speaker Activity Driven Neural Speech Extraction

ICASSP 2021accepted

Target speech extraction, which extracts the speech of a target speaker in a mixture given auxiliary speaker clues, has recently received increased interest. Various clues have been investigated such as pre-recorded enrollment utterances, direction information, or video of the target speaker. In thi…

Cited by 0SourceScholar
2020

A Dynamic Stream Weight Backprop Kalman Filter for Audiovisual Speaker Tracking

ICASSP 2020accepted

Audiovisual speaker tracking is an application that has been tackled by a wide range of classical approaches based on Gaussian filters, most notably the well-known Kalman filter. Recently, a specific Kalman filter implementation was proposed for this task, which incorporated dynamic stream weights t…

Cited by 0SourceScholar
2020

Beam-TasNet: Time-domain Audio Separation Network Meets Frequency-domain Beamformer

ICASSP 2020accepted

Recent studies have shown that acoustic beamforming using a microphone array plays an important role in the construction of high-performance automatic speech recognition (ASR) systems, especially for noisy and overlapping speech conditions. In parallel with the success of multichannel beamforming fo…

Cited by 0SourceScholar
2020

DNN-supported Mask-based Convolutional Beamforming for Simultaneous Denoising, Dereverberation, and Source Separation

ICASSP 2020accepted

In this article, we investigate an integrated mask-based convolutional beamforming method for performing simultaneous denoising, dereverberation, and source separation. Conventionally, it is difficult for neural network-supported mask-based source separation to perform denoising and dereverberation…

Cited by 25SourceScholar
2020

End-to-End Training of Time Domain Audio Separation and Recognition

ICASSP 2020accepted

The rising interest in single-channel multi-speaker speech separation sparked development of End-to-End (E2E) approaches to multi-speaker speech recognition. However, up until now, state-of-the-art neural network-based time domain source separation has not yet been combined with E2E speech recogniti…

Cited by 0SourceScholar
2020

Improving Noise Robust Automatic Speech Recognition with Single-Channel Time-Domain Enhancement Network

ICASSP 2020accepted

With the advent of deep learning, research on noise-robust automatic speech recognition (ASR) has progressed rapidly. However, ASR performance in noisy conditions of single-channel systems remains unsatisfactory. Indeed, most single-channel speech enhancement (SE) methods (denoising) have brought on…

Cited by 0SourceScholar
2020

Improving Speaker Discrimination of Target Speech Extraction With Time-Domain Speakerbeam

ICASSP 2020accepted

Target speech extraction, which extracts a single target source in a mixture given clues about the target speaker, has attracted increasing attention. We have recently proposed SpeakerBeam, which exploits an adaptation utterance of the target speaker to extract his/her voice characteristics that are…

Cited by 152SourceScholar
2020

Jointly Optimal Dereverberation and Beamforming

ICASSP 2020accepted

We previously proposed an optimal (in the maximum likelihood sense) convolutional beamformer that can perform simultaneous denoising and dereverberation, and showed its superiority over the widely used cascade of a Weighted Prediction Error (WPE) dereverberation filter and a conventional Minimum-Pow…

Cited by 0SourceScholar
2020

Tackling Real Noisy Reverberant Meetings with All-Neural Source Separation, Counting, and Diarization System

ICASSP 2020accepted

Automatic meeting analysis is an essential fundamental technology required to let, e.g. smart devices follow and respond to our conversations. To achieve an optimal automatic meeting analysis, we previously proposed an all-neural approach that jointly solves source separation, speaker diarization an…

Cited by 0SourceScholar
2019

A Unified Framework for Neural Speech Separation and Extraction

ICASSP 2019accepted

The development of deep learning techniques has triggered the active investigation of neural network-based speech enhancement approaches. In particular, single-channel blind (uninformed) speech separation and speaker-aware (informed) speech extraction have received increased interest. Blind speech s…

Cited by 21SourceScholar
2019

All-neural Online Source Separation, Counting, and Diarization for Meeting Analysis

ICASSP 2019accepted

Automatic meeting analysis comprises the tasks of speaker counting, speaker diarization, and the separation of overlapped speech, followed by automatic speech recognition. This all has to be carried out on arbitrarily long sessions and, ideally, in an online or block-online manner. While significant…

Cited by 0SourceScholar
2019

Compact Network for Speakerbeam Target Speaker Extraction

ICASSP 2019accepted

Speech separation that separates a mixture of speech signals into each of its sources has been an active research topic for a long time and has seen recent progress with the advent of deep learning. A related problem is target speaker extraction, i.e. extraction of only speech of a target speaker ou…

Cited by 0SourceScholar
2019

Estimation of Sampling Frequency Mismatch between Distributed Asynchronous Microphones under Existence of Source Movements with Stationary Time Periods Detection

ICASSP 2019accepted

In this paper, we propose a method of estimating the sampling frequency mismatch among asynchronous recording devices, even when the sources sometimes move. For a spatially stationary source, there is a method of estimating the sampling frequency mismatch, which appears in the drift of the time diff…

Cited by 14SourceScholar
2019

Joint Optimization of Neural Network-based WPE Dereverberation and Acoustic Model for Robust Online ASR

ICASSP 2019accepted

Signal dereverberation using the Weighted Prediction Error (WPE) method has been proven to be an effective means to raise the accuracy of far-field speech recognition. First proposed as an iterative algorithm, follow-up works have reformulated it as a recursive least squares algorithm and therefore…

Cited by 0SourceScholar
2019

Mask-based MVDR Beamformer for Noisy Multisource Environments: Introduction of Time-varying Spatial Covariance Model

ICASSP 2019accepted

This paper proposes a method for designing a time-varying minimum variance distortionless response (MVDR) beamformer using time-frequency masks, with the aim of improving speech enhancement in noisy multi-speaker environments. A key to successful beamforming is to estimate accurately a time-varying…

Cited by 0SourceScholar
2018

Dual Frequency- and Block-Permutation Alignment for Deep Learning Based Block-Online Blind Source Separation

ICASSP 2018accepted

Deep attractor networks (DANs) are a recently introduced method to blindly separate sources from spectral features of a monaural recording using bidirectional long short-term memory networks (BLSTMs). Due to the nature of BLSTMs, this is inherently not online-ready and resorting to operating on bloc…

Cited by 0SourceScholar
2018

Frame-by-Frame Closed-Form Update for Mask-Based Adaptive MVDR Beamforming

ICASSP 2018accepted

Beamforming approaches using time-frequency masks have recently been investigated and have shown promising results for noise robust automatic speech recognition (ASR) in many tasks. The time-frequency masks are estimated to compute the spatial statistics of target speech and noise signals, and then…

Cited by 0SourceScholar
2018

Listening to Each Speaker One by One with Recurrent Selective Hearing Networks

ICASSP 2018accepted

Deep learning-based single-channel source separation algorithms are currently being actively investigated. Among them, Deep Clustering (DC) and Deep Attractor Networks (DANs) have made it possible to separate an arbitrary number of speakers. In particular, they cleverly combine a neural network and…

Cited by 0SourceScholar
2018

Meeting Recognition with Asynchronous Distributed Microphone Array Using Block-Wise Refinement of Mask-Based MVDR Beamformer

ICASSP 2018accepted

This paper addresses a front-end system for speech recognition of spontaneous conversational speech signals that are recorded with asynchronous distributed microphones such as smartphones. In our previous work, we proposed combining blind synchronization and a state-of-the-art microphone array speec…

Cited by 0SourceScholar
2018

Optimization of Speaker-Aware Multichannel Speech Extraction with ASR Criterion

ICASSP 2018accepted

This paper addresses the problem of recognizing speech corrupted by overlapping speakers in a multichannel setting. To extract a target speaker from the mixture, we use a neural network based beamformer which uses masks estimated by a neural network to compute statistically optimal spatial filters.…

Cited by 0SourceScholar
2018

Single Channel Target Speaker Extraction and Recognition with Speaker Beam

ICASSP 2018accepted

This paper addresses the problem of single channel speech recognition of a target speaker in a mixture of speech signals. We propose to exploit auxiliary speaker information provided by an adaptation utterance from the target speaker to extract and recognize only that speaker. Using such auxiliary i…

Cited by 0SourceScholar
2017

Cumulative moving averaged bottleneck speaker vectors for online speaker adaptation of CNN-based acoustic models

ICASSP 2017accepted

Adapting acoustic models to speakers have shown to greatly improve performance for many tasks. Among the adaptation approaches, exploiting auxiliary features characterizing speakers or environments has received great attention because they allow rapid adaptation, i.e. adaptation with limited amount…

Cited by 0SourceScholar
2017

Deep mixture density network for statistical model-based feature enhancement

ICASSP 2017accepted

We propose a novel framework designed to extend conventional deep neural network (DNN)-based feature enhancement approaches. In general, the conventional DNN-based feature enhancement framework aims to map input noisy observation to clean speech or a binary/ soft mask in a deterministic way, assumin…

Cited by 13SourceScholar
2017

Integrating DNN-based and spatial clustering-based mask estimation for robust MVDR beamforming

ICASSP 2017accepted

Recently, time-frequency mask-based beamforming has been extensively studied as the frontend of deep neural network (DNN) based automatic speech recognition (ASR) in noisy environments. Two mask estimation approaches have been separately developed for this beamforming method, namely the the DNN-base…

Cited by 0SourceScholar
2017

Online environmental adaptation of CNN-based acoustic models using spatial diffuseness features

ICASSP 2017accepted

We propose a new concept for adapting CNN-based acoustic models using spatial diffuseness features as auxiliary information about the acoustic environment: the spatial diffuseness features are simultaneously employed as acoustic-model input features and to estimate environmental cues for context ada…

Cited by 0SourceScholar
2017

Unsupervised utterance-wise beamformer estimation with speech recognition-level criterion

ICASSP 2017accepted

In this paper, we perform beamforming with a speech recognition-level criterion. A beamformer is usually designed by optimizing signal-level criteria, e.g., by minimizing the beamformer output covariance or by maximizing the signal-to-noise ratio (SNR). Such signal-level criteria do not always guara…

Cited by 0SourceScholar
2016

Context adaptive deep neural networks for fast acoustic model adaptation in noisy conditions

ICASSP 2016accepted

Deep neural network (DNN) based acoustic models have greatly improved the performance of automatic speech recognition (ASR) for various tasks. Further performance improvements have been reported when making DNNs aware of the acoustic context (e.g. speaker or environment) for example by adding auxili…

Cited by 36SourceScholar
2015

Context adaptive deep neural networks for fast acoustic model adaptation

ICASSP 2015accepted

Deep neural networks (DNNs) are widely used for acoustic modeling in automatic speech recognition (ASR), since they greatly outperform legacy Gaussian mixture model-based systems. However, the levels of performance achieved by current DNN-based systems remain far too low in many tasks, e.g. when the…

Cited by 0SourceScholar
2015

Modeling inter-node acoustic dependencies with Restricted Boltzmann Machine for distributed microphone array based BSS

ICASSP 2015accepted

An accurate estimation of a source activity information is essential for many speech enhancement algorithms including blind source separation (BSS). In this paper, we propose a novel BSS method that accurately models and estimates the source activity in distributed microphone array (DMA) scenarios.…

Cited by 0SourceScholar