← Search

Shoko Araki

37 accepted papers

2026

Joint Enhancement and Classification using Coupled Diffusion Models of Signals and Logits

ICML 2026poster

Robust classification in noisy environments remains a fundamental challenge in machine learning. Standard approaches typically treat signal enhancement and classification as separate, sequential stages: first enhancing the signal and then applying a classifier. This approach fails to leverage the se…

Cited by 0SourceScholar
2025

30+ Years of Source Separation Research: Achievements and Future Challenges

ICASSP 2025accepted

Source separation (SS) of acoustic signals is a research field that emerged in the mid-1990s and has flourished ever since. On the occasion of ICASSP’s 50<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">th</sup> anniversary, we review the major contribut…

Cited by 0SourceScholar
2025

A Hybrid Probabilistic-Deterministic Model Recursively Enhancing Speech

ICASSP 2025accepted

This paper introduces Probabilistic-Deterministic Recursive Enhancement (PDRE), an innovative iterative Speech Enhancement (SE) approach that integrates probabilistic and deterministic methodologies. Recent advancements in diffusion models have demonstrated the exceptional effectiveness of probabili…

Cited by 0SourceScholar
2025

Mamba-based Segmentation Model for Speaker Diarization

ICASSP 2025accepted

Mamba is a newly proposed architecture that behaves like a recurrent neural network (RNN) with attention-like capabilities. These properties are promising for speaker diarization, as attention-based models have unsuitable memory requirements for long-form audio, and traditional RNN capabilities are…

Cited by 0SourceScholar
2025

SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model

ICASSP 2025accepted

Target sound extraction (TSE) consists of isolating a desired sound from a mixture of arbitrary sounds using clues to identify it. A TSE system requires solving two problems at once, identifying the target source and extracting the target signal from the mixture. For increased practicability, the sa…

Cited by 0SourceScholar
2025

TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models

ICASSP 2025accepted

Self-supervised learning (SSL) models have significantly advanced speech processing tasks, and several benchmarks have been proposed to validate their effectiveness. However, previous benchmarks have primarily focused on single-speaker scenarios, with less exploration of target-speaker tasks in nois…

Cited by 0SourceScholar
2024

How Does End-To-End Speech Recognition Training Impact Speech Enhancement Artifacts?

ICASSP 2024accepted

Jointly training a speech enhancement (SE) front-end and an automatic speech recognition (ASR) back-end has been investigated as a way to mitigate the influence of processing distortion generated by single-channel SE on ASR. In this paper, we investigate the effect of such joint training on the sign…

Cited by 0SourceScholar
2024

Neural Network-Based Virtual Microphone Estimation with Virtual Microphone and Beamformer-Level Multi-Task Loss

ICASSP 2024accepted

Array processing performance depends on the number of microphones available. Virtual microphone estimation (VME) has been proposed to increase the number of microphone signals artificially. Neural network-based VME (NN-VME) trains an NN with a VM-level loss to predict a signal at a microphone locati…

Cited by 0SourceScholar
2024

Online Target Sound Extraction with Knowledge Distillation from Partially Non-Causal Teacher

ICASSP 2024accepted

Target Sound Extraction (TSE) is a technique for extracting sound events belonging to a target sound class in a mixture using a Deep Neural Network (DNN). Offline TSE that uses non-causal models has achieved high extraction performance. However, many applications require online processing. Simply co…

Cited by 0SourceScholar
2024

Target Speech Extraction with Pre-Trained Self-Supervised Learning Models

ICASSP 2024accepted

Pre-trained self-supervised learning (SSL) models have achieved remarkable success in various speech tasks. However, their potential in target speech extraction (TSE) has not been fully exploited. TSE aims to extract the speech of a target speaker in a mixture guided by enrollment utterances. We exp…

Cited by 0SourceScholar
2023

Fast Online Source Steering Algorithm for Tracking Single Moving Source Using Online Independent Vector Analysis

ICASSP 2023accepted

We address the problem of separating moving sources using online independent vector analysis (IVA). To solve this problem, researchers have extended the iterative projection (IP) and iterative source steering (ISS) algorithms developed for batch auxiliary-function-based IVA (AuxIVA) to online scenar…

Cited by 0SourceScholar
2022

Lattice Rescoring Based on Large Ensemble of Complementary Neural Language Models

ICASSP 2022accepted

We investigate the effectiveness of using a large ensemble of advanced neural language models (NLMs) for lattice rescoring on automatic speech recognition (ASR) hypotheses. Previous studies have reported the effectiveness of combining a small number of NLMs. In contrast, in this study, we combine up…

Cited by 0SourceScholar
2021

Blind and Neural Network-Guided Convolutional Beamformer for Joint Denoising, Dereverberation, and Source Separation

ICASSP 2021accepted

This paper proposes an approach for optimizing a Convolutional BeamFormer (CBF) that can jointly perform denoising (DN), dereverberation (DR), and source separation (SS). First, we develop a blind CBF optimization algorithm that requires no prior information on the sources or the room acoustics, by…

Cited by 0SourceScholar
2021

Data Fusion for Audiovisual Speaker Localization: Extending Dynamic Stream Weights to the Spatial Domain

ICASSP 2021accepted

Estimating the positions of multiple speakers can be helpful for tasks like automatic speech recognition or speaker diarization. Both applications benefit from a known speaker position when, for instance, applying beamforming or assigning unique speaker identities. Recently, several approaches utili…

Cited by 0SourceScholar
2021

Low Latency Online Blind Source Separation Based on Joint Optimization with Blind Dereverberation

ICASSP 2021accepted

This paper presents a new low-latency online blind source separation (BSS) algorithm. Although algorithmic delay of a frequency domain online BSS can be reduced simply by shortening the short-time Fourier transform (STFT) frame length, it degrades the source separation performance in the presence of…

Cited by 0SourceScholar
2021

Neural Network-Based Virtual Microphone Estimator

ICASSP 2021accepted

Developing microphone array technologies for a small number of microphones is important due to the constraints of many devices. One direction to address this situation consists of virtually augmenting the number of microphone signals, e.g., based on several physical model assumptions. However, such…

Cited by 0SourceScholar
2020

A Dynamic Stream Weight Backprop Kalman Filter for Audiovisual Speaker Tracking

ICASSP 2020accepted

Audiovisual speaker tracking is an application that has been tackled by a wide range of classical approaches based on Gaussian filters, most notably the well-known Kalman filter. Recently, a specific Kalman filter implementation was proposed for this task, which incorporated dynamic stream weights t…

Cited by 0SourceScholar
2020

A Frequency-Domain BSS Method Based on ℓ1 Norm, Unitary Constraint, and Cayley Transform

ICASSP 2020accepted

We propose a frequency-domain blind source separation method that uses (a) the ℓ <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sub> norm of orthonormal vectors of estimated source signals as a sparsity measure and (b) Cayley transform for optimizin…

Cited by 0SourceScholar
2020

Beam-TasNet: Time-domain Audio Separation Network Meets Frequency-domain Beamformer

ICASSP 2020accepted

Recent studies have shown that acoustic beamforming using a microphone array plays an important role in the construction of high-performance automatic speech recognition (ASR) systems, especially for noisy and overlapping speech conditions. In parallel with the success of multichannel beamforming fo…

Cited by 0SourceScholar
2020

DNN-supported Mask-based Convolutional Beamforming for Simultaneous Denoising, Dereverberation, and Source Separation

ICASSP 2020accepted

In this article, we investigate an integrated mask-based convolutional beamforming method for performing simultaneous denoising, dereverberation, and source separation. Conventionally, it is difficult for neural network-supported mask-based source separation to perform denoising and dereverberation…

Cited by 0SourceScholar
2020

Improving Speaker Discrimination of Target Speech Extraction With Time-Domain Speakerbeam

ICASSP 2020accepted

Target speech extraction, which extracts a single target source in a mixture given clues about the target speaker, has attracted increasing attention. We have recently proposed SpeakerBeam, which exploits an adaptation utterance of the target speaker to extract his/her voice characteristics that are…

Cited by 0SourceScholar
2020

Tackling Real Noisy Reverberant Meetings with All-Neural Source Separation, Counting, and Diarization System

ICASSP 2020accepted

Automatic meeting analysis is an essential fundamental technology required to let, e.g. smart devices follow and respond to our conversations. To achieve an optimal automatic meeting analysis, we previously proposed an all-neural approach that jointly solves source separation, speaker diarization an…

Cited by 0SourceScholar
2019

All-neural Online Source Separation, Counting, and Diarization for Meeting Analysis

ICASSP 2019accepted

Automatic meeting analysis comprises the tasks of speaker counting, speaker diarization, and the separation of overlapped speech, followed by automatic speech recognition. This all has to be carried out on arbitrarily long sessions and, ideally, in an online or block-online manner. While significant…

Cited by 0SourceScholar
2019

Compact Network for Speakerbeam Target Speaker Extraction

ICASSP 2019accepted

Speech separation that separates a mixture of speech signals into each of its sources has been an active research topic for a long time and has seen recent progress with the advent of deep learning. A related problem is target speaker extraction, i.e. extraction of only speech of a target speaker ou…

Cited by 0SourceScholar
2019

Estimation of Sampling Frequency Mismatch between Distributed Asynchronous Microphones under Existence of Source Movements with Stationary Time Periods Detection

ICASSP 2019accepted

In this paper, we propose a method of estimating the sampling frequency mismatch among asynchronous recording devices, even when the sources sometimes move. For a spatially stationary source, there is a method of estimating the sampling frequency mismatch, which appears in the drift of the time diff…

Cited by 0SourceScholar
2019

Mask-based MVDR Beamformer for Noisy Multisource Environments: Introduction of Time-varying Spatial Covariance Model

ICASSP 2019accepted

This paper proposes a method for designing a time-varying minimum variance distortionless response (MVDR) beamformer using time-frequency masks, with the aim of improving speech enhancement in noisy multi-speaker environments. A key to successful beamforming is to estimate accurately a time-varying…

Cited by 0SourceScholar
2018

Maximum-Likelihood Online Speaker Diarization in Noisy Meetings Based on Categorical Mixture Model and Probabilistic Spatial Dictionary

ICASSP 2018accepted

In this paper, we propose a maximum-likelihood online diarization method based on a probabilistic spatial dictionary. This dictionary consists of the given probability distribution of spatial features for each possible direction of arrival (DOA) of source signals. Recently, we have developed an onli…

Cited by 0SourceScholar
2018

Meeting Recognition with Asynchronous Distributed Microphone Array Using Block-Wise Refinement of Mask-Based MVDR Beamformer

ICASSP 2018accepted

This paper addresses a front-end system for speech recognition of spontaneous conversational speech signals that are recorded with asynchronous distributed microphones such as smartphones. In our previous work, we proposed combining blind synchronization and a state-of-the-art microphone array speec…

Cited by 0SourceScholar
2018

Permutation-Free Cgmm: Complex Gaussian Mixture Model with Inverse Wishart Mixture Model Based Spatial Prior for Permutation-Free Source Separation and Source Counting

ICASSP 2018accepted

Here we propose a permutation-free cGMM (PF-cGMM), a new probabilistic model of observed mixtures, which can resolve permutation ambiguity between frequency bins, and is applicable even when the number of sources is unknown. A recently proposed complex Gaussian mixture model (cGMM) is highly effecti…

Cited by 0SourceScholar
2017

Integrating DNN-based and spatial clustering-based mask estimation for robust MVDR beamforming

ICASSP 2017accepted

Recently, time-frequency mask-based beamforming has been extensively studied as the frontend of deep neural network (DNN) based automatic speech recognition (ASR) in noisy environments. Two mask estimation approaches have been separately developed for this beamforming method, namely the the DNN-base…

Cited by 0SourceScholar
2017

Probabilistic spatial dictionary based online adaptive beamforming for meeting recognition in noisy and reverberant environments

ICASSP 2017accepted

Here we propose online adaptive beamforming for automatic speech recognition (ASR) in meetings in noisy, reverberant environments. The proposed method is based on recently developed mask-based beamforming, in which accurate mask estimation and diarization are paramount. Real-world experiments have s…

Cited by 0SourceScholar
2016

A generative-discriminative hybrid approach to multi-channel noise reduction for robust automatic speech recognition

ICASSP 2016accepted

In the recent years, discriminative models have become a very attractive utility and gained a lot of attention in the speech research community, encompassing both front and back-end methods, thanks to their prominent discriminative power and the availability of improved training strategies. When it…

Cited by 0SourceScholar
2016

Modeling audio directional statistics using a complex bingham mixture model for blind source extraction from diffuse noise

ICASSP 2016accepted

Mask estimation is a central task in blind signal processing including source separation, denoising, and multi-source localization. In this paper, we define a complex Bingham mixture model (cBMM), and propose it as a model of directional statistics for mask estimation. The complex Bingham distributi…

Cited by 0SourceScholar
2016

Real-time integration of statistical model-based speech enhancement with unsupervised noise PSD estimation using microphone array

ICASSP 2016accepted

We propose a technique of multi-channel speech enhancement based on integration of beamforming and statistical model-based speech enhancement to clearly extract the target speech, even in very noisy environments. Conventional microphone array-based techniques estimate speech and noise power spectral…

Cited by 0SourceScholar
2016

Spatial correlation model based observation vector clustering and MVDR beamforming for meeting recognition

ICASSP 2016accepted

This paper addresses a minimum variance distortionless response (MVDR) beamforming based speech enhancement approach for meeting speech recognition. In a meeting situation, speaker overlaps and noise signals are not negligible. To handle these issues, we employ MVDR beamforming, where accurate estim…

Cited by 0SourceScholar
2015

Exploring multi-channel features for denoising-autoencoder-based speech enhancement

ICASSP 2015accepted

This paper investigates a multi-channel denoising autoencoder (DAE)-based speech enhancement approach. In recent years, deep neural network (DNN)-based monaural speech enhancement and robust automatic speech recognition (ASR) approaches have attracted much attention due to their high performance. Al…

Cited by 0SourceScholar