← Search

Shoji Makino

13 accepted papers

2026

Entropy-Guided GRVQ for Ultra-Low Bitrate Neural Speech Codec

ICASSP 2026poster

Neural audio codec (NAC) is essential for reconstructing high-quality speech signals and generating discrete representations for downstream speech language models. However, ensuring accurate semantic modeling while maintaining high-fidelity reconstruction under ultra-low bitrate constraints remains…

Cited by 0SourcePDFScholar
2026

NEURAL NETWORK-BASED TIME-FREQUENCY-BIN-WISE LINEAR COMBINATION OF BEAMFORMERS FOR UNDERDETERMINED TARGET SOURCE EXTRACTION

ICASSP 2026poster

Extracting a target source from underdetermined mixtures is challenging for beamforming approaches. Recently proposed time-frequency-bin-wise switching (TFS) and linear combination (TFLC) strategies mitigate this by combining multiple beamformers in each time-frequency (TF) bin and choosing combinat…

Cited by 0SourcePDFScholar
2026

ROBUST ONLINE OVERDETERMINED INDEPENDENT VECTOR ANALYSIS BASED ON BILINEAR DECOMPOSITION

ICASSP 2026oral

Online blind source separation is essential for both speech communication and human-machine interaction. Among existing approaches, overdetermined independent vector analysis (OverIVA) delivers strong performance by exploiting the statistical independence of source signals and the orthogonality betw…

Cited by 0SourcePDFScholar
2024

A Computationally Efficient Semi-Blind Source Separation Approach for Nonlinear Echo Cancellation Based on an Element-Wise Iterative Source Steering

ICASSP 2024accepted

While the semi-blind source separation-based acoustic echo cancellation (SBSS-AEC) has received much research attention due to its promising performance during double-talk compared to the traditional adaptive algorithms, it suffers from system latency and nonlinear distortions. To circumvent these d…

Cited by 6SourceScholar
2024

Neural Network-Based Virtual Microphone Estimation with Virtual Microphone and Beamformer-Level Multi-Task Loss

ICASSP 2024accepted

Array processing performance depends on the number of microphones available. Virtual microphone estimation (VME) has been proposed to increase the number of microphone signals artificially. Neural network-based VME (NN-VME) trains an NN with a VM-level loss to predict a signal at a microphone locati…

Cited by 0SourceScholar
2024

Stereophonic Music Source Separation with Spatially-Informed Bridging Band-Split Network

ICASSP 2024accepted

Stereophonic music source separation (MSS) is a problem of extracting individual source tracks, e.g. bass, drums, vocals, from a stereo music recording. Deep neural network (DNN) based MSS systems have demonstrated great promise though spatial panning cues and time-frequency spectral structures in s…

Cited by 0SourceScholar
2024

Unrestricted Global Phase Bias-Aware Single-Channel Speech Enhancement with Conformer-Based Metric Gan

ICASSP 2024accepted

With the rapid development of neural networks in recent years, the ability of various networks to enhance the magnitude spectrum of noisy speech in the single-channel speech enhancement domain has become exceptionally outstanding. However, enhancing the phase spectrum using neural networks is often…

Cited by 0SourceScholar
2021

Low Latency Online Blind Source Separation Based on Joint Optimization with Blind Dereverberation

ICASSP 2021accepted

This paper presents a new low-latency online blind source separation (BSS) algorithm. Although algorithmic delay of a frequency domain online BSS can be reduced simply by shortening the short-time Fourier transform (STFT) frame length, it degrades the source separation performance in the presence of…

Cited by 0SourceScholar
2021

SepNet: A Deep Separation Matrix Prediction Network for Multichannel Audio Source Separation

ICASSP 2021accepted

In this paper, we propose SepNet, a deep neural network (DNN) designed to predict separation matrices from multichannel observations. One well-known approach to blind source separation (BSS) involves independent component analysis (ICA). A recently developed method called independent low-rank matrix…

Cited by 2SourceScholar
2021

Teacher-Student Learning for Low-Latency Online Speech Enhancement Using Wave-U-Net

ICASSP 2021accepted

In this paper, we propose a low-latency online extension of wave-U-net for single-channel speech enhancement, which utilizes teacher-student learning to reduce the system latency while keeping the enhancement performance high. Wave-U-net is a recently proposed end-to-end source separation method, wh…

Cited by 28SourceScholar
2019

Fast MVAE: Joint Separation and Classification of Mixed Sources Based on Multichannel Variational Autoencoder with Auxiliary Classifier

ICASSP 2019accepted

This paper proposes an alternative algorithm for the multi-channel variational autoencoder (MVAE), a recently proposed multichannel source separation approach. While MVAE is notable for its impressive source separation performance, its convergence-guaranteed optimization algorithm and the fact that…

Cited by 0SourceScholar
2019

Joint Separation and Dereverberation of Reverberant Mixtures with Multichannel Variational Autoencoder

ICASSP 2019accepted

In this paper, we deal with a multichannel source separation problem under a highly reverberant condition. The multichannel variational autoencoder (MVAE) is a recently proposed source separation method that employs the decoder distribution of a conditional VAE (CVAE) as the generative model for the…

Cited by 0SourceScholar
2019

Time-frequency-bin-wise Switching of Minimum Variance Distortionless Response Beamformer for Underdetermined Situations

ICASSP 2019accepted

In this paper, we present a speech enhancement method using two microphones in underdetermined situations. Time-frequency (TF) binary masking is a conventional method of enhancing speech in underdetermined situations by appropriately multiplying each TF component by zero or one. Extending this metho…

Cited by 0SourceScholar