← Search

Yoshiaki Bando

17 accepted papers

2025

Formula-Supervised Sound Event Detection: Pre-Training Without Real Data

ICASSP 2025accepted

In this paper, we propose a novel formula-driven supervised learning (FDSL) framework for pre-training an environmental sound analysis model by leveraging acoustic signals parametrically synthesized through formula-driven methods. Specifically, we outline detailed procedures and evaluate their effec…

Cited by 0SourceScholar
2025

Investigation of Spatial Self-Supervised Learning and Its Application to Target Speaker Speech Recognition

ICASSP 2025accepted

In this paper, we investigate spatial self-supervised learning for target speaker speech recognition. Neural separation models can be trained in a self-supervised manner by using only multichannel mixture signals. Such a framework is typically based on a physics-informed generative model, widely stu…

Cited by 0SourceScholar
2025

Source-Aware Spatial Self-Supervision for Sound Event Localization and Detection

ICASSP 2025accepted

This paper presents a source-aware spatial self-supervised learning (SSL) method for sound event localization and detection (SELD) based on neural blind source separation (BSS). SELD involves estimating both the temporal activations of sound classes and their directions of arrival (DOAs) from multic…

Cited by 0SourceScholar
2022

Direction-Aware Adaptive Online Neural Speech Enhancement with an Augmented Reality Headset in Real Noisy Conversational Environments

IROS 2022poster

This paper describes the practical response- and performance-aware development of online speech enhancement for an augmented reality (AR) headset that helps a user understand conversations made in real noisy echoic environments (e.g., cocktail party). One may use a state-of-the-art blind source sepa…

Cited by 6SourceScholar
2022

Flow-Based Fast Multichannel Nonnegative Matrix Factorization for Blind Source Separation

ICASSP 2022accepted

This paper describes a blind source separation method for multichannel audio signals, called NF-FastMNMF, based on the integration of the normalizing flow (NF) into the multichannel nonnegative matrix factorization with jointly-diagonalizable spatial covariance matrices, a.k.a. FastMNMF. Whereas the…

Cited by 0SourceScholar
2021

Autoregressive Fast Multichannel Nonnegative Matrix Factorization For Joint Blind Source Separation And Dereverberation

ICASSP 2021accepted

This paper describes a joint blind source separation and dereverberation method that works adaptively and efficiently in a reverberant noisy environment. The modern approach to blind source separation (BSS) is to formulate a probabilistic model of multichannel mixture signals that consists of a sour…

Cited by 0SourceScholar
2021

Pitch-Timbre Disentanglement Of Musical Instrument Sounds Based On Vae-Based Metric Learning

ICASSP 2021accepted

This paper describes a representation learning method for disentangling an arbitrary musical instrument sound into latent pitch and timbre representations. Although such pitch-timbre disentanglement has been achieved with a variational autoencoder (VAE), especially for a predefined set of musical in…

Cited by 0SourceScholar
2020

Self-supervised Neural Audio-Visual Sound Source Localization via Probabilistic Spatial Modeling

IROS 2020poster

Detecting sound source objects within visual observation is important for autonomous robots to comprehend surrounding environments. Since sounding objects have a large variety with different appearances in our living environments, labeling all sounding objects is impossible in practice. This calls f…

Cited by 21SourceScholar
2018

Statistical Speech Enhancement Based on Probabilistic Integration of Variational Autoencoder and Non-Negative Matrix Factorization

ICASSP 2018accepted

This paper presents a statistical method of single-channel speech enhancement that uses a variational autoencoder (VAE) as a prior distribution on clean speech. A standard approach to speech enhancement is to train a deep neural network (DNN) to take noisy speech as input and output clean speech. Al…

Cited by 0SourceScholar
2018

Unsupervised Beamforming Based on Multichannel Nonnegative Matrix Factorization for Noisy Speech Recognition

ICASSP 2018accepted

This paper presents unsupervised multichannel speech enhancement for noisy speech recognition. Time-frequency (TF) mask estimation has actively been studied for estimating the steering vectors and spatial covariance matrices of speech and noise used for beamforming. The state-of-the-art approach to…

Cited by 0SourceScholar
2017

Bayesian multichannel nonnegative matrix factorization for audio source separation and localization

ICASSP 2017accepted

This paper presents a Bayesian extension of multichannel nonnegative matrix factorization (MNMF) that decomposes the complex spectrograms of mixture signals recorded by a microphone array into basis spectra, their temporal activations, and the spatial correlation matrices of sources (directions) in…

Cited by 0SourceScholar
2017

Development of microphone-array-embedded UAV for search and rescue task

IROS 2017poster

This paper addresses online outdoor sound source localization using a microphone array embedded in an unmanned aerial vehicle (UAV). In addition to sound source localization, sound source enhancement and robust communication method are also described. This system is one instance of deployment of our…

Cited by 57SourceScholar
2016

Online simultaneous localization and mapping of multiple sound sources and asynchronous microphone arrays

IROS 2016poster

This paper presents an online method of simultaneous localization and mapping (SLAM) for estimating the positions of multiple moving sound sources and stationary robots and synchronizing microphone arrays attached to those robots. Since each robot with a microphone array can solely estimate the dire…

Cited by 18SourceScholar
2015

Audio-visual beat tracking based on a state-space model for a music robot dancing with humans

IROS 2015poster

This paper presents an audio-visual beat-tracking method for an entertainment robot that can dance in synchronization with music and human dancers. Conventional music robots have focused on either music audio signals or dancing movements of humans for detecting and predicting beat times in real time…

Cited by 17SourceScholar
2015

Challenges in deploying a microphone array to localize and separate sound sources in real auditory scenes

ICASSP 2015accepted

Analyzing the auditory scene of real environments is challenging partly because an unknown number and type of sound sources are observed at the same time and partly because these sounds are observed on a significantly different sound pressure level at the microphone. These are difficult problems eve…

Cited by 0SourceScholar
2015

Microphone-accelerometer based 3D posture estimation for a hose-shaped rescue robot

IROS 2015poster

3D posture estimation for a hose-shaped robot is critical in rescue activities due to complex physical environments. Conventional sound-based posture estimation assumes rather flat physical environments and focuses only on 2D, resulting in poor performance in real world environments with rubble. Thi…

Cited by 16SourceScholar
2015

Optimizing the layout of multiple mobile robots for cooperative sound source separation

IROS 2015poster

This paper presents a novel active audition method that enables multiple mobile robots to move to optimal positions for improving the performance of sound source separation. A main advantage of our distributed system is that each robot has its own microphone array and all mobile robots can collabora…

Cited by 7SourceScholar