← Search

Katsutoshi Itoyama

16 accepted papers

2022

Outdoor evaluation of sound source localization for drone groups using microphone arrays

IROS 2022poster

For robot and drone auditions, microphone arrays have been used for estimating sound source directions and sound source locations. By using sound source localization techniques, for example, drones can detect people calling for help even if the target person is not visible. Most sound source localiz…

Cited by 3SourceScholar
2022

Spotforming by NMF Using Multiple Microphone Arrays

IROS 2022poster

Sound source separation is a method to extract a target sound source from a mixture of various sound sources and noises. One of the typical sound source separation methods is beamforming, which can separate sound sources by direction based on the phase difference between channels from the recorded s…

Cited by 1SourceScholar
2020

Synchronization of Microphones Based on Rank Minimization of Warped Spectrum for Asynchronous Distributed Recording

IROS 2020poster

This paper describes a new method for synchronizing microphones based on spectral warping in an asynchronous microphone array. In an audio signal observed by an asynchronous microphone array, two factors are involved: the time lag caused by a mismatch of the sampling rate and offset between micropho…

Cited by 5SourceScholar
2019

Joint Transcription of Lead, Bass, and Rhythm Guitars Based on a Factorial Hidden Semi-Markov Model

ICASSP 2019accepted

This paper describes a statistical method for estimating musical scores for lead, bass, and rhythm guitars from polyphonic audio signals of typical band-style music. To perform multi-instrument transcription involving multi-pitch detection and part assignment, it is crucial to formulate a musical la…

Cited by 0SourceScholar
2018

Statistical Speech Enhancement Based on Probabilistic Integration of Variational Autoencoder and Non-Negative Matrix Factorization

ICASSP 2018accepted

This paper presents a statistical method of single-channel speech enhancement that uses a variational autoencoder (VAE) as a prior distribution on clean speech. A standard approach to speech enhancement is to train a deep neural network (DNN) to take noisy speech as input and output clean speech. Al…

Cited by 0SourceScholar
2018

Unsupervised Beamforming Based on Multichannel Nonnegative Matrix Factorization for Noisy Speech Recognition

ICASSP 2018accepted

This paper presents unsupervised multichannel speech enhancement for noisy speech recognition. Time-frequency (TF) mask estimation has actively been studied for estimating the steering vectors and spatial covariance matrices of speech and noise used for beamforming. The state-of-the-art approach to…

Cited by 0SourceScholar
2017

Bayesian multichannel nonnegative matrix factorization for audio source separation and localization

ICASSP 2017accepted

This paper presents a Bayesian extension of multichannel nonnegative matrix factorization (MNMF) that decomposes the complex spectrograms of mixture signals recorded by a microphone array into basis spectra, their temporal activations, and the spatial correlation matrices of sources (directions) in…

Cited by 0SourceScholar
2016

Online simultaneous localization and mapping of multiple sound sources and asynchronous microphone arrays

IROS 2016poster

This paper presents an online method of simultaneous localization and mapping (SLAM) for estimating the positions of multiple moving sound sources and stationary robots and synchronizing microphone arrays attached to those robots. Since each robot with a microphone array can solely estimate the dire…

Cited by 18SourceScholar
2016

Student's T nonnegative matrix factorization and positive semidefinite tensor factorization for single-channel audio source separation

ICASSP 2016accepted

This paper presents a robust variant of nonnegative matrix factorization (NMF) based on complex Student's t distributions (t-NMF) for source separation of single-channel audio signals. The Itakura-Saito divergence NMF (Gaussian NMF) is justified for this purpose under an assumption that the complex…

Cited by 0SourceScholar
2015

A feedback framework for improved chord recognition based on NMF-based approximate note transcription

ICASSP 2015accepted

This paper presents a feedback framework that can improve chord recognition for music audio signals by performing approximate note transcription with Bayesian non-negative matrix factorization (NMF) using prior knowledge on chords. Although the names and note compositions of chords are intrinsically…

Cited by 0SourceScholar
2015

Audio-visual beat tracking based on a state-space model for a music robot dancing with humans

IROS 2015poster

This paper presents an audio-visual beat-tracking method for an entertainment robot that can dance in synchronization with music and human dancers. Conventional music robots have focused on either music audio signals or dancing movements of humans for detecting and predicting beat times in real time…

Cited by 17SourceScholar
2015

Challenges in deploying a microphone array to localize and separate sound sources in real auditory scenes

ICASSP 2015accepted

Analyzing the auditory scene of real environments is challenging partly because an unknown number and type of sound sources are observed at the same time and partly because these sounds are observed on a significantly different sound pressure level at the microphone. These are difficult problems eve…

Cited by 0SourceScholar
2015

Microphone-accelerometer based 3D posture estimation for a hose-shaped rescue robot

IROS 2015poster

3D posture estimation for a hose-shaped robot is critical in rescue activities due to complex physical environments. Conventional sound-based posture estimation assumes rather flat physical environments and focuses only on 2D, resulting in poor performance in real world environments with rubble. Thi…

Cited by 16SourceScholar
2015

Optimizing the layout of multiple mobile robots for cooperative sound source separation

IROS 2015poster

This paper presents a novel active audition method that enables multiple mobile robots to move to optimal positions for improving the performance of sound source separation. A main advantage of our distributed system is that each robot has its own microphone array and all mobile robots can collabora…

Cited by 7SourceScholar
2015

Singing voice analysis and editing based on mutually dependent F0 estimation and source separation

ICASSP 2015accepted

This paper presents a novel framework that improves both vocal fundamental frequency (F0) estimation and singing voice separation by making effective use of the mutual dependency of those two tasks. A typical approach to singing voice separation is to estimate the vocal F0 contour from a target musi…

Cited by 0SourceScholar