← Search

Kazuyoshi Yoshii

31 accepted papers

2026

A DISTRIBUTION MATCHING APPROACH TO NEURAL PIANO TRANSCRIPTION WITH OPTIMAL TRANSPORT

ICASSP 2026poster

This paper describes a novel paradigm that formalizes automatic piano transcription (APT) as an optimal transport (OT) problem, not as a frame-level multi-label binary classification problem. Our method learns to minimize the cost of transporting a predicted distribution of note events to the ground…

Cited by 0SourcePDFScholar
2026

ABC-EVAL: BENCHMARKING LARGE LANGUAGE MODELS ON SYMBOLIC MUSIC UNDERSTANDING AND INSTRUCTION FOLLOWING

ICASSP 2026poster

As large language models continue to develop, the feasibility and significance of text-based symbolic music tasks have become increasingly prominent. While symbolic music has been widely used in generation tasks, LLM capabilities in understanding and reasoning about symbolic music remain largely und…

Cited by 0SourcePDFScholar
2026

SIRUP: A DIFFUSION-BASED VIRTUAL UPMIXER OF STEERING VECTORS FOR HIGHLY-DIRECTIVE SPATIALIZATION WITH FIRST-ORDER AMBISONICS

ICASSP 2026poster

This paper presents virtual upmixing of steering vectors captured by a fewer-channel spherical microphone array. This challenge has conventionally been addressed by recovering the directions and signals of sound sources from first-order ambisonics (FOA) data, and then rendering the higher-order ambi…

Cited by 0SourcePDFScholar
2022

Difficulty-Aware Neural Band-to-Piano Score Arrangement based on Note- and Statistic-Level Criteria

ICASSP 2022accepted

This paper describes a neural music arrangement method that converts a given band score into a piano score with an elementary or advanced level. The major challenge of this task lies in its ill-posed nature, i.e., various piano arrangements are plausible for a band score. In this paper, we take a sc…

Cited by 0SourceScholar
2022

Direction-Aware Adaptive Online Neural Speech Enhancement with an Augmented Reality Headset in Real Noisy Conversational Environments

IROS 2022poster

This paper describes the practical response- and performance-aware development of online speech enhancement for an augmented reality (AR) headset that helps a user understand conversations made in real noisy echoic environments (e.g., cocktail party). One may use a state-of-the-art blind source sepa…

Cited by 6SourceScholar
2022

Flow-Based Fast Multichannel Nonnegative Matrix Factorization for Blind Source Separation

ICASSP 2022accepted

This paper describes a blind source separation method for multichannel audio signals, called NF-FastMNMF, based on the integration of the normalizing flow (NF) into the multichannel nonnegative matrix factorization with jointly-diagonalizable spatial covariance matrices, a.k.a. FastMNMF. Whereas the…

Cited by 0SourceScholar
2021

Autoregressive Fast Multichannel Nonnegative Matrix Factorization For Joint Blind Source Separation And Dereverberation

ICASSP 2021accepted

This paper describes a joint blind source separation and dereverberation method that works adaptively and efficiently in a reverberant noisy environment. The modern approach to blind source separation (BSS) is to formulate a probabilistic model of multichannel mixture signals that consists of a sour…

Cited by 0SourceScholar
2021

Pitch-Timbre Disentanglement Of Musical Instrument Sounds Based On Vae-Based Metric Learning

ICASSP 2021accepted

This paper describes a representation learning method for disentangling an arbitrary musical instrument sound into latent pitch and timbre representations. Although such pitch-timbre disentanglement has been achieved with a variational autoencoder (VAE), especially for a predefined set of musical in…

Cited by 0SourceScholar
2021

Statistical Correction of Transcribed Melody Notes Based on Probabilistic Integration of a Music Language Model and a Transcription Error Model

ICASSP 2021accepted

This paper describes a statistical post-processing method for automatic singing transcription that corrects pitch and rhythm errors included in a transcribed note sequence. Although the performance of frame-level pitch estimation has been improved drastically by deep learning techniques, note-level…

Cited by 0SourceScholar
2019

Automatic Singing Transcription Based on Encoder-decoder Recurrent Neural Networks with a Weakly-supervised Attention Mechanism

ICASSP 2019accepted

This paper describes neural singing transcription that estimates a sequence of musical notes directly from the audio signal of singing voice in an end-to-end manner without time-aligned training data. A conventional approach to singing transcription is to perform vocal F0 estimation followed by musi…

Cited by 27SourceScholar
2019

Bayesian Drum Transcription Based on Nonnegative Matrix Factor Decomposition with a Deep Score Prior

ICASSP 2019accepted

This paper describes a statistical method of automatic drum transcription that estimates a musical score of bass and snare drums and hi-hats from a drum signal separated from a popular music signal. One of the most effective approaches for this problem is to apply nonnegative matrix factor deconvolu…

Cited by 0SourceScholar
2019

Improved Metrical Alignment of Midi Performance Based on a Repetition-aware Online-adapted Grammar

ICASSP 2019accepted

This paper presents an improvement on an existing grammar-based method for metrical structure detection and alignment, a task which involves aligning a repeated tree structure with an input stream of musical notes. The previous method achieves state-of-the-art results, but performs poorly when it la…

Cited by 1SourceScholar
2019

Joint Transcription of Lead, Bass, and Rhythm Guitars Based on a Factorial Hidden Semi-Markov Model

ICASSP 2019accepted

This paper describes a statistical method for estimating musical scores for lead, bass, and rhythm guitars from polyphonic audio signals of typical band-style music. To perform multi-instrument transcription involving multi-pitch detection and part assignment, it is crucial to formulate a musical la…

Cited by 0SourceScholar
2018

An End-to-End Approach to Joint Social Signal Detection and Automatic Speech Recognition

ICASSP 2018accepted

Social signals such as laughter and fillers are often observed in natural conversation, and they play various roles in human-to-human communication. Detecting these events is useful for transcription systems to generate rich transcription and for dialogue systems to behave as we do such as synchroni…

Cited by 0SourceScholar
2018

Statistical Speech Enhancement Based on Probabilistic Integration of Variational Autoencoder and Non-Negative Matrix Factorization

ICASSP 2018accepted

This paper presents a statistical method of single-channel speech enhancement that uses a variational autoencoder (VAE) as a prior distribution on clean speech. A standard approach to speech enhancement is to train a deep neural network (DNN) to take noisy speech as input and output clean speech. Al…

Cited by 0SourceScholar
2018

Towards Complete Polyphonic Music Transcription: Integrating Multi-Pitch Detection and Rhythm Quantization

ICASSP 2018accepted

Most work on automatic transcription produces “piano roll” data with no musical interpretation of the rhythm or pitches. We present a polyphonic transcription method that converts a music audio signal into a human-readable musical score, by integrating multi-pitch detection and rhythm quantization m…

Cited by 0SourceScholar
2018

Unsupervised Beamforming Based on Multichannel Nonnegative Matrix Factorization for Noisy Speech Recognition

ICASSP 2018accepted

This paper presents unsupervised multichannel speech enhancement for noisy speech recognition. Time-frequency (TF) mask estimation has actively been studied for estimating the steering vectors and spatial covariance matrices of speech and noise used for beamforming. The state-of-the-art approach to…

Cited by 0SourceScholar
2017

Bayesian multichannel nonnegative matrix factorization for audio source separation and localization

ICASSP 2017accepted

This paper presents a Bayesian extension of multichannel nonnegative matrix factorization (MNMF) that decomposes the complex spectrograms of mixture signals recorded by a microphone array into basis spectra, their temporal activations, and the spatial correlation matrices of sources (directions) in…

Cited by 0SourceScholar
2016

Online simultaneous localization and mapping of multiple sound sources and asynchronous microphone arrays

IROS 2016poster

This paper presents an online method of simultaneous localization and mapping (SLAM) for estimating the positions of multiple moving sound sources and stationary robots and synchronizing microphone arrays attached to those robots. Since each robot with a microphone array can solely estimate the dire…

Cited by 18SourceScholar
2016

Student's T nonnegative matrix factorization and positive semidefinite tensor factorization for single-channel audio source separation

ICASSP 2016accepted

This paper presents a robust variant of nonnegative matrix factorization (NMF) based on complex Student's t distributions (t-NMF) for source separation of single-channel audio signals. The Itakura-Saito divergence NMF (Gaussian NMF) is justified for this purpose under an assumption that the complex…

Cited by 0SourceScholar
2016

Tree-structured probabilistic model of monophonic written music based on the generative theory of tonal music

ICASSP 2016accepted

This paper presents a probabilistic formulation of music language modelling based on the generative theory of tonal music (GTTM) named probabilistic GTTM (PGTTM). GTTM is a well-known music theory that describes the tree structure of written music in analogy with the phrase structure grammar of natu…

Cited by 22SourceScholar
2015

A feedback framework for improved chord recognition based on NMF-based approximate note transcription

ICASSP 2015accepted

This paper presents a feedback framework that can improve chord recognition for music audio signals by performing approximate note transcription with Bayesian non-negative matrix factorization (NMF) using prior knowledge on chords. Although the names and note compositions of chords are intrinsically…

Cited by 0SourceScholar
2015

Audio-visual beat tracking based on a state-space model for a music robot dancing with humans

IROS 2015poster

This paper presents an audio-visual beat-tracking method for an entertainment robot that can dance in synchronization with music and human dancers. Conventional music robots have focused on either music audio signals or dancing movements of humans for detecting and predicting beat times in real time…

Cited by 17SourceScholar
2015

Challenges in deploying a microphone array to localize and separate sound sources in real auditory scenes

ICASSP 2015accepted

Analyzing the auditory scene of real environments is challenging partly because an unknown number and type of sound sources are observed at the same time and partly because these sounds are observed on a significantly different sound pressure level at the microphone. These are difficult problems eve…

Cited by 0SourceScholar
2015

Microphone-accelerometer based 3D posture estimation for a hose-shaped rescue robot

IROS 2015poster

3D posture estimation for a hose-shaped robot is critical in rescue activities due to complex physical environments. Conventional sound-based posture estimation assumes rather flat physical environments and focuses only on 2D, resulting in poor performance in real world environments with rubble. Thi…

Cited by 16SourceScholar
2015

Optimizing the layout of multiple mobile robots for cooperative sound source separation

IROS 2015poster

This paper presents a novel active audition method that enables multiple mobile robots to move to optimal positions for improving the performance of sound source separation. A main advantage of our distributed system is that each robot has its own microphone array and all mobile robots can collabora…

Cited by 7SourceScholar
2015

Singing voice analysis and editing based on mutually dependent F0 estimation and source separation

ICASSP 2015accepted

This paper presents a novel framework that improves both vocal fundamental frequency (F0) estimation and singing voice separation by making effective use of the mutual dependency of those two tasks. A typical approach to singing voice separation is to estimate the vocal F0 contour from a target musi…

Cited by 0SourceScholar