← Search

Satoru Fukayama

7 accepted papers

2026

ADVANCED MODELING OF INTERLANGUAGE SPEECH INTELLIGIBILITY BENEFIT WITH L1-L2 MULTI-TASK LEARNING USING DIFFERENTIABLE K-MEANS FOR ACCENT-ROBUST DISCRETE TOKEN-BASED ASR

ICASSP 2026poster

Building ASR systems robust to foreign-accented speech is an important challenge in today's globalized world. A prior study explored the way to enhance the performance of phonetic token-based ASR on accented speech by reproducing the phenomenon known as interlanguage speech intelligibility benefit (…

Cited by 0SourcePDFScholar
2025

Investigation of Spatial Self-Supervised Learning and Its Application to Target Speaker Speech Recognition

ICASSP 2025accepted

In this paper, we investigate spatial self-supervised learning for target speaker speech recognition. Neural separation models can be trained in a self-supervised manner by using only multichannel mixture signals. Such a framework is typically based on a physics-informed generative model, widely stu…

Cited by 0SourceScholar
2023

jaCappella Corpus: A Japanese a Cappella Vocal Ensemble Corpus

ICASSP 2023accepted

We construct a corpus of Japanese a cappella vocal ensembles (ja-Cappella corpus) for vocal ensemble separation and synthesis. It consists of 35 copyright-cleared vocal ensemble songs and their audio recordings of individual voice parts. These songs were arranged from out-of-copyright Japanese child…

Cited by 0SourceScholar
2019

Automatic Singing Transcription Based on Encoder-decoder Recurrent Neural Networks with a Weakly-supervised Attention Mechanism

ICASSP 2019accepted

This paper describes neural singing transcription that estimates a sequence of musical notes directly from the audio signal of singing voice in an end-to-end manner without time-aligned training data. A conventional approach to singing transcription is to perform vocal F0 estimation followed by musi…

Cited by 27SourceScholar
2019

Joint Transcription of Lead, Bass, and Rhythm Guitars Based on a Factorial Hidden Semi-Markov Model

ICASSP 2019accepted

This paper describes a statistical method for estimating musical scores for lead, bass, and rhythm guitars from polyphonic audio signals of typical band-style music. To perform multi-instrument transcription involving multi-pitch detection and part assignment, it is crucial to formulate a musical la…

Cited by 0SourceScholar
2019

Transdrums: A Drum Pattern Transfer System Preserving Global Pattern Structure

ICASSP 2019accepted

This paper presents TransDrums, which is a system that transfers drum patterns from a drum-pattern-source song (D-song) to a base song (B-song) and synthesizes the audio with the substituted drum pattern. Typical drum parts consist of multiple drum patterns that are concatenated to form a structure…

Cited by 0SourceScholar