← Search

Ryosuke Sawata

7 accepted papers

2024

Hearing Anything Anywhere

CVPR 2024poster

Recent years have seen immense progress in 3D computer vision and computer graphics with emerging tools that can virtualize real-world 3D environments for numerous Mixed Reality (XR) applications. However alongside immersive visual experiences immersive auditory experiences are equally vital to our…

2023

Diffroll: Diffusion-Based Generative Music Transcription with Unsupervised Pretraining Capability

ICASSP 2023accepted

In this paper we propose a novel generative approach, DiffRoll, to tackle automatic music transcription (AMT). Instead of treating AMT as a discriminative task in which the model is trained to convert spectrograms into piano rolls, we think of it as a conditional generative task where we train our m…

Cited by 0SourceScholar
2022

Improving Character Error Rate is Not Equal to Having Clean Speech: Speech Enhancement for ASR Systems with Black-Box Acoustic Models

ICASSP 2022accepted

A deep neural network (DNN)-based speech enhancement (SE) aiming to maximize the performance of an automatic speech recognition (ASR) system is proposed in this paper. In order to optimize the DNN-based SE model in terms of the character error rate (CER), which is one of the metric to evaluate the A…

Cited by 0SourceScholar
2021

All For One And One For All: Improving Music Separation By Bridging Networks

ICASSP 2021accepted

This paper proposes several improvements for music separation with deep neural networks (DNNs), namely a multi-domain loss (MDL) and two combination schemes. First, by using MDL we take advantage of the frequency and time domain representation of audio signals. Next, we utilize the relationship amon…

Cited by 0SourceScholar
2021

Human-Centered Favorite Music Classification Using EEG-Based Individual Music Preference Via Deep Time-Series CCA

ICASSP 2021accepted

A method to classify a user’s like or dislike musical pieces based on the extraction of his or her music preference is proposed in this paper. New scheme of Canonical Correlation Analysis (CCA), called Deep Time-series CCA (DTCCA), which can consider the correlation between two sets of input feature…

Cited by 0SourceScholar
2016

Novel favorite music classification using EEG-based optimal audio features selected via KDLPCCA

ICASSP 2016accepted

This paper presents a novel method of favorite music classification using EEG-based optimal audio features. To select audio features related to user's music preference, our method utilizes a relationship between EEG features obtained from the user's EEG signals during listening to music and their co…

Cited by 0SourceScholar