← Search

Masataka Goto

17 accepted papers

2025

A Data-Driven Method for Analyzing and Quantifying Lyrics-Dance Motion Relationships

NAACL 2025long

Dancing to music with lyrics is a popular form of expression. While it is generally accepted that there are relationships between lyrics and dance motions, previous studies have not explored these relationships. A major challenge is that the relationships between lyrics and dance motions are not con…

Cited by 0SourcePDFScholar
2025

Constrained Preferential Bayesian Optimization and Its Application in Banner Ad Design

IJCAI 2025

Preferential Bayesian optimization (PBO) is a variant of Bayesian optimization that observes relative preferences (e.g., pairwise comparisons) instead of direct objective values, making it especially suitable for human-in-the-loop scenarios. However, real-world optimization tasks often involve inequ

Cited by 0SourcePDFScholar
2024

Rail DRAGON: Long-Reach Bendable Modularized Rail Structure for Constant Observation Inside PCV

RA-L 2024

To reduce errors in the remote control of robots during decommissioning, we developed a Rail DRAGON, which enables continuous observation of the work environment. The Rail DRAGON is constructed by assembling and pushing a long rail structure inside the primary containment vessel (PCV), and then repe

Cited by 1SourceScholar
2021

Tool- and Domain-Agnostic Parameterization of Style Transfer Effects Leveraging Pretrained Perceptual Metrics

IJCAI 2021poster

Current deep learning techniques for style transfer would not be optimal for design support since their "one-shot" transfer does not fit exploratory design processes. To overcome this gap, we propose parametric transcription, which transcribes an end-to-end style transfer effect into parameter value…

Cited by 7SourcePDFScholar
2019

Automatic Singing Transcription Based on Encoder-decoder Recurrent Neural Networks with a Weakly-supervised Attention Mechanism

ICASSP 2019accepted

This paper describes neural singing transcription that estimates a sequence of musical notes directly from the audio signal of singing voice in an end-to-end manner without time-aligned training data. A conventional approach to singing transcription is to perform vocal F0 estimation followed by musi…

Cited by 27SourceScholar
2019

Joint Transcription of Lead, Bass, and Rhythm Guitars Based on a Factorial Hidden Semi-Markov Model

ICASSP 2019accepted

This paper describes a statistical method for estimating musical scores for lead, bass, and rhythm guitars from polyphonic audio signals of typical band-style music. To perform multi-instrument transcription involving multi-pitch detection and part assignment, it is crucial to formulate a musical la…

Cited by 0SourceScholar
2019

Transdrums: A Drum Pattern Transfer System Preserving Global Pattern Structure

ICASSP 2019accepted

This paper presents TransDrums, which is a system that transfers drum patterns from a drum-pattern-source song (D-song) to a base song (B-song) and synthesizes the audio with the substituted drum pattern. Typical drum parts consist of multiple drum patterns that are concatenated to form a structure…

Cited by 0SourceScholar
2019

Zero-mean Convolutional Network with Data Augmentation for Sound Level Invariant Singing Voice Separation

ICASSP 2019accepted

We address an issue of separating singing voices from polyphonic music signals regardless of sound level variance of the mixture input. Using a standard separation quality assessment tool BSS Eval 4.0, we found that the separation quality of a singing voice separation (SVS) system based on a dilatab…

Cited by 0SourceScholar
2018

Instlistener: An Expressive Parameter Estimation System Imitating Human Performances of Monophonic Musical Instruments

ICASSP 2018accepted

We present InstListener, a system that takes an expressive monophonic solo instrument performance by a human performer as the input and imitates its audio recordings by using an existing MIDI (Musical Instrument Digital Interface) synthesizer. It automatically analyzes the input and estimates, for e…

Cited by 0SourceScholar
2018

Music Structure Boundary Detection and Labelling by a Deconvolution of Path-Enhanced Self-Similarity Matrix

ICASSP 2018accepted

We propose a music structure analysis method that converts a path-enhanced self-similarity matrix (SSM) into a block-enhanced SSM using non-negative matrix factor 2-D deconvolution (NMF2D). With a non-negative constraint, the deconvolution intuitively corresponds to the repeated stripes in the path-…

Cited by 0SourceScholar
2016

An estimation method of voice timbre evaluation values using feature extraction with Gaussian mixture model based on reference singer

ICASSP 2016accepted

This paper presents an estimation method of voice timbre evaluation values for arbitrary singer's singing voices generated with a singing voice synthesis system towards the development of a singing voice retrieval system. The voice timbre evaluation values are numerical values corresponding to voice…

Cited by 0SourceScholar
2016

Student's T nonnegative matrix factorization and positive semidefinite tensor factorization for single-channel audio source separation

ICASSP 2016accepted

This paper presents a robust variant of nonnegative matrix factorization (NMF) based on complex Student's t distributions (t-NMF) for source separation of single-channel audio signals. The Itakura-Saito divergence NMF (Gaussian NMF) is justified for this purpose under an assumption that the complex…

Cited by 0SourceScholar
2015

A feedback framework for improved chord recognition based on NMF-based approximate note transcription

ICASSP 2015accepted

This paper presents a feedback framework that can improve chord recognition for music audio signals by performing approximate note transcription with Bayesian non-negative matrix factorization (NMF) using prior knowledge on chords. Although the names and note compositions of chords are intrinsically…

Cited by 0SourceScholar