← Search

Yi-Hsuan Yang

25 accepted papers

2026

How Does Instrumental Music Help SingFake Detection?

ICASSP 2026poster

Although many models exist to detect singing voice deepfakes (SingFake), how these models operate, particularly with instrumental accompaniment, is unclear. We investigate how instrumental music affects SingFake detection from two perspectives. To investigate the behavioral effect, we test different…

Cited by 0SourcePDFScholar
2026

SYNTHCLONER: SYNTHESIZER-STYLE AUDIO TRANSFER VIA FACTORIZED CODEC WITH ADSR ENVELOPE CONTROL

ICASSP 2026poster

Electronic synthesizer sounds are controlled by parameter settings that yield complex timbral characteristics and ADSR envelopes, making synthesizer-style audio transfer particularly challenging. Recent approaches to timbre transfer often rely on spectral objectives or implicit style matching, offer…

Cited by 0SourcePDFScholar
2025

DDSP Guitar Amp: Interpretable Guitar Amplifier Modeling

ICASSP 2025accepted

Neural network models for guitar amplifier emulation, while being effective, often demand high computational cost and lack interpretability. Drawing ideas from physical amplifier design, this paper aims to address these issues with a new differentiable digital signal processing (DDSP)-based model, c…

Cited by 0SourceScholar
2025

METEOR: Melody-aware Texture-controllable Symbolic Music Re-Orchestration via Transformer VAE

IJCAI 2025

Re-orchestration is the process of adapting a music piece for a different set of instruments. By altering the original instrumentation, the orchestrator often modifies the musical texture while preserving a recognizable melodic line and ensures that each part is playable within the technical and exp

2025

MuseControlLite: Multifunctional Music Generation with Lightweight Conditioners

ICML 2025poster

We propose MuseControlLite, a lightweight mechanism designed to fine-tune text-to-music generation models for precise conditioning using various time-varying musical attributes and reference audio signals. The key finding is that positional embeddings, which have been seldom used by text-to-music ge…

2023

Compose & Embellish: Well-Structured Piano Performance Generation via A Two-Stage Approach

ICASSP 2023accepted

Even with strong sequence models like Transformers, generating expressive piano performances with long-range musical structures remains challenging. Meanwhile, methods to compose well-structured melodies or lead sheets (melody + chords), i.e., simpler forms of music, gained more success. Observing t…

Cited by 0SourceScholar
2022

Automatic DJ Transitions with Differentiable Audio Effects and Generative Adversarial Networks

ICASSP 2022accepted

A central task of a Disc Jockey (DJ) is to create a mixset of music with seamless transitions between adjacent tracks. In this paper, we explore a data-driven approach that uses a generative adversarial network to create the song transition by learning from real-world DJ mixes. The generator uses tw…

Cited by 0SourceScholar
2022

KaraSinger: Score-Free Singing Voice Synthesis with VQ-VAE Using Mel-Spectrograms

ICASSP 2022accepted

In this paper, we propose a novel neural network model called KaraSinger for a less-studied singing voice synthesis (SVS) task named score-free SVS, in which the prosody and melody are spontaneously decided by machine. KaraSinger comprises a vector-quantized variational autoencoder (VQ-VAE) that com…

Cited by 0SourceScholar
2022

Towards Automatic Transcription of Polyphonic Electric Guitar Music: A New Dataset and a Multi-Loss Transformer Model

ICASSP 2022accepted

In this paper, we propose a new dataset named EGDB, that contains transcriptions of the electric guitar performance of 240 tablatures rendered with different tones. Moreover, we benchmark the performance of two well-known transcription models proposed originally for the piano on this dataset, along…

Cited by 28SourceScholar
2021

Compound Word Transformer: Learning to Compose Full-Song Music over Dynamic Directed Hypergraphs

AAAI 2021technical

To apply neural sequence models such as the Transformers to music generation tasks, one has to represent a piece of music by a sequence of tokens drawn from a finite set of pre-defined vocabulary. Such a vocabulary usually involves tokens of various types. For example, to describe a musical note, on…

2021

Relative Positional Encoding for Transformers with Linear Complexity

ICML 2021oral

Recent advances in Transformer models allow for unprecedented sequence lengths, due to linear space and time complexity. In the meantime, relative positional encoding (RPE) was proposed as beneficial for classical Transformers and consists in exploiting lags instead of absolute positions for inferen…

2020

A Comparative Study of Western and Chinese Classical Music Based on Soundscape Models

ICASSP 2020accepted

Whether literally or suggestively, the concept of soundscape is alluded in both modern and ancient music. In this study, we examine whether we can analyze and compare Western and Chinese classical music based on soundscape models. We addressed this question through a comparative study. Specifically,…

Cited by 12SourceScholar
2020

Addressing The Confounds Of Accompaniments In Singer Identification

ICASSP 2020accepted

Identifying singers is an important task with many applications. However, the task remains challenging due to many issues. One major issue is related to the confounding factors from the background instrumental music that is mixed with the vocals in music production. A singer identification model may…

Cited by 18SourceScholar
2019

Learning to Match Transient Sound Events Using Attentional Similarity for Few-shot Sound Recognition

ICASSP 2019accepted

In this paper, we introduce a novel attentional similarity module for the problem of few-shot sound recognition. Given a few examples of an unseen sound event, a classifier must be quickly adapted to recognize the new sound event without much fine-tuning. The proposed attentional similarity module c…

Cited by 0SourceScholar
2017

Automatic conversion of Pop music into chiptunes for 8-bit pixel art

ICASSP 2017accepted

In this paper, we propose an audio mosaicing method that converts Pop songs into a specific music style called “chiptune,” or “8-bit music.” The goal is to reproduce Pop songs by using the sound of the chips on the old game consoles in 1980s/1990s. The proposed method goes through a procedure that f…

Cited by 0SourceScholar
2017

Deep-net fusion to classify shots in concert videos

ICASSP 2017accepted

Varying types of shots is a fundamental element in the language of film, commonly used by a visual storytelling director to convey the emotion, ideas, and art. To classify such types of shots from images, we present a new framework that facilitates the intriguing task by addressing two key issues. W…

Cited by 0SourceScholar
2017

Polyphonic piano note transcription with non-negative matrix factorization of differential spectrogram

ICASSP 2017accepted

Automatic music transcription is usually approached by using a time-frequency (TF) representation such as the short-time Fourier transform (STFT) spectrogram or the constant-Q transform. In this paper, we propose a novel yet simple TF representation that capitalizes the effectiveness of spectral flu…

Cited by 0SourceScholar
2017

Revisiting the problem of audio-based hit song prediction using convolutional neural networks

ICASSP 2017accepted

Being able to predict whether a song can be a hit has important applications in the music industry. Although it is true that the popularity of a song can be greatly affected by external factors such as social and commercial influences, to which degree audio features computed from musical signals (wh…

Cited by 0SourceScholar
2017

Weakly-supervised audio event detection using event-specific Gaussian filters and fully convolutional networks

ICASSP 2017accepted

Audio event detection aims at discovering the elements inside an audio clip. In addition to labeling the clips with the audio events, we want to find out the temporal locations of these events. However, creating clearly annotated training data can be time-consuming. Therefore, we provide a model bas…

Cited by 0SourceScholar
2015

Informed monaural source separation of music based on convolutional sparse coding

ICASSP 2015accepted

Monaural source separation is a challenging problem that has many important applications in music information retrieval. In this paper, we focus on the score-informed variant of this problem. While non-negative matrix factorization and some other approaches have been shown effective, few existing ap…

Cited by 0SourceScholar
2015

Vocal activity informed singing voice separation with the iKala dataset

ICASSP 2015accepted

A new algorithm is proposed for robust principal component analysis with predefined sparsity patterns. The algorithm is then applied to separate the singing voice from the instrumental accompaniment using vocal activity information. To evaluate its performance, we construct a new publicly available…

Cited by 0SourceScholar