← Search

Jyh-Shing Roger Jang

10 accepted papers

2026

How Does Instrumental Music Help SingFake Detection?

ICASSP 2026poster

Although many models exist to detect singing voice deepfakes (SingFake), how these models operate, particularly with instrumental accompaniment, is unclear. We investigate how instrumental music affects SingFake detection from two perspectives. To investigate the behavioral effect, we test different…

Cited by 0SourcePDFScholar
2025

Similarity-based Accent Recognition with Continuous and Discrete Self-supervised Speech Representations

ICASSP 2025accepted

The primary challenge in accent recognition lies in data scarcity due to the high diversity of accents, which make the collection of large-scale training data for each accent almost impossible in practice. To overcome this challenge, we propose a simple solution that leverages both continuous and di…

Cited by 1SourceScholar
2024

MIR-MLPop: A Multilingual Pop Music Dataset with Time-Aligned Lyrics and Audio

ICASSP 2024accepted

We introduce MIR-MLPop, a publicly available multilingual pop music dataset designed for automatic lyrics transcription and lyrics alignment in polyphonic music. The dataset comprises 90 pop music tracks in Mandarin, Cantonese, and Taiwanese Hokkien, with manually annotated time-aligned lyrics with…

Cited by 0SourceScholar
2024

Multimodal Transformer Distillation for Audio-Visual Synchronization

ICASSP 2024accepted

Audio-visual synchronization aims to determine whether the mouth movements and speech in the video are synchronized. VocaLiST reaches state-of-the-art performance by incorporating multimodal Transformers to model audio-visual interact information. However, it requires high computing resources, makin…

Cited by 0SourceScholar
2022

Towards Automatic Transcription of Polyphonic Electric Guitar Music: A New Dataset and a Multi-Loss Transformer Model

ICASSP 2022accepted

In this paper, we propose a new dataset named EGDB, that contains transcriptions of the electric guitar performance of 240 tablatures rendered with different tones. Moreover, we benchmark the performance of two well-known transcription models proposed originally for the piano on this dataset, along…

Cited by 28SourceScholar
2021

On the Preparation and Validation of a Large-Scale Dataset of Singing Transcription

ICASSP 2021accepted

This paper proposes a large-scale dataset for singing transcription, along with some methods for fine-tuning and validating its contents. The dataset is named MIR-ST500, which consists of more than 160,000 notes from 500 pop songs. To create this large-scale dataset, we set some labeling criteria an…

Cited by 39SourceScholar
2019

Learning to Match Transient Sound Events Using Attentional Similarity for Few-shot Sound Recognition

ICASSP 2019accepted

In this paper, we introduce a novel attentional similarity module for the problem of few-shot sound recognition. Given a few examples of an unseen sound event, a classifier must be quickly adapted to recognize the new sound event without much fine-tuning. The proposed attentional similarity module c…

Cited by 0SourceScholar
2016

An efficient method for polyphonic audio-to-score alignment using onset detection and constant Q transform

ICASSP 2016accepted

This paper proposes an innovative method that aligns a polyphonic audio recording of music to its corresponding symbolic score. In the first step, we perform onset detection and then apply constant Q transform around each onset. A similarity matrix is computed by using a scoring function which evalu…

Cited by 0SourceScholar
2015

Vocal activity informed singing voice separation with the iKala dataset

ICASSP 2015accepted

A new algorithm is proposed for robust principal component analysis with predefined sparsity patterns. The algorithm is then applied to separate the singing voice from the instrumental accompaniment using vocal activity information. To evaluate its performance, we construct a new publicly available…

Cited by 0SourceScholar