← Search

Bilei Zhu

14 accepted papers

2024

ByteHum: Fast and Accurate Query-by-Humming in the Wild

ICASSP 2024accepted

Query by Humming (QBH) is a practically meaningful task, while most existing methods struggle to scale to real-life applications due to the complex preprocessing for building the database and the limited search speed. In this paper, we propose the ByteHum system, a fast and efficient humming retriev…

Cited by 0SourceScholar
2024

Joint Music and Language Attention Models for Zero-Shot Music Tagging

ICASSP 2024accepted

Music tagging is a task to predict the tags of music recordings. However, previous music tagging research primarily focuses on close-set music tagging tasks which can not be generalized to new tags. In this work, we propose a zero-shot music tagging system modeled by a joint music and language atten…

Cited by 0SourceScholar
2023

Bytecover3: Accurate Cover Song Identification On Short Queries

ICASSP 2023accepted

Deep learning based methods have become a paradigm for cover song identification (CSI) in recent years, where the ByteCover systems have achieved state-of-the-art results on all the mainstream datasets of CSI. However, with the burgeon of short videos, many real-world applications require matching s…

Cited by 0SourceScholar
2022

Bytecover2: Towards Dimensionality Reduction of Latent Embedding for Efficient Cover Song Identification

ICASSP 2022accepted

Convolutional neural network (CNN)-based methods have dominated the recent research of cover song identification (CSI). A typical example is the ByteCover system we proposed, which has achieved state-of-the-art results on all the mainstream datasets of CSI. In this paper, we propose an up-graded ver…

Cited by 0SourceScholar
2022

HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection

ICASSP 2022accepted

Audio classification is an important task of mapping audio samples into their corresponding labels. Recently, the transformer model with self-attention mechanisms has been adopted in this field. However, existing audio transformers require large GPU memories and long training time, meanwhile relying…

Cited by 0SourceScholar
2022

S3T: Self-Supervised Pre-Training with Swin Transformer For Music Classification

ICASSP 2022accepted

In this paper, we propose S3T, a self-supervised pre-training method with Swin Transformer for music classification, aiming to learn meaningful music representations from massive easily accessible unlabeled music data. S3T introduces a momentum-based paradigm, MoCo, with Swin Transformer as its feat…

Cited by 0SourceScholar
2022

Zero-Shot Audio Source Separation through Query-Based Learning from Weakly-Labeled Data

AAAI 2022technical

Deep learning techniques for separating audio into different sound sources face several challenges. Standard architectures require training separate models for different types of audio sources. Although some universal separators employ a single model to target multiple sources, they have difficulty…

2021

An Hrnet-Blstm Model With Two-Stage Training For Singing Melody Extraction

ICASSP 2021accepted

Well-labeled datasets available for melody extraction are scarce, which limits the further advancement of deep learning based methods. To overcome this problem, we propose to use a pitch refinement method to refine the semitone-level pitch sequences decoded from massive melody MIDI files to generate…

Cited by 0SourceScholar
2021

Bytecover: Cover Song Identification Via Multi-Loss Training

ICASSP 2021accepted

We present in this paper ByteCover, which is a new feature learning method for cover song identification (CSI). Byte-Cover is built based on the classical ResNet model, and two major improvements are designed to further enhance the capability of the model for CSI. In the first improvement, we introd…

Cited by 0SourceScholar
2021

Rule-Embedded Network for Audio-Visual Voice Activity Detection in Live Musical Video Streams

ICASSP 2021accepted

Detecting anchor’s voice in live musical streams is an important preprocessing step for music and speech signal processing. Existing approaches to voice activity detection (VAD) primarily rely on audio, however, audio-based VAD is difficult to effectively focus on the target voice in noisy environme…

Cited by 0SourceScholar
2021

Singing Melody Extraction from Polyphonic Music based on Spectral Correlation Modeling

ICASSP 2021accepted

Convolutional neural network (CNN) based methods have achieved state-of-the-art performance for singing melody extraction from polyphonic music. However, most of these methods focus on the learning of local features, while relationships among spectral components locating far apart are often neglecte…

Cited by 0SourceScholar
2019

Vocal Melody Extraction via DNN-based Pitch Estimation and Salience-based Pitch Refinement

ICASSP 2019accepted

Data-driven methods for melody extraction from polyphonic music generally require large amounts of labeled data for model training. However, musical data with annotations of melody fundamental frequency (F0) are rare and hard to obtain. To overcome this limitation, in this paper we propose to use me…

Cited by 0SourceScholar
2017

Fusing transcription results from polyphonic and monophonic audio for singing melody transcription in polyphonic music

ICASSP 2017accepted

This paper presents a new system for singing melody transcription from polyphonic songs. Instead of operating solely on polyphonic audio of each song to be processed (as most existing systems do), our system takes as inputs additionally multiple monophonic recordings of people singing the song. To t…

Cited by 0SourceScholar
2015

Latent time-frequency component analysis: A novel pitch-based approach for singing voice separation

ICASSP 2015accepted

Monaural singing voice separation has aroused considerable attention. Many pitch-based methods have been proposed to address this task, but generally have limited performance. The most crucial difficulties lie in the inaccurate judgment on voiced pitches and the failed recognition on unvoiced singin…

Cited by 0SourceScholar