← Search

Jiarui Hai

6 accepted papers

2026

Summary of The Inaugural Music Source Restoration Challenge

ICASSP 2026poster

Music Source Restoration (MSR) aims to recover original, unprocessed instrument stems from professionally mixed and degraded audio, requiring the reversal of both production effects and real-world degradations. We present the inaugural MSR Challenge, which features objective evaluation on studio-pro…

Cited by 0SourcePDFScholar
2025

SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis

ICASSP 2025accepted

In this paper, we introduce SSR-Speech, a neural codec autoregressive model designed for stable, safe, and robust zero-shot text-based speech editing and text-to-speech synthesis. SSR-Speech is built on a Transformer decoder and incorporates classifier-free guidance to enhance the stability of the g…

Cited by 0SourceScholar
2025

SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer

ICASSP 2025accepted

In this paper, we introduce SoloAudio, a novel diffusion-based generative model for target sound extraction (TSE). Our approach trains latent diffusion models on audio, replacing the previous U-Net backbone with a skip-connected Transformer that operates on latent features. SoloAudio supports both a…

Cited by 0SourceScholar
2024

DPM-TSE: A Diffusion Probabilistic Model for Target Sound Extraction

ICASSP 2024accepted

Common target sound extraction (TSE) approaches primarily relied on discriminative approaches in order to separate the target sound while minimizing interference from the unwanted sources, with varying success in separating the target from the background. This study introduces DPM-TSE, a generative…

Cited by 0SourceScholar
2024

Investigating Self-Supervised Deep Representations for EEG-Based Auditory Attention Decoding

ICASSP 2024accepted

Auditory Attention Decoding (AAD) algorithms play a crucial role in isolating desired sound sources within challenging acoustic environments directly from brain activity. Although recent research has shown promise in AAD using shallow representations such as auditory envelope and spectrogram, there…

Cited by 0SourceScholar
2022

Progressive Teacher-Student Training Framework for Music Tagging

ICASSP 2022accepted

Music tagging is the task of predicting multiple tags of a music excerpt, and plays an important role in modern music recommendation systems. To obtain superior performance, recent approaches of music tagging focus on developing sophisticated models or exploiting additional multi-modal information.…

Cited by 0SourceScholar