← Search

Jisi Zhang

4 accepted papers

2025

Diffusion based Text-to-Music Generation with Global and Local Text based Conditioning

ICASSP 2025accepted

Diffusion based Text-To-Music (TTM) models generate music corresponding to text descriptions. Typically UNet based diffusion models condition on text embeddings generated from a pre-trained large language model or from a cross-modality audio-language representation model. This work proposes a diffus…

Cited by 0SourceScholar
2025

Retrieval Augmented Generation based context discovery for ASR

EMNLP 2025

This work investigates retrieval augmented generation as an efficient strategy for automatic context discovery in context-aware Automatic Speech Recognition (ASR) system, in order to improve transcription accuracy in the presence of rare or out-of-vocabulary terms. However, identifying the right con

Cited by 0SourcePDFScholar
2021

Time-Domain Speech Extraction with Spatial Information and Multi Speaker Conditioning Mechanism

ICASSP 2021accepted

In this paper, we present a novel multi-channel speech extraction system to simultaneously extract multiple clean individual sources from a mixture in noisy and reverberant environments. The proposed method is built on an improved multi-channel time-domain speech separation network which employs spe…

Cited by 0SourceScholar
2020

On End-to-end Multi-channel Time Domain Speech Separation in Reverberant Environments

ICASSP 2020accepted

This paper introduces a new method for multi-channel time domain speech separation in reverberant environments. A fully-convolutional neural network structure has been used to directly separate speech from multiple microphone recordings, with no need of conventional spatial feature extraction. To re…

Cited by 0SourceScholar