← Search

Jianwei Yu

38 accepted papers

2026

THE SJTU X-LANCE LAB SYSTEM FOR MSR CHALLENGE 2025

ICASSP 2026poster

This report describes the system submitted to the music source restoration (MSR) Challenge 2025. Our approach is composed of sequential BS-RoFormers, each dealing with a single task including music source separation (MSS), denoise and dereverb. To support 8 instruments given in the task, we utilize…

Cited by 0SourcePDFScholar
2026

VibeVoice: Expressive Podcast Generation with Next-Token Diffusion

ICLR 2026oral

Generating long-form, multi-speaker conversational audio like podcasts poses significant challenges for traditional Text-to-Speech (TTS) systems, particularly in scalability, speaker consistency, and natural turn-taking. We present VibeVoice , a novel model designed to synthesize expressive, long-fo…

Cited by 0SourceScholar
2026

YuE: Scaling Open Foundation Models for Long-Form Music Generation

ICLR 2026poster

We tackle the task of long-form music generation, particularly the challenging \textbf{lyrics-to-song} problem, by introducing \textbf{YuE (乐)}, a family of open-source music generation foundation models. Specifically, YuE scales to trillions of tokens and generates up to five minutes of music while…

Cited by 0SourcecodeScholar
2025

CoVoMix2: Advancing Zero-Shot Dialogue Generation with Fully Non-Autoregressive Flow Matching

NeurIPS 2025poster

Generating natural-sounding, multi-speaker dialogue is crucial for applications such as podcast creation, virtual agents, and multimedia content generation. However, existing systems struggle to maintain speaker consistency, model overlapping speech, and synthesize coherent conversations efficiently…

Cited by 0SourceScholar
2025

MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

NeurIPS 2025poster

We introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,000 meticulously curated audio-question-answer triplets, collected from real-world internet videos and refined through ite…

Cited by 0SourcecodeScholar
2025

MoonCast: High-Quality Zero-Shot Podcast Generation

NeurIPS 2025poster

Recent advances in text-to-speech synthesis have achieved notable success in generating high-quality short utterances for individual speakers. However, these systems still face challenges when extending their capabilities to long, multi-speaker, and spontaneous dialogues, typical of real-world scena…

Cited by 0SourcecodeScholar
2025

Preference Alignment Improves Language Model-Based TTS

ICASSP 2025accepted

Recent advancements in text-to-speech (TTS) have shown that language model (LM)-based systems offer competitive performance to their counterparts. Further optimization can be achieved through preference alignment algorithms, which adjust LMs to align with the preferences of reward models, enhancing…

Cited by 0SourceScholar
2025

SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion Refinement

NeurIPS 2025poster

Generating music with coherent structure, harmonious instrumental and vocal elements remains a significant challenge in song generation. Existing language models and diffusion-based methods often struggle to balance global coherence with local fidelity, resulting in outputs that lack musicality or s…

Cited by 0SourcecodeScholar
2025

SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task Editor

AAAI 2025technical

The emergence of novel generative modeling paradigms, particularly audio language models, has significantly advanced the field of song generation. Although state-of-the-art models are capable of synthesizing both vocals and accompaniment tracks up to several minutes long concurrently, research about…

2024

AutoPrep: An Automatic Preprocessing Framework for In-The-Wild Speech Data

ICASSP 2024accepted

Recently, the utilization of extensive open-sourced text data has significantly advanced the performance of text-based large language models (LLMs). However, the use of in-the-wild large-scale speech data in the speech technology community remains constrained. One reason for this limitation is that…

Cited by 0SourceScholar
2024

Consistent and Relevant: Rethink the Query Embedding in General Sound Separation

ICASSP 2024accepted

The query-based audio separation usually employs specific queries to extract target sources from a mixture of audio signals. Currently, most query-based separation models need additional networks to obtain query embedding. In this way, separation model is optimized to be adapted to the distribution…

Cited by 0SourceScholar
2024

Leveraging in-the-wild Data for Effective Self-supervised Pretraining in Speaker Recognition

ICASSP 2024accepted

Current speaker recognition systems primarily rely on supervised approaches, constrained by the scale of labeled datasets. To boost the system performance, researchers leverage large pretrained models such as WavLM to transfer learned high-level features to the downstream speaker recognition task. H…

Cited by 2SourceScholar
2024

SECap: Speech Emotion Captioning with Large Language Model

AAAI 2024technical

Speech emotions are crucial in human communication and are extensively used in fields like speech synthesis and natural language understanding. Most prior studies, such as speech emotion recognition, have categorized speech emotions into a fixed set of classes. Yet, emotions expressed in human spee…

2023

BAYES RISK CTC: CONTROLLABLE CTC ALIGNMENT IN SEQUENCE-TO-SEQUENCE TASKS

ICLR 2023poster

Sequence-to-Sequence (seq2seq) tasks transcribe the input sequence to a target sequence. The Connectionist Temporal Classification (CTC) criterion is widely used in multiple seq2seq tasks. Besides predicting the target sequence, a side product of CTC is to predict the alignment, which is the most pr…

Cited by 10SourcePDFScholar
2023

TSpeech-AI System Description to the 5th Deep Noise Suppression (DNS) Challenge

ICASSP 2023accepted

This report presents the development of Tencent AI Lab’s personalized speech enhancement system for the 2023 ICASSP Signal Processing Grand Challenge – deep noise suppression (DNS) challenge <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sup> , whic…

Cited by 0SourceScholar
2022

Audio-Visual Multi-Channel Speech Separation, Dereverberation and Recognition

ICASSP 2022accepted

Despite the rapid advance of automatic speech recognition (ASR) technologies, accurate recognition of cocktail party speech characterised by the interference from overlapping speakers, background noise and room reverberation remains a highly challenging task to date. Motivated by the invariance of v…

Cited by 0SourceScholar
2022

Consistent Training and Decoding for End-to-End Speech Recognition Using Lattice-Free MMI

ICASSP 2022accepted

Recently, End-to-End (E2E) frameworks have achieved remarkable results on various Automatic Speech Recognition (ASR) tasks. However, Lattice-Free Maximum Mutual Information (LF-MMI), as one of the discriminative training criteria that show superior performance in hybrid ASR systems, is rarely adopte…

Cited by 0SourceScholar
2022

Mixed Precision DNN Quantization for Overlapped Speech Separation and Recognition

ICASSP 2022accepted

Recognition of overlapped speech has been a highly challenging task to date. State-of-the-art multi-channel speech separation system are becoming increasingly complex and expensive for practical applications. To this end, low-bit neural network quantization provides a powerful solution to dramatical…

Cited by 0SourceScholar
2022

Multi-Channel Speaker Diarization Using Spatial Features for Meetings

ICASSP 2022accepted

Speaker identification for overlapped speech presents a great challenge for speaker diarization tasks in meeting scenarios. In order to overcome such challenges, several overlap-aware resegmentation methods based on deep learning have been integrated into speaker diarization systems. In this paper w…

Cited by 0SourceScholar
2021

A Comparative Study of Acoustic and Linguistic Features Classification for Alzheimer's Disease Detection

ICASSP 2021accepted

With the global population ageing rapidly, Alzheimer's disease (AD) is particularly prominent in older adults, which has an insidious onset followed by gradual, irreversible deterioration in cognitive domains (memory, communication, etc). Thus the detection of Alzheimer's disease is crucial for time…

Cited by 0SourceScholar
2021

A Joint Training Framework of Multi-Look Separator and Speaker Embedding Extractor for Overlapped Speech

ICASSP 2021accepted

In multi-talker cases, overlapped speech degrades the speaker verification (SV) performance dramatically. To tackle this challenging problem, speech separation with multi-channel techniques can be adopted to extract each speaker’s signals to improve the SV performance. In this paper, a joint trainin…

Cited by 0SourceScholar
2021

Bayesian Transformer Language Models for Speech Recognition

ICASSP 2021accepted

State-of-the-art neural language models (LMs) represented by Transformers are highly complex. Their use of fixed, deterministic parameter estimates fail to account for model uncertainty and lead to over-fitting and poor generalization when given limited training data. In order to address these issue…

Cited by 0SourceScholar
2021

Development of the Cuhk Elderly Speech Recognition System for Neurocognitive Disorder Detection Using the Dementiabank Corpus

ICASSP 2021accepted

Early diagnosis of Neurocognitive Disorder (NCD) is crucial in facilitating preventive care and timely treatment to delay further progression. This paper presents the development of a state-of-the-art automatic speech recognition (ASR) system built on the Dementia-Bank Pitt corpus for automatic NCD…

Cited by 54SourceScholar
2021

Mixed Precision Quantization of Transformer Language Models for Speech Recognition

ICASSP 2021accepted

State-of-the-art neural language models represented by Transformers are becoming increasingly complex and expensive for practical applications. Low-bit deep neural network quantization techniques provides a powerful solution to dramatically reduce their model size. Current low-bit quantization metho…

Cited by 0SourceScholar
2020

Adversarial Attacks on GMM I-Vector Based Speaker Verification Systems

ICASSP 2020accepted

This work investigates the vulnerability of Gaussian Mixture Model (GMM) i-vector based speaker verification systems to adversarial attacks, and the transferability of adversarial samples crafted from GMM i-vector based systems to x-vector based systems. In detail, we formulate the GMM i-vector syst…

Cited by 0SourceScholar
2020

Audio-Visual Recognition of Overlapped Speech for the LRS2 Dataset

ICASSP 2020accepted

Automatic recognition of overlapped speech remains a highly challenging task to date. Motivated by the bimodal nature of human speech perception, this paper investigates the use of audio-visual technologies for overlapped speech recognition. Three issues associated with the construction of audio-vis…

Cited by 82SourceScholar
2020

Dirichlet Graph Variational Autoencoder

NeurIPS 2020poster

Graph Neural Networks (GNN) and Variational Autoencoders (VAEs) have been widely used in modeling and generating graphs with latent factors. However there is no clear explanation of what these latent factors are and why they perform well. In this work, we present Dirichlet Graph Variational Autoenco…

Cited by 55SourcePDFScholar
2020

End-To-End Voice Conversion Via Cross-Modal Knowledge Distillation for Dysarthric Speech Reconstruction

ICASSP 2020accepted

Dysarthric speech reconstruction (DSR) is a challenging task due to difficulties in repairing unstable prosody and correcting imprecise articulation. Inspired by the success of sequence-to-sequence (seq2seq) based text-to-speech (TTS) synthesis and knowledge distillation (KD) techniques, this paper…

Cited by 0SourceScholar
2020

Low-bit Quantization of Recurrent Neural Network Language Models Using Alternating Direction Methods of Multipliers

ICASSP 2020accepted

The high memory consumption and computational costs of Recurrent neural network language models (RNNLMs) limit their wider application on resource constrained devices. In recent years, neural network quantization techniques that are capable of producing extremely low-bit compression, for example, bi…

Cited by 0SourceScholar
2019

Bayesian and Gaussian Process Neural Networks for Large Vocabulary Continuous Speech Recognition

ICASSP 2019accepted

The hidden activation functions inside deep neural networks (DNNs) play a vital role in learning high level discriminative features and controlling the information flows to track longer history. However, the fixed model parameters used in standard DNNs can lead to over-fitting and poor generalizatio…

Cited by 0SourceScholar
2019

End-to-end Code-switched TTS with Mix of Monolingual Recordings

ICASSP 2019accepted

State-of-the-art text-to-speech (TTS) synthesis models can produce monolingual speech with high intelligibility and naturalness. However, when the models are applied to synthesize code-switched (CS) speech, the performance declines seriously. Conventionally, developing a CS TTS system requires multi…

Cited by 0SourceScholar
2019

Gaussian Process Lstm Recurrent Neural Network Language Models for Speech Recognition

ICASSP 2019accepted

Recurrent neural network language models (RNNLMs) have shown superior performance across a range of speech recognition tasks. At the heart of all RNNLMs, the activation functions play a vital role to control the information flows and tracking longer history contexts that are useful for predicting th…

Cited by 0SourceScholar
2019

Recurrent Neural Network Language Model Training Using Natural Gradient

ICASSP 2019accepted

Recurrent neural network language models (RNNLMs) have become an increasing popular choice for state-of-the-art speech recognition systems. RNNLMs are normally trained by minimizing the cross entropy (CE) using the stochastic gradient descent (SGD) algorithm. However, the SGD method doesn't consider…

Cited by 0SourceScholar
2019

Speech Emotion Recognition Using Capsule Networks

ICASSP 2019accepted

Speech emotion recognition (SER) is a fundamental step towards fluent human-machine interaction. One challenging problem in SER is obtaining utterance-level feature representation for classification. Recent works on SER have made significant progress by using spectrogram features and introducing neu…

Cited by 0SourceScholar
2018

Limited-Memory BFGS Optimization of Recurrent Neural Network Language Models for Speech Recognition

ICASSP 2018accepted

Recurrent neural network language models (RNNLM) have become an increasingly popular choice for state-of-the-art speech recognition systems. RNNLMs are normally trained by minimizing the cross entropy (CE) using the stochastic gradient descent (SGD) algorithm. The SGD method only uses first-order de…

Cited by 0SourceScholar