← Search

Yuejiao Wang

5 accepted papers

2025

Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC

ICASSP 2025accepted

Multi-talker speech recognition (MTASR) faces unique challenges in disentangling and transcribing overlapping speech. To address these challenges, this paper investigates the role of Connectionist Temporal Classification (CTC) in speaker disentanglement when incorporated with Serialized Output Train…

Cited by 0SourceScholar
2025

Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions

ICASSP 2025accepted

Recent advancements in large language models (LLMs) have revolutionized various domains, bringing significant progress and new opportunities. Despite progress in speech-related tasks, LLMs have not been sufficiently explored in multi-talker scenarios. In this work, we present a pioneering effort to…

Cited by 0SourceScholar
2024

Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction

ICASSP 2024accepted

Dysarthric speech reconstruction (DSR) aims to transform dysarthric speech into normal speech by improving the intelligibility and naturalness. This is a challenging task especially for patients with severe dysarthria and speaking in complex, noisy acoustic environments. To address these challenges,…

Cited by 0SourceScholar
2024

UNIT-DSR: Dysarthric Speech Reconstruction System Using Speech Unit Normalization

ICASSP 2024accepted

Dysarthric speech reconstruction (DSR) systems aim to automatically convert dysarthric speech into normal-sounding speech. The technology eases communication with speakers affected by the neuromotor disorder and enhances their social inclusion. NED-based (Neural Encoder-Decoder) systems have signifi…

Cited by 0SourceScholar
2023

A Sidecar Separator Can Convert A Single-Talker Speech Recognition System to A Multi-Talker One

ICASSP 2023accepted

Although automatic speech recognition (ASR) can perform well in common non-overlapping environments, sustaining performance in multi-talker overlapping speech recognition remains challenging. Recent research revealed that ASR model’s encoder captures different levels of information with different la…

Cited by 0SourceScholar