← Search

Qibing Bai

3 accepted papers

2026

CosyAccent: Duration-Controllable Accent Normalization Using Source-Synthesis Training Data

ICASSP 2026poster

Accent normalization (AN) systems often struggle with unnatural outputs and undesired content distortion, stemming from both suboptimal training data and rigid duration modeling. In this paper, we propose a "source-synthesis" methodology for training data construction. By generating source L2 speech…

Cited by 0SourcePDFScholar
2024

LLaST: Improved End-to-end Speech Translation System Leveraged by Large Language Models

ACL 2024findings

We introduces ***LLaST***, a framework for building high-performance Large Language model based Speech-to-text Translation systems. We address the limitations of end-to-end speech translation (E2E ST) models by exploring model architecture design and optimization techniques tailored for LLMs. Our ap…

2024

Leveraging in-the-wild Data for Effective Self-supervised Pretraining in Speaker Recognition

ICASSP 2024accepted

Current speaker recognition systems primarily rely on supervised approaches, constrained by the scale of labeled datasets. To boost the system performance, researchers leverage large pretrained models such as WavLM to transfer learned high-level features to the downstream speaker recognition task. H…

Cited by 2SourceScholar