← Search

Yi-Cheng Wang

5 accepted papers

2025

ConPCO: Preserving Phoneme Characteristics For Automatic Pronunciation Assessment Leveraging Contrastive Ordinal Regularization

ICASSP 2025accepted

Automatic pronunciation assessment (APA) manages to evaluate the pronunciation proficiency of a second language (L2) learner in a target language. Existing efforts typically draw on regression models for proficiency score prediction, wherein the models are trained to estimate target values without e…

Cited by 0SourceScholar
2024

An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder Disentanglement

ICASSP 2024accepted

With the massive developments of end-to-end (E2E) neural networks, recent years have witnessed unprecedented breakthroughs in automatic speech recognition (ASR). However, the code-switching phenomenon remains a major obstacle that hinders ASR from perfection, as the lack of labeled data and the vari…

Cited by 0SourceScholar
2024

An Effective Pronunciation Assessment Approach Leveraging Hierarchical Transformers and Pre-training Strategies

ACL 2024long

Automatic pronunciation assessment (APA) manages to quantify a second language (L2) learner’s pronunciation proficiency in a target language by providing fine-grained feedback with multiple pronunciation aspect scores at various linguistic levels. Most existing efforts on APA typically parallelize t…

2024

DANCER: Entity Description Augmented Named Entity Corrector for Automatic Speech Recognition

COLING 2024main

End-to-end automatic speech recognition (E2E ASR) systems often suffer from mistranscription of domain-specific phrases, such as named entities, sometimes leading to catastrophic failures in downstream tasks. A family of fast and lightweight named entity correction (NEC) models for ASR have recently…

2023

Effective Graph-Based Modeling of Articulation Traits for Mispronunciation Detection and Diagnosis

ICASSP 2023accepted

Mispronunciation detection and diagnosis (MDD) manages to pinpoint phone-level erroneous pronunciation segmentations and provide instant and informative diagnostic feedback to L2 (second-language) learners. Among the various modeling paradigms for MDD, dictation-based neural methods have recently be…

Cited by 0SourceScholar