← Search

Guangpu Huang

4 accepted papers

2024

Investigating the Clusters Discovered By Pre-Trained AV-HuBERT

ICASSP 2024accepted

Self-supervised models, such as HuBERT and its audio-visual version AV-HuBERT, have demonstrated excellent performance on various tasks. The main factor for their success is the pre-training procedure, which requires only raw data without human transcription. During the self-supervised pre-training…

Cited by 0SourceScholar
2017

An investigation into language model data augmentation for low-resourced STT and KWS

ICASSP 2017accepted

This paper reports on investigations using two techniques for language model text data augmentation for low-resourced automatic speech recognition and keyword search. Lowresourced languages are characterized by limited training materials, which typically results in high out-of-vocabulary (OOV) rates…

Cited by 0SourceScholar
2017

Effective keyword search for low-resourced conversational speech

ICASSP 2017accepted

In this paper we aim to enhance keyword search for conversational telephone speech under low-resourced conditions. Two techniques to improve the detection of out-of-vocabulary keywords are assessed in this study: using extra text resources to augment the lexicon and language model, and via subword u…

Cited by 0SourceScholar
2016

Machine translation based data augmentation for Cantonese keyword spotting

ICASSP 2016accepted

This paper presents a method to improve a language model for a limited-resourced language using statistical machine translation from a related language to generate data for the target language. In this work, the machine translation model is trained on a corpus of parallel Mandarin-Cantonese subtitle…

Cited by 0SourceScholar