IMPROVING CONTEXTUAL ASR VIA MULTI-GRAINED FUSION WITH LARGE LANGUAGE MODELS
While end-to-end Automatic Speech Recognition (ASR) models have shown impressive performance in transcribing general speech, they often struggle to accurately recognize contextually relevant keywords, such as proper nouns or user-specific entities. Previous approaches have explored leveraging keywor…