← Search

Mingbin Xu

4 accepted papers

2025

Contextualization of ASR with LLM using phonetic retrieval-based augmentation

ICASSP 2025accepted

Large language models (LLMs) have shown superb capability of modeling multimodal signals including audio and text, allowing the model to generate spoken or textual response given a speech input. However, it remains a challenge for the model to recognize personal named entities, such as contacts in a…

Cited by 0SourceScholar
2024

Conformer-Based Speech Recognition On Extreme Edge-Computing Devices

NAACL 2024industry

With increasingly more powerful compute capabilities and resources in today’s devices, traditionally compute-intensive automatic speech recognition (ASR) has been moving from the cloud to devices to better protect user privacy. However, it is still challenging to implement on-device ASR on resource-…

Cited by 4SourcePDFScholar
2024

Personalization of CTC-Based End-to-End Speech Recognition Using Pronunciation-Driven Subword Tokenization

ICASSP 2024accepted

Recent advances in deep learning and automatic speech recognition have improved the accuracy of end-to-end speech recognition systems, but recognition of personal content such as contact names remains a challenge. In this work, we describe our personalization solution for an end-to-end speech recogn…

Cited by 0SourceScholar
2023

Training Large-Vocabulary Neural Language Models by Private Federated Learning for Resource-Constrained Devices

ICASSP 2023accepted

Federated Learning (FL) is a technique to train models on distributed edge devices with local data samples. Differential Privacy (DP) can be applied with FL to provide a formal privacy guarantee for sensitive data on device. Our goal is to train a large neural network language model (NNLM) on comput…

Cited by 0SourceScholar