← Search

Ernest Pusateri

4 accepted papers

2025

Contextualization of ASR with LLM using phonetic retrieval-based augmentation

ICASSP 2025accepted

Large language models (LLMs) have shown superb capability of modeling multimodal signals including audio and text, allowing the model to generate spoken or textual response given a speech input. However, it remains a challenge for the model to recognize personal named entities, such as contacts in a…

Cited by 0SourceScholar
2025

Retrieval Augmented Correction of Named Entity Speech Recognition Errors

ICASSP 2025accepted

In recent years, end-to-end automatic speech recognition (ASR) systems have proven themselves remarkably accurate and performant, but these systems still have a significant error rate for entity names which appear infrequently in their training data. In parallel to the rise of end-to-end ASR systems…

Cited by 0SourceScholar
2024

Personalization of CTC-Based End-to-End Speech Recognition Using Pronunciation-Driven Subword Tokenization

ICASSP 2024accepted

Recent advances in deep learning and automatic speech recognition have improved the accuracy of end-to-end speech recognition systems, but recognition of personal content such as contact names remains a challenge. In this work, we describe our personalization solution for an end-to-end speech recogn…

Cited by 0SourceScholar
2021

Error-Driven Pruning of Language Models for Virtual Assistants

ICASSP 2021accepted

Language models (LMs) for virtual assistants (VAs) are typically trained on large amounts of data, resulting in prohibitively large models which require excessive memory and/or cannot be used to serve user requests in real-time. Entropy pruning results in smaller models but with significant degradat…

Cited by 0SourceScholar