← Search

Petar S. Aleksic

5 accepted papers

2021

Improving Entity Recall in Automatic Speech Recognition with Neural Embeddings

ICASSP 2021accepted

Automatic speech recognition (ASR) systems often have difficulty recognizing long-tail entities such as contact names and local restaurant names, which usually do not occur, or occur infrequently, in the system’s training data. In this work, we present a method which uses learned text embeddings and…

Cited by 7SourceScholar
2020

Incorporating Written Domain Numeric Grammars into End-To-End Contextual Speech Recognition Systems for Improved Recognition of Numeric Sequences

ICASSP 2020accepted

Accurate recognition of numeric sequences is crucial for many contextual speech recognition applications. For example, a user might create a calendar event and be prompted by a virtual assistant for the time, date, and duration of the event. We propose a modular and scalable solution for improved re…

Cited by 0SourceScholar
2020

Multistate Encoding with End-To-End Speech RNN Transducer Network

ICASSP 2020accepted

Recurrent Neural Network Transducer (RNN-T) models [1] for automatic speech recognition (ASR) provide high accuracy speech recognition. Such end-to-end (E2E) models combine acoustic, pronunciation and language models (AM, PM, LM) of a conventional ASR system into a single neural network, dramaticall…

Cited by 0SourceScholar
2018

Cross-Lingual Phoneme Mapping for Language Robust Contextual Speech Recognition

ICASSP 2018accepted

Standard automatic speech recognition (ASR) systems are increasingly expected to recognize foreign entities, yet doing so while preserving accuracy on native words remains a challenge. We describe a novel approach for recognizing foreign words by injecting them with appropriate pronunciations into t…

Cited by 0SourceScholar
2015

Improved recognition of contact names in voice commands

ICASSP 2015accepted

The recognition of contact names in mobile-device voice commands is a challenging problem. Some of the difficulties include potentially infinite vocabularies, low probability of contact tokens in the language model (LM), increased false triggering of contact voice commands when none are spoken, and…

Cited by 0SourceScholar