← Search

Dhananjaya Gowda

6 accepted papers

2024

Data Driven Grapheme-to-Phoneme Representations for a Lexicon-Free Text-to-Speech

ICASSP 2024accepted

Grapheme-to-Phoneme (G2P) is an essential first step in any modern, high-quality Text-to-Speech (TTS) system. Most of the current G2P systems rely on carefully hand-crafted lexicons developed by experts. This poses a two-fold problem. Firstly, the lexicons are generated using a fixed phoneme set, us…

Cited by 0SourceScholar
2023

A Transformer-Based E2E SLU Model for Improved Semantic Parsing

ICASSP 2023accepted

Spoken Language Understanding (SLU) is an essential part of voice and speech assistant tools. End-to-End (E2E) SLU models attempt to automatically extract semantic meanings from the speech signal without the need for an intermediate transcription of speech. However, SLU is a challenging task mainly…

Cited by 0SourceScholar
2023

Self-Supervised Accent Learning for Under-Resourced Accents Using Native Language Data

ICASSP 2023accepted

In this paper, we propose a novel method to improve the accuracy of an English speech recognizer for a target accent using the corresponding native language data. Collecting labeled data for all accents of English to train an end-to-end neural speech recognizer for English is a difficult and expensi…

Cited by 0SourceScholar
2021

Neural Utterance Confidence Measure for RNN-Transducers and Two Pass Models

ICASSP 2021accepted

In this paper, we propose methods to compute confidence score on the predictions made by an end-to-end speech recognition model in a 2-pass framework. We use RNN-Transducer for a streaming model, and an attention-based decoder for the second pass model. We use neural technique to compute the confide…

Cited by 0SourceScholar
2021

Streaming End-to-End Speech Recognition with Jointly Trained Neural Feature Enhancement

ICASSP 2021accepted

In this paper, we present a streaming end-to-end speech recognition model based on Monotonic Chunkwise Attention (MoCha) jointly trained with enhancement layers. Even though the MoCha attention enables streaming speech recognition with recognition accuracy comparable to a full attention-based approa…

Cited by 0SourceScholar