← Search

Pranay Dighe

10 accepted papers

2024

Leveraging Large Language Models for Exploiting ASR Uncertainty

ICASSP 2024accepted

While large language models excel in a variety of natural language processing (NLP) tasks, to perform well on spoken language understanding (SLU) tasks, they must either rely on off-the-shelf automatic speech recognition (ASR) systems for transcription, or be equipped with an in-built speech modalit…

Cited by 0SourceScholar
2024

Modality Drop-Out for Multimodal Device Directed Speech Detection Using Verbal and Non-Verbal Features

ICASSP 2024accepted

Device-directed speech detection (DDSD) is the binary classification task of distinguishing between queries directed at a voice assistant versus side conversation or background speech. State-of-the-art DDSD systems use verbal cues, e.g acoustic, text and/or automatic speech recognition system (ASR)…

Cited by 0SourceScholar
2023

Audio-to-Intent Using Acoustic-Textual Subword Representations from End-to-End ASR

ICASSP 2023accepted

Accurate prediction of the user intent to interact with a voice assistant (VA) on a device (e.g. a smartphone) is critical for achieving naturalistic, engaging, and privacy-centric interactions with the VA. To this end, we present a novel approach to predict the user intention (whether the user is s…

Cited by 0SourceScholar
2023

Less Is More: A Unified Architecture for Device-Directed Speech Detection with Multiple Invocation Types

ICASSP 2023accepted

Suppressing unintended invocation of the device because of the speech that sounds like wake-word, or accidental button presses, is critical for a good user experience, and is referred to as False-Trigger-Mitigation (FTM). In case of multiple invocation options, the traditional approach to FTM is to…

Cited by 0SourceScholar
2022

Streaming on-Device Detection of Device Directed Speech from Voice and Touch-Based Invocation

ICASSP 2022accepted

When interacting with smart devices such as mobile-phones or wearables, the user typically invokes a virtual assistant (VA) by saying a keyword or by pressing a button on the device. However, in many cases, the VA can accidentally be invoked by the keyword-like speech or accidental button press, whi…

Cited by 0SourceScholar
2021

Knowledge Transfer for Efficient on-Device False Trigger Mitigation

ICASSP 2021accepted

In this paper, we address the task of determining whether a given utterance is directed towards a voice-enabled smart-assistant device or not. An undirected utterance is termed as a "false trigger" and false trigger mitigation (FTM) is essential for designing a privacy-centric non-intrusive smart as…

Cited by 0SourceScholar
2020

Lattice-Based Improvements for Voice Triggering Using Graph Neural Networks

ICASSP 2020accepted

Voice-triggered smart assistants often rely on detection of a trigger-phrase before they start listening for the user request. Mitigation of false triggers is an important aspect of building a privacy-centric non-intrusive smart assistant. In this paper, we address the task of false trigger mitigati…

Cited by 0SourceScholar
2016

Exploiting low-dimensional structures to enhance DNN based acoustic modeling in speech recognition

ICASSP 2016accepted

We propose to model the acoustic space of deep neural network (DNN) class-conditional posterior probabilities as a union of low-dimensional subspaces. To that end, the training posteriors are used for dictionary learning and sparse coding. Sparse representation of the test posteriors using this dict…

Cited by 0SourceScholar