← Search

Vineet Garg

5 accepted papers

2024

Leveraging Large Language Models for Exploiting ASR Uncertainty

ICASSP 2024accepted

While large language models excel in a variety of natural language processing (NLP) tasks, to perform well on spoken language understanding (SLU) tasks, they must either rely on off-the-shelf automatic speech recognition (ASR) systems for transcription, or be equipped with an in-built speech modalit…

Cited by 0SourceScholar
2024

Streaming Anchor Loss: Augmenting Supervision with Temporal Significance

ICASSP 2024accepted

Streaming neural network models for fast frame-wise responses to various speech and sensory signals are widely adopted on resource-constrained platforms. Hence, increasing the learning capacity of such streaming models (i.e., by adding more parameters) to improve the predictive power may not be viab…

Cited by 2SourceScholar
2023

Less Is More: A Unified Architecture for Device-Directed Speech Detection with Multiple Invocation Types

ICASSP 2023accepted

Suppressing unintended invocation of the device because of the speech that sounds like wake-word, or accidental button presses, is critical for a good user experience, and is referred to as False-Trigger-Mitigation (FTM). In case of multiple invocation options, the traditional approach to FTM is to…

Cited by 0SourceScholar
2022

Streaming on-Device Detection of Device Directed Speech from Voice and Touch-Based Invocation

ICASSP 2022accepted

When interacting with smart devices such as mobile-phones or wearables, the user typically invokes a virtual assistant (VA) by saying a keyword or by pressing a button on the device. However, in many cases, the VA can accidentally be invoked by the keyword-like speech or accidental button press, whi…

Cited by 0SourceScholar
2021

Progressive Voice Trigger Detection: Accuracy vs Latency

ICASSP 2021accepted

We present an architecture for voice trigger detection for virtual assistants. The main idea in this work is to exploit information in words that immediately follow the trigger phrase. We first demonstrate that by including more audio context after a detected trigger phrase, we can indeed get a more…

Cited by 0SourceScholar