← Search

Devang Naik

14 accepted papers

2025

An Efficient and Streaming Audio Visual Active Speaker Detection System

ICASSP 2025accepted

This paper delves into the challenging task of active speaker detection (asd), where the system needs to determine in real-time whether a person is speaking or not in a series of video frames. While previous works have made significant strides in improving network architectures and learning effectiv…

Cited by 0SourceScholar
2024

Flexible Keyword Spotting Based on Homogeneous Audio-Text Embedding

ICASSP 2024accepted

Spotting user-defined/flexible keywords represented in text frequently uses an expensive text encoder for joint analysis with an audio encoder in an embedding space, which can suffer from heterogeneous modality representation (i.e., large mismatch) and increased complexity. In this work, we propose…

Cited by 0SourceScholar
2024

Improving Vision-Inspired Keyword Spotting Using Dynamic Module Skipping in Streaming Conformer Encoder

ICASSP 2024accepted

Using a vision-inspired keyword spotting framework, we propose an architecture with input-dependent dynamic depth capable of processing streaming audio. Specifically, we extend a conformer encoder with trainable binary gates that allow us to dynamically skip network modules according to the input au…

Cited by 0SourceScholar
2024

KV-Runahead: Scalable Causal LLM Inference by Parallel Key-Value Cache Generation

ICML 2024poster

Large Language Model or LLM inference has two phases, the prompt (or prefill) phase to output the first token and the extension (or decoding) phase to the generate subsequent tokens. In this work, we propose an efficient parallelization scheme, KV-Runahead to accelerate the prompt phase. The key obs…

Cited by 3SourcePDFScholar
2023

HEiMDaL: Highly Efficient Method for Detection and Localization of Wake-Words

ICASSP 2023accepted

Streaming keyword spotting is a widely used solution for activating voice assistants. Methods based on Deep Neural Networks with Hidden Markov Model (DNN-HMM) have proven to be efficient and widely adopted in this space, primarily because of the ability to detect and identify the start and end of th…

Cited by 0SourceScholar
2023

I See What You Hear: A Vision-Inspired Method to Localize Words

ICASSP 2023accepted

This paper explores the possibility of using visual object detection techniques for word localization in speech data. Object detection has been thoroughly studied in the contemporary literature for visual data. Noting that an audio can be interpreted as a 1-dimensional image, object localization tec…

Cited by 0SourceScholar
2021

Knowledge Transfer for Efficient on-Device False Trigger Mitigation

ICASSP 2021accepted

In this paper, we address the task of determining whether a given utterance is directed towards a voice-enabled smart-assistant device or not. An undirected utterance is termed as a "false trigger" and false trigger mitigation (FTM) is essential for designing a privacy-centric non-intrusive smart as…

Cited by 0SourceScholar
2021

On The Role of Visual Cues in Audiovisual Speech Enhancement

ICASSP 2021accepted

We present an introspection of an audiovisual speech enhancement model. In particular, we focus on interpreting how a neural audiovisual speech enhancement model uses visual cues to improve the quality of the target speech signal. We show that visual cues provide not only high-level information abou…

Cited by 0SourceScholar
2021

Optimize What Matters: Training DNN-Hmm Keyword Spotting Model Using End Metric

ICASSP 2021accepted

Deep Neural Network–Hidden Markov Model (DNN-HMM) based methods have been successfully used for many always-on keyword spotting algorithms that detect a wake word to trigger a device. The DNN predicts the state probabilities of a given speech frame, while HMM decoder combines the DNN predictions of…

Cited by 0SourceScholar
2020

Detecting Emotion Primitives from Speech and Their Use in Discerning Categorical Emotions

ICASSP 2020accepted

Emotion plays an essential role in human-to-human communication, enabling us to convey feelings such as happiness, frustration, and sincerity. While modern speech technologies rely heavily on speech recognition and natural language understanding for speech content understanding, the investigation of…

Cited by 0SourceScholar
2020

Lattice-Based Improvements for Voice Triggering Using Graph Neural Networks

ICASSP 2020accepted

Voice-triggered smart assistants often rely on detection of a trigger-phrase before they start listening for the user request. Mitigation of false triggers is an important aspect of building a privacy-centric non-intrusive smart assistant. In this paper, we address the task of false trigger mitigati…

Cited by 0SourceScholar
2020

Multi-Task Learning for Speaker Verification and Voice Trigger Detection

ICASSP 2020accepted

Automatic speech transcription and speaker recognition are usually treated as separate tasks even though they are interdependent. In this study, we investigate training a single network to perform both tasks jointly. We train the network in a supervised multi-task learning setup, where the speech tr…

Cited by 0SourceScholar