← Search

Sachin Kajarekar

8 accepted papers

2023

Less Is More: A Unified Architecture for Device-Directed Speech Detection with Multiple Invocation Types

ICASSP 2023accepted

Suppressing unintended invocation of the device because of the speech that sounds like wake-word, or accidental button presses, is critical for a good user experience, and is referred to as False-Trigger-Mitigation (FTM). In case of multiple invocation options, the traditional approach to FTM is to…

Cited by 0SourceScholar
2022

Streaming on-Device Detection of Device Directed Speech from Voice and Touch-Based Invocation

ICASSP 2022accepted

When interacting with smart devices such as mobile-phones or wearables, the user typically invokes a virtual assistant (VA) by saying a keyword or by pressing a button on the device. However, in many cases, the VA can accidentally be invoked by the keyword-like speech or accidental button press, whi…

Cited by 0SourceScholar
2021

Knowledge Transfer for Efficient on-Device False Trigger Mitigation

ICASSP 2021accepted

In this paper, we address the task of determining whether a given utterance is directed towards a voice-enabled smart-assistant device or not. An undirected utterance is termed as a "false trigger" and false trigger mitigation (FTM) is essential for designing a privacy-centric non-intrusive smart as…

Cited by 0SourceScholar
2021

On The Role of Visual Cues in Audiovisual Speech Enhancement

ICASSP 2021accepted

We present an introspection of an audiovisual speech enhancement model. In particular, we focus on interpreting how a neural audiovisual speech enhancement model uses visual cues to improve the quality of the target speech signal. We show that visual cues provide not only high-level information abou…

Cited by 0SourceScholar
2021

SEP-28k: A Dataset for Stuttering Event Detection from Podcasts with People Who Stutter

ICASSP 2021accepted

The ability to automatically detect stuttering events in speech could help speech pathologists track an individual’s fluency over time or help improve speech recognition systems for people with atypical speech patterns. Despite increasing interest in this area, existing public datasets are too small…

Cited by 0SourceScholar
2020

Detecting Emotion Primitives from Speech and Their Use in Discerning Categorical Emotions

ICASSP 2020accepted

Emotion plays an essential role in human-to-human communication, enabling us to convey feelings such as happiness, frustration, and sincerity. While modern speech technologies rely heavily on speech recognition and natural language understanding for speech content understanding, the investigation of…

Cited by 17SourceScholar
2020

Multi-Task Learning for Speaker Verification and Voice Trigger Detection

ICASSP 2020accepted

Automatic speech transcription and speaker recognition are usually treated as separate tasks even though they are interdependent. In this study, we investigate training a single network to perform both tasks jointly. We train the network in a supervised multi-task learning setup, where the speech tr…

Cited by 0SourceScholar
2018

Generalised Discriminative Transform via Curriculum Learning for Speaker Recognition

ICASSP 2018accepted

In this paper we introduce a speaker verification system deployed on mobile devices that can be used to personalise a keyword spotter. We describe a baseline DNN system that maps an utterance to a speaker embedding, which is used to measure speaker differences via cosine similarity. We then introduc…

Cited by 0SourceScholar